斯坦福论文:Agent框架对模型性能影响显著,DuMateBench基准测试发布
New Stanford and other top research lab paper shows that the framework around an…
Agent开发者必读,这篇论文用真实数据量化了框架对性能的边际贡献,Opus-4.8在不同框架下近30分的差距极具参考价值,建议重新审视你的Agent架构选型。
New Stanford and other top research lab paper shows that the framework around an LLM can change agent performance dramatically.
斯坦福大学及其他顶级研究实验室的新论文表明,围绕大语言模型(LLM)构建的框架可以显著改变智能体的性能。
A strong LLM does not guarantee a strong agent
强大的 LLM 并不保证能造就强大的智能体。
Their DuMateBench benchmark uses 200 tasks rebuilt from real user sessions. They mix things agents actually do together: coding, web research, document work, and content creation. The environment also includes missing dependencies, flaky networks, and distracting files.
他们的 DuMateBench 基准测试使用了从真实用户会话中重建的 200 个任务。 他们将智能体实际执行的任务混合在一起:编码、网络研究、文档处理和内容创作。环境还包含缺失的依赖项、不稳定的网络以及干扰性的文件。
The clearest result is how much the same model changes across agent frameworks. With Opus-4.8, the final score ranges from 0.5821 with OpenClaw to 0.8548 with DuMate, a 27.27 percentage-point gap.
最清晰的结果是同一模型在不同智能体框架下的表现差异。使用 Opus-4.8 时,最终得分在 OpenClaw 的 0.5821 到 DuMate 的 0.8548 之间波动,差距达 27.27 个百分点。
– arxiv. org/abs/2608.26546
– arxiv.org/abs/2608.26546
Title: "DuMateBench: Evaluating Autonomous Agents in Complex Real-World Workflows"
标题:《DuMateBench:评估复杂现实工作流中的自主智能体》
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力