跳到主内容
精选85elvis论文研究

论文揭示:Agent 基准分数受测试框架影响远超模型本身

Great paper on why agent leaderboard comparisons are hard to trust.

原文
推荐理由

做 Agent 评测或参考排行榜选模型的同学必看,这篇论文用数据证明框架对分数的影响远超模型本身,建议读原文并关注 Harness Card 提案,避免被榜单误导。

Great paper on why agent leaderboard comparisons are hard to trust.

关于为何智能体排行榜比较难以信赖的精彩论文。

It's on the hot topic of how much of an agent benchmark score actually belongs to the harness.

它聚焦于热门话题:智能体基准测试分数中有多少实际上归功于“脚手架”(harness)。

The harness is the layer between the model and the task. It builds the context the model sees, mediates tool calls, validates outputs, and decides when to retry or stop. Every score comes out of a model and a harness together, but only the model gets reported.

脚手架是模型与任务之间的层。它构建模型所见的上下文,调解工具调用,验证输出,并决定何时重试或停止。每个分数都来自模型与脚手架的组合,但只有模型被报告出来。

The authors ran a controlled grid to measure this. Three frontier models, three harness configurations, 100 tasks from SWE-bench Verified, with task order, execution environment, step budget, and evaluation script all held fixed.

作者运行了一个受控网格来测量这一点。三个前沿模型,三种脚手架配置,来自SWE-bench Verified的100个任务,任务顺序、执行环境、步骤预算和评估脚本均保持不变。

Swapping the harness moved GLM-5.1 by 13.0 points. Swapping the model inside a fixed harness moved scores by 3.0, 2.5, and 5.0 points.

更换脚手架使GLM-5.1的分数变动了13.0点。在固定脚手架内更换模型,分数变动分别为3.0、2.5和5.0点。

Harness-induced variance came out 7.8x larger than model-induced variance, and 6 of 9 model-pair comparisons flipped their ranking depending on which harness ran.

由脚手架引起的方差比模型引起的方差大7.8倍,9对模型比较中有6对的排名因运行脚手架的不同而翻转。

Public leaderboards show the same thing. On SWE-bench Verified Mini, HAL reports a 34 point swing for Claude Sonnet 4.5 across scaffolds and nearly 48 points for o4-mini.

公开排行榜也显示了同样的情况。在SWE-bench Verified Mini上,HAL报告Claude Sonnet 4.5在不同脚手架上的分数波动达34点,o4-mini则接近48点。

They propose a Harness Card, a structured disclosure across seven layers, so you can tell whether a score gap came from the model, the harness, or the interaction.

他们提出了“脚手架卡片”(Harness Card),一种跨七个层的结构化披露方式,以便你能判断分数差距是来自模型、脚手架还是两者的交互。

Paper: https://arxiv.org/abs/2605.23950

论文:https://arxiv.org/abs/2605.23950

Track more trending AI papers in our academy: https://academy.dair.ai/

在我们的学院追踪更多热门AI论文:https://academy.dair.ai/

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

关联讨论

同一事件的更多信源

相似阅读

另一事件,读法相近