论文指出长程Agent评测中Harness影响超模型,建议披露基准环境
For long-horizon agents, this paper argues the harness can matter more than the…
Agent评测方法论的重要纠偏,揭示了Benchmark结果可能受环境干扰而非模型能力,做Agent研发的同学务必关注此结论以优化评估体系。
For long-horizon agents, this paper argues the harness can matter more than the model, so benchmark scores should not be compared without disclosing or controlling the harness.
对于长视界智能体,本文认为 harness(运行环境/框架)的影响可能比模型更大,因此在未披露或控制 harness 的情况下,不应比较基准分数。
In a controlled 3-model × 3-harness study on 100 SWE-bench Verified tasks, average harness-induced variance was 7.80X model-induced variance.
在针对 100 个 SWE-bench Verified 任务的受控 3 模型 × 3 harness 研究中,由 harness 引起的平均方差是模型引起方差的 7.80 倍。
Keeping the model fixed, moving from the minimal to full harness changed pass@1 by 8.5 to 13.0 percentage points. Keeping the harness fixed, switching models changed it by only 2.5 to 5.0 points.
保持模型固定,从最小化 harness 切换到完整 harness 使 pass@1 指标变化了 8.5 到 13.0 个百分点;保持 harness 固定,切换模型仅使其变化了 2.5 到 5.0 个点。
The ranking itself was unstable too: 6 of 9 model-pair/harness-pair comparisons reversed under another harness.
排名本身也不稳定:在另一种 harness 下,9 组模型对/harness 对比较中有 6 组的排序发生了反转。
– arxiv. org/abs/2605.23950
– arxiv.org/abs/2605.23950
Title: "Stop Comparing LLM Agents Without Disclosing the Harness"
标题:《未在披露 Harness 的情况下停止比较 LLM 智能体》
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力