跳到主内容
精选88Rohan Paul论文研究

论文指出长程Agent评测中Harness影响超模型,建议披露基准环境

For long-horizon agents, this paper argues the harness can matter more than the…

原文
推荐理由

Agent评测方法论的重要纠偏,揭示了Benchmark结果可能受环境干扰而非模型能力,做Agent研发的同学务必关注此结论以优化评估体系。

For long-horizon agents, this paper argues the harness can matter more than the model, so benchmark scores should not be compared without disclosing or controlling the harness.

对于长视界智能体,本文认为 harness(运行环境/框架)的影响可能比模型更大,因此在未披露或控制 harness 的情况下,不应比较基准分数。

In a controlled 3-model × 3-harness study on 100 SWE-bench Verified tasks, average harness-induced variance was 7.80X model-induced variance.

在针对 100 个 SWE-bench Verified 任务的受控 3 模型 × 3 harness 研究中,由 harness 引起的平均方差是模型引起方差的 7.80 倍。

Keeping the model fixed, moving from the minimal to full harness changed pass@1 by 8.5 to 13.0 percentage points. Keeping the harness fixed, switching models changed it by only 2.5 to 5.0 points.

保持模型固定,从最小化 harness 切换到完整 harness 使 pass@1 指标变化了 8.5 到 13.0 个百分点;保持 harness 固定,切换模型仅使其变化了 2.5 到 5.0 个点。

The ranking itself was unstable too: 6 of 9 model-pair/harness-pair comparisons reversed under another harness.

排名本身也不稳定:在另一种 harness 下,9 组模型对/harness 对比较中有 6 组的排序发生了反转。

– arxiv. org/abs/2605.23950

– arxiv.org/abs/2605.23950

Title: "Stop Comparing LLM Agents Without Disclosing the Harness"

标题:《未在披露 Harness 的情况下停止比较 LLM 智能体》

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

相似阅读

另一事件,读法相近