SWE-bench Science:编码智能体更擅长修表象而非根因
Coding agents fix the symptom you show them far more often than the defect under…
Coding agents fix the symptom you show them far more often than the defect underneath.
编码代理修复你展示给它们的症状,远比修复底层缺陷更频繁。
Agents are good at making a visible failure disappear, but this paper finds them much weaker at restoring the science underneath, and supplied domain knowledge does not reliably help.
代理擅长让可见的失败消失,但本文发现它们在恢复底层科学方面要弱得多,并且提供的领域知识并不总是可靠地帮助。
SWE-bench Science gives agents real defects from open scientific repositories and scores public tests they can iterate against separately from private tests they never see.
SWE-bench Science 给代理提供来自开放科学仓库的真实缺陷,并评分它们可以迭代的公共测试,与它们从未见过的私有测试分开。
The best configuration, Claude Code with Opus-5, clears 96.64% of public tests but reaches 47.90% Pass@1, which requires every hidden test to pass.
最佳配置,Claude Code 与 Opus-5,清除了 96.64% 的公共测试,但达到了 47.90% 的 Pass@1,这要求所有隐藏测试都通过。
So the limit looks like verification rather than knowledge: an agent that cannot check supplied science against executable evidence can anchor on the explanation instead of testing it.
因此,限制看起来是验证而非知识:一个无法对照可执行证据检查所提供科学的代理,可能会锚定在解释上而不是测试它。
– arxiv. org/abs/2608.19799
– arxiv. org/abs/2608.19799
Title: "SWE-bench Science: Can Coding Agents Resolve Engineering Tasks in Science?"
标题:“SWE-bench Science:编码代理能否解决科学中的工程任务?”
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力