跳到主内容
@wquguru
精选75Rohan Paul论文研究

SWE-bench Science:编码智能体更擅长修表象而非根因

Coding agents fix the symptom you show them far more often than the defect under…

原文
发到 X

Coding agents fix the symptom you show them far more often than the defect underneath.

编码代理修复你展示给它们的症状,远比修复底层缺陷更频繁。

Agents are good at making a visible failure disappear, but this paper finds them much weaker at restoring the science underneath, and supplied domain knowledge does not reliably help.

代理擅长让可见的失败消失,但本文发现它们在恢复底层科学方面要弱得多,并且提供的领域知识并不总是可靠地帮助。

SWE-bench Science gives agents real defects from open scientific repositories and scores public tests they can iterate against separately from private tests they never see.

SWE-bench Science 给代理提供来自开放科学仓库的真实缺陷,并评分它们可以迭代的公共测试,与它们从未见过的私有测试分开。

The best configuration, Claude Code with Opus-5, clears 96.64% of public tests but reaches 47.90% Pass@1, which requires every hidden test to pass.

最佳配置,Claude Code 与 Opus-5,清除了 96.64% 的公共测试,但达到了 47.90% 的 Pass@1,这要求所有隐藏测试都通过。

So the limit looks like verification rather than knowledge: an agent that cannot check supplied science against executable evidence can anchor on the explanation instead of testing it.

因此,限制看起来是验证而非知识:一个无法对照可执行证据检查所提供科学的代理,可能会锚定在解释上而不是测试它。

– arxiv. org/abs/2608.19799

– arxiv. org/abs/2608.19799

Title: "SWE-bench Science: Can Coding Agents Resolve Engineering Tasks in Science?"

标题:“SWE-bench Science:编码代理能否解决科学中的工程任务?”

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

相似阅读

另一事件,读法相近