跳到主内容
精选75elvis论文研究

27B智能体在论文复现任务上超越Claude Opus 4.8与GPT-5.5

A 27B agent just beat Claude Opus 4.8 and GPT-5.5 on held-out research replicati…

原文

A 27B agent just beat Claude Opus 4.8 and GPT-5.5 on held-out research replication.

一个27B的智能体在保留的研究复现任务上击败了Claude Opus 4.8和GPT-5.5。

Replica turns paper replication into a scalable RL task space. Replicating a paper forces the same hypothesis-driven exploration as open research, and it surfaces details the original authors left underspecified.

Replica将论文复现转化为一个可扩展的强化学习任务空间。复现一篇论文迫使进行与开放研究相同的假设驱动探索,并揭示了原作者未详细说明的细节。

The reward signal comes from an auto-generated rubric judge that runs low-noise and agrees with human assessment of replication quality.

奖励信号来自一个自动生成的评分标准评判器,该评判器运行低噪声,并与人类对复现质量的评估一致。

Faraday, the resulting 27B agent, calls coding agents as tools. Rollout analysis shows it takes a more scientifically principled approach rather than gaming the rubric.

由此产生的27B智能体Faraday将编码智能体作为工具调用。展开分析显示,它采取了更科学严谨的方法,而不是玩弄评分标准。

The authors argue this points toward long-horizon scientific capability trained into weights, without requiring complex harnesses.

作者认为,这指向了无需复杂框架即可在权重中训练出的长周期科学能力。

Paper: https://arxiv.org/abs/2608.13331

论文:https://arxiv.org/abs/2608.13331

Track more trending AI papers in our academy: https://academy.dair.ai/

在我们的学院中追踪更多热门AI论文:https://academy.dair.ai/

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

关联讨论

同一事件的更多信源

相似阅读

另一事件,读法相近