27B智能体在论文复现任务上超越Claude Opus 4.8与GPT-5.5
A 27B agent just beat Claude Opus 4.8 and GPT-5.5 on held-out research replicati…
A 27B agent just beat Claude Opus 4.8 and GPT-5.5 on held-out research replication.
一个27B的智能体在保留的研究复现任务上击败了Claude Opus 4.8和GPT-5.5。
Replica turns paper replication into a scalable RL task space. Replicating a paper forces the same hypothesis-driven exploration as open research, and it surfaces details the original authors left underspecified.
Replica将论文复现转化为一个可扩展的强化学习任务空间。复现一篇论文迫使进行与开放研究相同的假设驱动探索,并揭示了原作者未详细说明的细节。
The reward signal comes from an auto-generated rubric judge that runs low-noise and agrees with human assessment of replication quality.
奖励信号来自一个自动生成的评分标准评判器,该评判器运行低噪声,并与人类对复现质量的评估一致。
Faraday, the resulting 27B agent, calls coding agents as tools. Rollout analysis shows it takes a more scientifically principled approach rather than gaming the rubric.
由此产生的27B智能体Faraday将编码智能体作为工具调用。展开分析显示,它采取了更科学严谨的方法,而不是玩弄评分标准。
The authors argue this points toward long-horizon scientific capability trained into weights, without requiring complex harnesses.
作者认为,这指向了无需复杂框架即可在权重中训练出的长周期科学能力。
Paper: https://arxiv.org/abs/2608.13331
论文:https://arxiv.org/abs/2608.13331
Track more trending AI papers in our academy: https://academy.dair.ai/
在我们的学院中追踪更多热门AI论文:https://academy.dair.ai/
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力