27B研究智能体Faraday超越Claude Opus 4.8与GPT-5.5
A 27B research agent outscored Claude Opus 4.8 and GPT-5.5 on held-out paper rep…
A 27B research agent outscored Claude Opus 4.8 and GPT-5.5 on held-out paper replication by learning how to direct the research while outsourcing the coding.
一个27B的研究智能体通过学会如何指导研究同时外包编码,在保留论文复现任务上得分超过了Claude Opus 4.8和GPT-5.5。
Huge implication, maybe you do not need one giant model to do everything.
巨大的启示,也许你不需要一个巨型模型来做所有事情。
You can train one model to think like the researcher, then let stronger coding models do the implementation underneath it.
你可以训练一个模型像研究员一样思考,然后让更强的编码模型在它下面执行实现。
Replica, proposed in this paper, a scalable task space for paper replication, gives it a surprisingly simple way to practice that job.
本文提出的Replica,一个可扩展的论文复现任务空间,为练习这项工作提供了一种出奇简单的方法。
Take a research paper, remove one results figure, and tell the agent: recreate this result by actually running the experiment.
拿一篇研究论文,移除一个结果图,告诉智能体:通过实际运行实验来重现这个结果。
Now the agent has to figure out all the messy stuff papers leave out: what to implement, what to simplify, what experiments to run, and whether the result is believable.
现在智能体必须弄清楚论文中遗漏的所有杂乱细节:要实现什么,简化什么,运行哪些实验,以及结果是否可信。
That gives the researchers something they can repeatedly train on, with an automated rubric grading each attempt.
这给了研究人员可以反复训练的东西,并有一个自动评分标准来评估每次尝试。
After training on 242 of these tasks, Faraday uses GPT-5.5 as its coding agent but decides what research to do.
在242个这样的任务上训练后,Faraday使用GPT-5.5作为其编码智能体,但决定要做什么研究。
On 68 unseen AI-for-science tasks, Faraday scored 0.791 versus 0.748 for Claude Opus 4.8 and 0.729 for GPT-5.5, beating both on 60% of tasks.
在68个未见过的AI用于科学任务上,Faraday得分0.791,而Claude Opus 4.8为0.748,GPT-5.5为0.729,在60%的任务上击败了这两者。
And the difference was not just prettier plots.
而且差异不仅仅是更漂亮的图表。
Faraday was more likely to actually test the mechanism in the paper instead of taking shortcuts that produced the expected-looking answer.
Faraday更可能实际测试论文中的机制,而不是采取产生预期外观答案的捷径。
– arxiv. org/abs/2608.13331
– arxiv.org/abs/2608.13331
Title: "Training AI Scientists to Replicate Research"
标题:“训练AI科学家复现研究”
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力