Google ScientistOne:用证据链解决AI研究可信度问题
Google’s ScientistOne paper tackles a basic problem with AI-generated research:
Google’s ScientistOne paper tackles a basic problem with AI-generated research:
谷歌的ScientistOne论文解决了一个AI生成研究的基本问题:
The result can look credible even when its evidence chain is broken.
即使证据链断裂,结果也可能看起来可信。
AI research agents are getting good enough at solving benchmark problems that the new bottleneck is whether you can trust the paper they write afterward.
AI研究代理在解决基准问题方面已经足够出色,以至于新的瓶颈在于你是否能信任它们随后撰写的论文。
Google Cloud AI Research audited 75 papers from five autonomous research systems on five ADRS tasks, and every baseline showed at least one systematic evidence failure.
谷歌云AI研究审计了五个自主研究系统在五个ADRS任务上的75篇论文,每个基线都显示出至少一个系统性的证据失败。
Some fabricated citations, some reported scores that did not reproduce, and some described algorithms that were simply not in the submitted code.
有些捏造了引用,有些报告了无法复现的分数,还有些描述的算法根本不在提交的代码中。
ScientistOne attacks that gap with “Chain-of-Evidence”: citations must trace to retrieved papers, numerical claims to evaluator logs, and method claims to implementation artifacts before the manuscript is finalized.
ScientistOne通过“证据链”来弥补这一差距:在稿件定稿之前,引用必须追溯到检索到的论文,数值声明必须追溯到评估器日志,方法声明必须追溯到实现工件。
– arxiv. org/abs/2605.26340
– arxiv.org/abs/2605.26340
Title: "ScientistOne: Towards Human-Level Autonomous Research via Chain-of-Evidence"
标题:“ScientistOne:通过证据链迈向人类水平的自主研究”
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力