跳到主内容
精选88Rohan Paul论文研究

Google论文揭示自主AI科研易产生严重结果幻觉

New Google paper shows that autonomous AI research can go badly wrong even when…

原文
推荐理由

揭示了当前自主AI科研的核心风险:高幻觉率。提供具体数据证明日志校验的有效性,对构建可靠AI Agent极具参考价值。

New Google paper shows that autonomous AI research can go badly wrong even when the final paper looks convincing:

Google 最新论文表明,即使最终论文看起来令人信服,自主 AI 研究也可能出现严重偏差:

severe result hallucinations appeared in 90% of Agent Laboratory papers and 46% of Co-Scientist papers when its reliability modules were removed.

当移除其可靠性模块时,Agent Laboratory 的论文中有 90% 出现了严重的结果幻觉,Co-Scientist 的论文中这一比例为 46%。

With Co-Scientist checking manuscript claims against the actual execution logs, that rate dropped to just 4%, and complete data fabrication fell to 0%.

通过 Co-Scientist 将手稿中的声明与实际执行日志进行核对,该比例降至仅 4%,且完全的数据伪造率降至 0%。

– arxiv. org/abs/2608.26701

– arxiv.org/abs/2608.26701

Title: "Accelerating Scientific Research with Gemini in the Real-World"

标题:《在现实世界中利用 Gemini 加速科学研究》

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

相似阅读

另一事件,读法相近