Google论文揭示自主AI科研易产生严重结果幻觉
New Google paper shows that autonomous AI research can go badly wrong even when…
揭示了当前自主AI科研的核心风险:高幻觉率。提供具体数据证明日志校验的有效性,对构建可靠AI Agent极具参考价值。
New Google paper shows that autonomous AI research can go badly wrong even when the final paper looks convincing:
Google 最新论文表明,即使最终论文看起来令人信服,自主 AI 研究也可能出现严重偏差:
severe result hallucinations appeared in 90% of Agent Laboratory papers and 46% of Co-Scientist papers when its reliability modules were removed.
当移除其可靠性模块时,Agent Laboratory 的论文中有 90% 出现了严重的结果幻觉,Co-Scientist 的论文中这一比例为 46%。
With Co-Scientist checking manuscript claims against the actual execution logs, that rate dropped to just 4%, and complete data fabrication fell to 0%.
通过 Co-Scientist 将手稿中的声明与实际执行日志进行核对,该比例降至仅 4%,且完全的数据伪造率降至 0%。
– arxiv. org/abs/2608.26701
– arxiv.org/abs/2608.26701
Title: "Accelerating Scientific Research with Gemini in the Real-World"
标题:《在现实世界中利用 Gemini 加速科学研究》
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力