跳到主内容
@wquguru
精选88elvis论文研究

Google发布Stellar Colosseum多智能体数学证明系统

Banger paper from Google on agent harnesses for long-horizon tasks.

原文
发到 X
推荐理由

Agent协作解决高难度数学证明的完整工作流与参数表现,值得收藏参考。

Banger paper from Google on agent harnesses for long-horizon tasks.

Google 关于智能体工具集在长周期任务中应用的精彩论文。

(bookmark it)

(收藏它)

Google Research built a many-agent harness for long mathematical proofs, and it produced new results on open problems from FOCS and JMLR papers.

Google Research 构建了一个用于长数学证明的多智能体工具集,它在 FOCS 和 JMLR 论文中的开放问题上产生了新的成果。

Stellar Colosseum works in stages.

Stellar Colosseum 分阶段工作。

It explores several proof strategies, waits for a readiness gate before breaking a route into section-level subproblems, and sends each verifier finding back to the section it affects.

它探索多种证明策略,在进入下一环节前等待就绪门控,将路径分解为节级子问题,并将每个验证器的发现反馈回受影响的章节。

Inside each stage, candidates are generated in parallel, attacked with targeted falsification, and merged together with their critiques.

在每个阶段内,候选方案并行生成,通过定向证伪进行攻击,并与各自的批评意见合并。

With Gemini 3.1 Pro and Gemini 3.7 Flash it reaches 71.0% on TCS-Bench, a set of research-level theorem-proving tasks from FOCS, STOC and SODA papers. With execution feedback it solves 218 of 222 Codeforces problems.

使用 Gemini 3.1 Pro 和 Gemini 3.7 Flash,它在 TCS-Bench(一组来自 FOCS、STOC 和 SODA 论文的研究级定理证明任务)上达到 71.0% 的准确率;结合执行反馈,它解决了 222 道 Codeforces 题目中的 218 道。

Paper: https://arxiv.org/abs/2609.15983

论文:https://arxiv.org/abs/2609.15983

Chat with Paper: https://academy.dair.ai/papers/stellar-colosseum-a-many-agent-harness-for-long-horizon-research-in-mathematics-2609.15983

与论文对话:https://academy.dair.ai/papers/stellar-colosseum-a-many-agent-harness-for-long-horizon-research-in-mathematics-2609.15983

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

相似阅读

关联信息,但可能不是同一事件