Google发布Stellar Colosseum多智能体数学证明系统
Banger paper from Google on agent harnesses for long-horizon tasks.
Agent协作解决高难度数学证明的完整工作流与参数表现,值得收藏参考。
Banger paper from Google on agent harnesses for long-horizon tasks.
Google 关于智能体工具集在长周期任务中应用的精彩论文。
(bookmark it)
(收藏它)
Google Research built a many-agent harness for long mathematical proofs, and it produced new results on open problems from FOCS and JMLR papers.
Google Research 构建了一个用于长数学证明的多智能体工具集,它在 FOCS 和 JMLR 论文中的开放问题上产生了新的成果。
Stellar Colosseum works in stages.
Stellar Colosseum 分阶段工作。
It explores several proof strategies, waits for a readiness gate before breaking a route into section-level subproblems, and sends each verifier finding back to the section it affects.
它探索多种证明策略,在进入下一环节前等待就绪门控,将路径分解为节级子问题,并将每个验证器的发现反馈回受影响的章节。
Inside each stage, candidates are generated in parallel, attacked with targeted falsification, and merged together with their critiques.
在每个阶段内,候选方案并行生成,通过定向证伪进行攻击,并与各自的批评意见合并。
With Gemini 3.1 Pro and Gemini 3.7 Flash it reaches 71.0% on TCS-Bench, a set of research-level theorem-proving tasks from FOCS, STOC and SODA papers. With execution feedback it solves 218 of 222 Codeforces problems.
使用 Gemini 3.1 Pro 和 Gemini 3.7 Flash,它在 TCS-Bench(一组来自 FOCS、STOC 和 SODA 论文的研究级定理证明任务)上达到 71.0% 的准确率;结合执行反馈,它解决了 222 道 Codeforces 题目中的 218 道。
Paper: https://arxiv.org/abs/2609.15983
论文:https://arxiv.org/abs/2609.15983
Chat with Paper: https://academy.dair.ai/papers/stellar-colosseum-a-many-agent-harness-for-long-horizon-research-in-mathematics-2609.15983
与论文对话:https://academy.dair.ai/papers/stellar-colosseum-a-many-agent-harness-for-long-horizon-research-in-mathematics-2609.15983
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力