斯坦福与Together AI论文:Agent协作优于辩论投票
Stop making your AI agents debate and vote.
这篇论文揭示了Agent协作的新范式,用角色分工和历史复盘替代低效的辩论投票,效果显著提升,值得做Multi-Agent系统的同学参考其设计思路。
Stop making your AI agents debate and vote.
别再让你的 AI 智能体进行辩论和投票了。
New Stanford+Together AI paper shows teams that learn to check each other's work solve problems none of them got right alone.
斯坦福大学与 Together AI 联合发表的新论文显示,那些学会互相检查工作成果的团队,能够解决单个成员无法独立解决的难题。
Many agent setups just have models debate, then vote. That mostly picks an answer someone already had.
许多智能体设置仅让模型进行辩论然后投票,这通常只是从已有的答案中选择一个。
Here, 1 model reviewed past team chats and rewrote the team's playbook, like who double-checks whom and who plays devil's advocate. Learning this took just 15 practice problems.
在此方法中,一个模型审查过去的团队聊天记录并重写了团队的协作手册,例如规定谁需要复核谁的工作、谁扮演反对者角色。学习这一过程仅需 15 道练习题。
On math and physics tests, the 3-model team averaged 66.7%, versus 48.8% for its best model alone. It even beat perfectly choosing among the models' own answers, so teamwork created right answers none of them had.
在数学和物理测试中,由 3 个模型组成的团队平均准确率为 66.7%,而其表现最好的单个模型仅为 48.8%。该团队的表现甚至优于直接从各模型答案中选择最优解的策略,表明团队协作创造了原本无人能得出的正确答案。
So skip the debate-and-vote script: give agents clear jobs like checker and challenger, and let past runs improve them.
因此,请摒弃“辩论加投票”的模式:为智能体分配明确的职责(如检查者和挑战者),并利用过往运行结果来优化它们。
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力