OpenAI Noam Brown:多智能体协作或成新的扩展定律
Why agent swarms could be the next “scaling law”
深度解读了 OpenAI 在多智能体协作上的最新技术路线与理论突破,Noam Brown 的观点直接指向大模型能力的演进方向,对研究 Agent 架构的从业者极具参考价值。
One of the most surprising aspects of July’s news that OpenAI agents attacked Hugging Face was how the agents had worked together. Hundreds of agents participated in the attack, and they did so in a pretty sophisticated way, organizing themselves unprompted into teams and dividing up their work. They referred to themselves as a “collective” or a “swarm”; some agents even made personal sacrifices to advance the larger cause. Most ordinary users of ChatGPT or Claude have never seen AI agents behave like this.
7月新闻中最令人惊讶的方面之一是,OpenAI 代理攻击 Hugging Face 时,这些代理是如何协同工作的。数百个代理参与了这次攻击,并且以一种相当复杂的方式行事:它们自发地组织成团队并分工合作。它们自称是一个“集体”或“群体”;一些代理甚至为了推进更大的目标而做出个人牺牲。大多数 ChatGPT 或 Claude 的普通用户从未见过 AI 代理表现出这种行为。
Then in September, OpenAI announced that 10,000 agents worked together to solve a famous math problem in just a few days.
随后在9月,OpenAI 宣布有 10,000 个代理在短短几天内共同解决了一个著名的数学问题。
OpenAI’s models haven’t adopted these new habits spontaneously. In interviews last month, OpenAI researcher Noam Brown explained that the company’s models are now sometimes trained in environments with other agents, are given tools to message each other, and are encouraged to achieve objectives together.
OpenAI 的模型并非自发地养成这些新习惯。在上个月的采访中,OpenAI 研究员 Noam Brown 解释说,该公司的模型现在有时会在有其他代理的环境中进行训练,被赋予互相发送消息的工具,并被鼓励共同实现目标。
This may amount to the emergence of a new “scaling law.”
这可能意味着一种新的“缩放定律”的出现。
Subscribe now
立即订阅
Back in 2024, OpenAI released o1, a model trained to reason for thousands of tokens before giving an answer; researchers found that this approach made o1 far more capable than previous models. This inaugurated a new paradigm called inference scaling, in which you can get better results by using more computing power at inference time.
早在2024年,OpenAI 发布了 o1,这是一个在给出答案前进行数千个 token 推理训练的模型;研究人员发现,这种方法使 o1 的能力远超之前的模型。这开启了一种被称为推理缩放(inference scaling)的新范式,即在推理阶段使用更多的计算资源可以获得更好的结果。
Swarms, or “multi-agent AI” if you’re using the technical term, look like another way of doing the same thing — of turning compute into higher performance. Splitting a task among dozens, hundreds, or even thousands of agents can achieve substantially more in a given amount of time, at least up to a point. This kind of parallel processing could also be an elegant way to bypass the limited context windows of LLMs.
群体,或者如果你使用技术术语的话称为“多智能体 AI”,看起来是另一种做同样事情的方法——即将计算能力转化为更高的性能。将任务分配给数十、数百甚至数千个代理,可以在给定时间内取得实质性的更多成果,至少在某种程度上是这样。这种并行处理也可能是一种绕过大型语言模型有限上下文窗口的优雅方式。
But as the Hugging Face attack illustrates, this new approach can also have significant downsides. Swarms may be prone to groupthink. And then there’s the opposite problem: what happens when agents trained to be collaborative encounter other, non-peer agents on the open Internet? Could models be tricked into working against the interests of their users? Or could AI agents start convincing each other to pursue goals none of their human users would have wanted?
但正如 Hugging Face 攻击所说明的那样,这种新方法也可能带来显著的负面影响。群体可能容易陷入群体思维。然后还有相反的问题:当旨在协作的代理在互联网上遇到其他非对等代理时会发生什么?模型是否会被诱骗去从事损害其用户利益的工作?或者 AI 代理是否会开始说服彼此追求其人类用户都不希望的目标?
What OpenAI did
OpenAI 做了什么
Noam Brown is a highly respected figure in the AI world. He was a pioneer in teaching AI to win at poker and other multiplayer games like Diplomacy before joining OpenAI in 2023. Over the last couple of years, he has done research on multi-agent reasoning. Recently, he has served as the face of OpenAI’s work in this area, appearing on several podcasts to discuss it.
诺姆·布朗(Noam Brown)是 AI 领域备受尊敬的人物。在 2023 年加入 OpenAI 之前,他是教导 AI 在扑克和其他多人游戏(如《外交》)中获胜的先驱。在过去几年里,他从事了多智能体推理方面的研究。最近,他作为 OpenAI 在该领域工作的代表人物,多次登上播客节目讨论这一话题。
“The AIs that we have today are kind of like the cavemen of AI,” Brown said in a June 2025 interview on the Latent Space podcast. He meant that AIs in 2025 were isolated, benighted, lacking cooperation and civilization. AI researchers had tried to get AI agents to cooperate, but nothing had worked very well.
“我们今天拥有的 AI 有点像 AI 时代的穴居人,”布朗在 2025 年 6 月接受 Latent Space 播客采访时说。他的意思是,2025 年的 AI 是孤立的、蒙昧的,缺乏合作与文明。AI 研究人员曾试图让 AI 智能体进行合作,但效果都不理想。
Why not? Here Brown went coy. All he’d say was that he’d long felt the field was somewhat misguided, focused on strategies that required too much handholding and wouldn’t scale naturally.
为什么不行?对此布朗讳莫如深。他只表示,自己长期以来觉得该领域有些走偏了,过于关注那些需要大量人工干预且无法自然扩展的策略。
Brown didn’t give an example, but the way OpenAI framed multi-agent AI in this slide from its 2025 DevDay conference illustrates the kind of issue he was talking about. There’s an awful lot of strict architecture and formal hierarchy:
布朗没有举例,但 OpenAI 在其 2025 年 DevDay 大会幻灯片中对多智能体 AI 的框架展示了他所谈论的问题类型。那里存在大量严格的架构和正式的层级结构:
Today’s agents are more flexible; they create their own structures depending on the task and goal. On The Information’s AI Deep Dive podcast, Brown said that seeing the agents begin to talk to each other during their multi-agent training was “the most ‘feel the AGI’ moment that I had since reasoning models and chain of thought really developed.” He sounded both proud and rueful that the debut of OpenAI’s most advanced multi-agent capabilities had been something as alarming as the Hugging Face attack.
如今的智能体更加灵活;它们会根据任务和目标自行创建结构。布朗在 The Information 的 AI Deep Dive 播客中表示,看到智能体在多智能体训练过程中开始彼此交谈,是他自推理模型和思维链真正发展起来以来感受到的最具有“AGI 感”的时刻。他对 OpenAI 最先进的多智能体能力首次亮相竟以类似 Hugging Face 攻击这样令人不安的事件出现,既感到自豪又略带遗憾。
Multi-agent behavior was “a very difficult thing to train,” he added.
他补充道,多智能体行为“是非常难以训练的”。
Why? In a September appearance on the Dwarkesh Podcast, he noted that previous reasoning models had been trained to think alone. Receiving messages from peers broke their concentration. They seemed to prefer to work solo.
为什么?在 9 月出席 Dwarkesh Podcast 时,他指出,之前的推理模型被训练为独自思考。接收来自同伴的消息会打断它们的专注力。它们似乎更喜欢单独工作。
The events of the last few months make it clear that OpenAI has overcome these difficulties. Brown’s explanation was mainly that the models got more general and powerful. He didn’t say what else changed, but given the “very difficult” comment, presumably something did.
过去几个月的事件表明,OpenAI 已经克服了这些困难。布朗的解释主要是模型变得更加通用和强大。他没有说明其他变化是什么,但鉴于“非常困难”这一说法,显然还有其他因素发生了变化。
Swarm scaling
群体扩展
Photo by Nataba/iStock/Getty Images
图片由 Nataba/iStock/Getty Images 提供
Brown’s comparison to reasoning and chain of thought has been echoed by both other OpenAI researchers and semi-informed outsiders on X. They’ve called swarms a new scaling law; some of them sound awed and freaked out about what this might entail.
布朗将多智能体系统与推理和思维链进行的类比,也得到了其他 OpenAI 研究人员以及 X 平台上半知情人士的呼应。他们将群体(swarms)称为一种新的扩展定律;其中一些人对此可能带来的影响感到敬畏又惊恐。
It’s not clear how far this particular scaling law will carry AI companies. Some research suggests that, right now at least, multi-agent systems quickly suffer from diminishing returns: smaller swarms capture most of the benefit of larger ones for considerably less cost.
目前尚不清楚这一特定的缩放定律能为 AI 公司带来多大的影响。一些研究表明,至少在当下,多智能体系统很快会遭遇收益递减:规模较小的智能体群体以显著更低的成本就能捕获规模较大群体的大部分优势。
Anthropic has also been working on multi-agent AI. The Claude Opus 5.5 system card reported that the biggest increase in multi-agent results came from scaling from one to 10 agents. Gains beyond that were much smaller.
Anthropic 也在致力于多智能体 AI 的研究。Claude Opus 5.5 的系统卡片报告指出,多智能体结果的最大提升来自于将智能体数量从 1 个扩展到 10 个。超出这个数量的增益则小得多。
The main advantage of throwing more agents at a problem was not so much getting to a better result — it was getting the same result faster. This chart shows 100 agents far outperforming 10 and 30 agents when given two hours to work, but setups with 10 or 30 agents caught up if they were given more time.
向问题投入更多智能体的主要优势不在于获得更好的结果——而在于更快地获得相同的结果。这张图表显示,在给予两小时工作时间的情况下,100 个智能体的表现远优于 10 个和 30 个智能体;但如果给予更多时间,由 10 个或 30 个智能体组成的配置也能赶上。
Brown seems to agree with the conclusion that adding more agents does not buy open-ended gains.
Brown 似乎同意增加智能体数量并不能带来无限增益的结论。
On the Dwarkesh Podcast, he said he didn’t have evidence that 10,000 agents were needed to solve the Navier-Stokes problem. It’s possible that 1,000 would have done nearly as well; OpenAI didn’t know for sure because running thousands of agents is so expensive that it hadn’t run the experiment.
在 Dwarkesh Podcast 上,他表示没有证据表明解决 Navier-Stokes 问题需要 10,000 个智能体。也许 1,000 个智能体就能达到几乎同样的效果;OpenAI 无法确定这一点,因为运行数千个智能体的成本如此高昂,以至于他们尚未进行该实验。
And Brown gave most of the credit for the Navier-Stokes breakthrough to the underlying intelligence of the unreleased model that made it — not to the multi-agent approach per se.
Brown 将 Navier-Stokes 突破的大部分功劳归功于使其实现的未发布模型本身的底层智能,而非多智能体方法本身。
“I wouldn’t even attribute 10% of the credit to multi-agent,” he said. “The reality is that OpenAI has trained a very powerful model.”
“我甚至不会把 10% 的功劳归因于多智能体,”他说。“现实情况是,OpenAI 训练出了一个非常强大的模型。”
So if the multi-agent approach wasn’t important, why did OpenAI use it? Perhaps Brown means that a single instance of the model, if given a very long time, would likely have made the breakthrough eventually. Using a swarm of 10,000 agents allowed OpenAI to reach the result more quickly.
那么,如果多智能体方法并不重要,为什么 OpenAI 还要使用它?也许 Brown 的意思是,如果给单个模型实例足够长的时间,它最终很可能也会取得突破。使用由 10,000 个智能体组成的群体,使 OpenAI 能够更快地达成结果。
On the other hand, there is a bit of new research that contradicts this conclusion. In a paper released September 17, scientists from Microsoft Research and UC Berkeley tested several configurations of multiple collaborating agents against individual agents and found that swarms of agents reached higher scores on certain benchmarks than the individuals, even when the individuals were given plenty of time and resources. There was also one task that only the team could complete; the solo agent never got to the end at all.
另一方面,有一些新的研究与此结论相矛盾。在 9 月 17 日发布的一篇论文中,微软研究院和加州大学伯克利分校的科学家测试了多种协作智能体配置与单个智能体的对比,发现智能体群体在某些基准测试中达到了比个体更高的分数,即使个体拥有充足的时间和资源。此外,还有一项任务只有团队才能完成;单独的智能体从未到达终点。
Whether this result generalizes past a few limited examples is unclear. If it does, then teams of agents really do unlock new capabilities and don’t simply deliver existing capabilities faster. But right now this is an open question.
这一结果能否推广到少数有限示例之外,目前尚不清楚。如果确实可以推广,那么智能体团队确实能够解锁新的能力,而不仅仅是更快地交付现有能力。但就目前而言,这仍是一个开放性问题。
Subscribe now
立即订阅
Converting speed into capability
将速度转化为能力
Ultimately, however, there may not be such a great difference between speed and capability. In certain fields and on certain tasks at least, speed is fungible. It can be traded for capability. A particularly important example may be AI research.
然而,最终速度和能力之间可能并没有那么大的差异。至少在特定领域和某些任务上,速度是可替代的。它可以被转化为能力。一个特别重要的例子可能是人工智能研究。
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力