跳到主内容
@wquguru
精选88Dwarkesh Podcast(RSS)模型发布/更新

OpenAI Noam Brown:万智能体协作解千禧年难题与推理模型扩展

Noam Brown – Agent swarms, alignment, & recursive self-improvement

原文
发到 X
推荐理由

OpenAI首次公开万级智能体协作解决重大数学难题的工程细节,揭示了多智能体并行扩展的真实效能边界,对理解下一代推理模型架构极具参考价值。

New episode with Noam Brown.

新一期节目,嘉宾是 Noam Brown。

We talk about multi-agent, Navier-Stokes, and what the current explosion of maths progress tells us about what happens once you automate AI research.

我们讨论了多智能体、纳维-斯托克斯方程,以及当前数学领域的爆发式进展告诉我们,一旦实现 AI 研究自动化,将会发生什么。

And we also discuss how we will know if the models are actually aligned before we kick off RSI.

我们还讨论了在启动 RSI(递归自我改进)之前,我们将如何知道模型是否真正实现了对齐。

Watch on YouTube; listen on Apple Podcasts or Spotify.

在 YouTube 上观看;在 Apple Podcasts 或 Spotify 上收听。

Sponsors

赞助商

  • Jane Street has been interested in AI for a lot longer than you’d think, and not just for trading. In 2011, a full year before AlexNet and over a decade before ChatGPT launched, they hosted the first FOOM Debate between Eliezer Yudkowsky and Robin Hanson on whether AI would lead to an intelligence explosion. Now Jane Street is revisiting the question with a new panel: Daniel Kokotajlo, Ege Erdil, Ryan Greenblatt, and Jaime Sevilla, hosted by Ron Minsky in San Francisco this October. I expect it to be a truly excellent conversation. Register at janestreet.com/dwarkesh
  • Grok Bot has made handing off work super easy. It runs on its own cloud computer, where it installs the tools it needs to handle tasks end-to-end. For the podcast, we use Grok Bot to help produce our videos. You may have noticed that our ads feature animations of real websites. Getting these pixel-perfect used to mean running a convoluted, multi-step workflow ourselves. Now we just let Grok Bot handle it. Best of all, Grok Bot has learned all of our specs and preferences, so we don’t have to redescribe the task each time! Try Grok Bot for yourself at x.ai/bot
  • Antithesis gives you the confidence of a giant test suite without actually having to write one. Say you’re doing a major backend refactor: building enough tests to trust it could take weeks. Antithesis solves this by running your software through countless simulated worlds, injecting faults and hunting for failures. On any PR, you can turn a dial to decide exactly how much testing you want. And because every run is fully deterministic, agents can branch off the moment a bug appears, rewind it, inspect memory, and replay it, all while the original test keeps running. Learn more at antithesis.com/dwarkesh
  • Jane Street 对 AI 的兴趣远比人们想象的要早,而且不仅仅局限于交易领域。2011 年,即在 AlexNet 发布前整整一年,以及在 ChatGPT 推出前十多年,他们举办了首届 FOOM 辩论,由 Eliezer Yudkowsky 和 Robin Hanson 就“AI 是否会导致智力爆炸”展开辩论。如今,Jane Street 正在通过一个新的专家组重新探讨这一问题:Daniel Kokotajlo、Ege Erdil、Ryan Greenblatt 和 Jaime Sevilla,由 Ron Minsky 主持,将于今年十月在旧金山举行。我期待这将是一场真正精彩的对话。请在 janestreet.com/dwarkesh 注册。
  • Grok Bot 让工作交接变得超级简单。它在自己的云计算机上运行,安装处理端到端任务所需的工具。对于播客,我们使用 Grok Bot 来帮助制作视频。你可能已经注意到,我们的广告中包含了真实网站的动画。以前要获得这些像素级完美的画面,意味着我们要自己运行一个复杂的多步骤工作流程。现在,我们只需让 Grok Bot 来处理。最重要的是,Grok Bot 已经学习了我们的所有规格和偏好,因此我们不必每次都重新描述任务!请前往 x.ai/bot 亲自体验 Grok Bot。
  • Antithesis 让你拥有大型测试套件的信心,而无需实际编写任何测试。假设你正在进行重大的后端重构:构建足够的测试以建立信任可能需要数周时间。Antithesis 通过在无数模拟世界中运行你的软件、注入故障并寻找失败来解决了这个问题。在任何 PR(拉取请求)上,你都可以调节一个旋钮来决定你想要多少测试量。而且由于每次运行都是完全确定性的,一旦出现 bug,代理可以立即分支,回溯它,检查内存并重放它,同时原始测试仍在运行。更多信息请访问 antithesis.com/dwarkesh。

Timestamps

时间戳

(00:00:00) – Multi-agent and Navier-Stokes

(00:00:00) – 多智能体和纳维-斯托克斯方程

(00:15:28) – How will AI firms work?

(00:15:28) – AI 公司将如何运作?

(00:22:02) – What math progress tells us about recursive self improvement

(00:22:02) – 数学进展告诉我们关于递归自我改进的什么信息

(00:40:22) – Hugging Face and alignment

(00:40:22) – Hugging Face 和对齐

(01:01:18) – The internal/external model gap

(01:01:18) – 内部/外部模型差距

(01:08:34) – Chain of thought is degrading

(01:08:34) – 思维链正在退化

(01:14:12) – How will we know when alignment is solved?

(01:14:12) – 我们将如何知道对齐问题何时得到解决?

Transcript

文字稿

00:00:00 – Multi-agent and Navier-Stokes

00:00:00 – 多智能体和纳维-斯托克斯方程

Dwarkesh Patel

Dwarkesh Patel

Today, I’m chatting with Noam Brown, who is a researcher at OpenAI. He was one of the foundational contributors to what became o1 and the reasoning models. Now he’s working on multi-agent systems. Speaking of which, you guys announced last week that you solved one of the Millennium Prize Problems with a system of 10,000 different AI agents that spent 130 billion tokens over 88 hours.

今天,我与 OpenAI 的研究员 Noam Brown 进行了一次对话。他是 o1 及推理模型的核心贡献者之一。如今,他正致力于多智能体系统(multi-agent systems)的研究。说到这个,你们上周宣布,通过一个由 10,000 个不同 AI 智能体组成的系统,在 88 小时内消耗了 1300 亿个 token,解决了一个千禧年大奖难题。

One of the reasons I’m interested in talking to you is that you were among the first people, maybe two or three years ago, who were thinking about how the reasoning models would allow us to see into the future. Because if you scale up inference compute, you can see what the base capabilities of the models will be a few years in the future.

我之所以有兴趣与你交谈,是因为早在两三年前,你就是最早一批思考推理模型如何让我们预见未来的人之一。因为如果增加推理计算量,你就能看出模型在未来几年将具备哪些基础能力。

I feel like you’re in a similar position now to help us understand what future capabilities will look like, given the enormous scaling of agent sizes that we can do right now.

我觉得你现在处于类似的位置,能够帮助我们理解未来的能力将呈现何种面貌,鉴于我们现在能够对智能体规模进行如此巨大的扩展。

Noam Brown

Noam Brown

The way I think about it, when you plot the performance of these reasoning models with test-time compute on the x-axis and performance on basically any reasoning benchmark on the y-axis, you see a very clear pattern where the longer these models take to think about their answer, the better they do. This is a very natural thing. It’s the same thing with people. If you’re taking the SATs and you have five minutes to go through the entire exam, you’re not going to do very well. If you have five hours, you’re probably going to do a lot better.

我是这样看待这个问题的:当你在 x 轴上绘制这些推理模型在测试时计算(test-time compute)下的性能,在 y 轴上绘制其在几乎所有推理基准测试中的性能时,你会看到一个非常清晰的模式:模型思考答案的时间越长,表现就越好。这非常自然。人类也是如此。如果你参加 SAT 考试,只有五分钟时间完成整张试卷,你的成绩不会很好。如果你有五个小时,你的成绩可能会好得多。

The AI models are pretty similar. They’ll spend that time doing this monologue to themselves, figuring things out, going through different cases, ruling out different possibilities, building on some of their previous discoveries.

AI 模型的情况也大同小异。它们会花费这段时间进行自我独白,理清思路,探讨不同的情况,排除各种可能性,并基于之前的发现进一步构建。

The problem is that as you push that further and further, you hit a latency bottleneck. You don’t want to sit around for three years waiting for a response. So what you can do is what a lot of people do. They parallelize. They just get a team of people. If you’re going to found a company, you want to get a group of people together so you can go faster. It’s the same thing with these AI models. It helps to just have multiple agents working on something because they can go faster.

问题在于,当你不断推高这一过程时,会遇到延迟瓶颈。你不想干等三年才得到一个回复。因此,你可以采取许多人的做法:并行化。组建一个团队。如果你要创办一家公司,你需要聚集一群人以便更快地推进工作。对于 AI 模型也是如此,让多个智能体协同工作有助于加快进度。

So multi-agent is a way of scaling test-time compute in parallel instead of purely serially. It is less efficient, because it’s not like a single agent has all the context to itself. But it is a very effective way of scaling test-time compute if it’s done well.

因此,多智能体是一种以并行而非纯串行方式扩展测试时计算的方法。它的效率较低,因为不像单个智能体那样拥有完整的上下文信息。但如果执行得当,这是一种非常有效的扩展测试时计算的方式。

Dwarkesh Patel

Dwarkesh Patel

I’m going to ask a bunch of naive questions. This is an unreleased model, so we haven’t publicly seen how these systems work. I just have a bunch of ways in which I’m confused about what the qualitative properties of such systems are.

我要问一堆天真的问题。这是一个尚未发布的模型,所以我们还没有公开看到这些系统是如何工作的。我只是对这类系统的定性特性有很多困惑之处。

I am shocked by the scale of cognitive effort that you can concentrate in such a short period of time. Think about what 130 billion tokens are. If it were a single human thinking as a full-time job, stretched back to back, 130 billion tokens would be a human thinking for 4,000 years. Eight hours a day, working a normal work week. Starting from ancient Sumeria up till today, a single sequential human thinking that long, concentrated in 88 hours.

我对在如此短的时间内能够集中的认知努力规模感到震惊。想想1300亿个token意味着什么。如果是一个人类将其作为全职工作,连续不断地思考,1300亿个token相当于一个人类思考了4000年。每天工作8小时,按正常的工作周计算。从古代的苏美尔一直到现在,一个顺序思考的人类花了这么长的时间,却集中在88小时内完成。

I feel like qualitatively, that is a super important consideration. I’m surprised that there isn’t a bigger parallelization penalty. You can just have 10,000 agents collaborate. Maybe because the agents are better at collaborating than humans might be, they’re going much faster. They can actually productively collaborate at such a big scale. Or maybe there is a big parallelization penalty.

我觉得从定性角度来看,这是一个非常重要的考量因素。我惊讶于并行化惩罚并没有那么大。你可以让1万个代理协同工作。也许是因为代理比人类更擅长协作,所以它们的速度要快得多。它们实际上可以在如此大的规模上进行有效的协作。或者也许存在很大的并行化惩罚。

Noam Brown

Noam Brown

Let’s talk about the parallelization penalty, and then we can talk about the qualitative stuff. The truth is that we don’t have very good science on multi-agent scaling up to this kind of scale. When we released 5.6, I think that was the first time that we had a proper multi-agent system in our models. We actually did show some plots in the blog post of the scaling performance of multi-agent systems, because we have it as an option. It’s Ultra Mode. The default is four agents, but you can set that higher.

让我们谈谈并行化惩罚,然后我们可以讨论定性方面的问题。事实上,我们在多智能体扩展到这种规模方面的科学研究还不够完善。当我们发布5.6版本时,我认为那是我们首次在模型中拥有真正的多智能体系统。我们确实在博客文章中展示了一些关于多智能体系统扩展性能的图表,因为这是一个可选功能。它是Ultra模式。默认是四个代理,但你可以将其设置得更高。

In the plot, we show what the performance looks like on some benchmarks for one agent, for four agents working together, for 16 agents working together. It depends on the benchmark, but for some of the benchmarks, what you see is that if you have four agents working on the problem, it is done twice as fast. Because there are four agents working for half as long, you’re paying 2x more to get an answer twice as quickly. If you go to 16 agents, you see a similar pattern. It’s a little less efficient, but you continue to see that performance.

在图表中,我们展示了在一个基准测试上,单个代理、四个代理协同工作以及16个代理协同工作的性能表现。这取决于具体的基准测试,但对于某些基准测试,你会发现如果有四个代理共同解决问题,速度会快两倍。因为有四个代理以一半的时间工作,你支付了2倍的费用以获得两倍的回答速度。如果你增加到16个代理,你会看到类似的模式。效率稍低一些,但你仍然能看到这种性能提升。

Dwarkesh Patel

Dwarkesh Patel

Is it a linear serial time speedup or a sublinear speedup as you increase the number of parallel agents?

随着并行代理数量的增加,这是线性串行时间加速还是次线性加速?

Noam Brown

Noam Brown

It’s slightly sublinear, though it does depend a lot on the problem. Math, for example, is quite parallelizable. It’s not the most parallelizable thing, but it is very parallelizable. Web search, things like doing a Deep Research report where you have to look through a bunch of sources, is extremely parallelizable. I suspect that something like writing a novel would be very unparallelizable. You would probably not see a big benefit from having 10,000 agents working on a novel together, in the same way that you’d probably not get a big benefit from having 10,000 people work on a novel together.

虽然它的增长速度略低于线性,但这在很大程度上取决于具体问题。例如,数学问题具有相当高的并行性。它并非最具并行性的任务,但确实非常适合并行处理。像网络搜索、撰写深度研究报告(需要查阅大量资料)这类任务则具有极高的并行性。我怀疑写小说这类任务几乎无法并行化。让1万个智能体共同创作一部小说,可能并不会带来显著收益,正如让1万人共同写一部小说也不太可能获得巨大好处一样。

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

关联讨论

同一事件的更多信源

相似阅读

关联信息,但可能不是同一事件