DeepMind前首席科学家等探讨递归自我改进的技术瓶颈
AI researchers debate how close we are to recursive self-improvement
汇聚 OpenAI 创始人与一线实验室负责人的深度观点,直击 RSI 可行性这一核心争议,适合关注 AGI 路径的研究者收藏参考。
New episode with John Schulman, Beren Millidge and Charlie O’Neill.
新一期节目,嘉宾包括 John Schulman、Beren Millidge 和 Charlie O’Neill。
I got together with some of the most insightful AI researchers I know who are at the openish companies, because I wanted to hear the details of what's actually happening at the frontier and what comes next.
我与一些我所知的、在相对开放的公司工作的极具洞察力的 AI 研究人员聚在一起,因为我想知道前沿领域实际正在发生什么以及接下来会发生什么。
Watch on YouTube; listen on Apple Podcasts or Spotify.
在 YouTube 上观看;在 Apple Podcasts 或 Spotify 上收听。
Sponsors
赞助商
- Antithesis helps you trust your code. As agents generate more and more of your software, the bottleneck shifts from your engineers actually writing code to verifying it. Antithesis does that testing for you. Ron Minsky, who co-leads Jane Street’s tech group, told me that Antithesis was able to help his team shake out bugs in software that had already undergone heavy review. If you want to see how it fits into your development process, go to antithesis.com/dwarkesh
- Grok Bot has been a great way to hand off tasks. My team uses it as a producer: whenever my editor posts a rough cut of an interview in Slack, Grok Bot opens the transcript on its own computer, matches my notes to the exact moments they refer to, and uses a file of my preferences to suggest edits. Then it sends me its top clip candidates so I can review everything from my phone, which saves my editors from sorting through hours of footage. Try Grok Bot for yourself at x.ai/bot
- Jane Street just launched its most ambitious competition yet: design a protocol-emulator ASIC. Basically, if you have a chip you want to test outside of a live system, you should be able to connect it to your design and have it simulate realistic traffic. Jane Street wants general-purpose, reprogrammable designs that can work across multiple protocols and remain useful as new ones emerge. The most novel submissions will actually get taped out, and the winners will receive a physical copy! The competition is open until January 18, 2027, and teams are encouraged. To get started download the template code at janestreet.com/dwarkesh
- Antithesis 帮助你信任你的代码。随着智能体生成越来越多的软件,瓶颈已从工程师实际编写代码转移到验证代码。Antithesis 为你进行这项测试工作。Jane Street 技术团队的联合负责人 Ron Minsky 告诉我,Antithesis 帮助他的团队找出了已经经过严格审查的软件中的漏洞。如果你想了解它如何融入你的开发流程,请访问 antithesis.com/dwarkesh
- Grok Bot 是移交任务的好方法。我的团队将其作为制作人使用:每当我的编辑在 Slack 上发布采访的粗剪版本时,Grok Bot 就会在其计算机上自动打开转录文本,将我的笔记与它们所指的具体时刻匹配,并根据我的偏好文件建议修改。然后它会发送给我最推荐的几个片段候选项,让我可以在手机上审阅所有内容,这节省了我的编辑从数小时素材中进行筛选的时间。亲自试用 Grok Bot,访问 x.ai/bot
- Jane Street 刚刚推出了他们最具雄心的比赛:设计一个协议仿真器 ASIC。基本上,如果你有一个想在实时系统之外测试的芯片,你应该能够将其连接到你的设计中,并让它模拟真实的流量。Jane Street 希望得到通用的、可重新编程的设计,这些设计可以跨多种协议工作,并且在新协议出现时仍能保持有用性。最有创意的提交作品实际上会被流片,获胜者将获得实物拷贝!比赛截止日期为 2027 年 1 月 18 日,鼓励团队参加。要开始参与,请在 janestreet.com/dwarkesh 下载模板代码
Timestamps
时间戳
(00:00:00) – Steelmanning the case against RSI
(00:00:00) – 为反对 RSI 的观点做最强辩护
(00:18:39) – What’s driving the Chinese labs’ progress
(00:18:39) – 中国实验室进展的驱动力是什么
(00:28:06) – How will automated AI researchers be trained
(00:28:06) – 自动化 AI 研究人员将如何被训练
(00:33:51) – Will long-horizon RL elicit AGI?
(00:33:51) – 长视界强化学习会催生 AGI 吗?
(00:45:24) – The sim-to-real gap
(00:45:24) – 仿真到现实的差距
(01:00:33) – How much progress is explained by data?
(01:00:33) – 有多少进展是由数据解释的?
(01:18:03) – Why is RL working so well?
(01:18:03) – 为什么强化学习如此有效?
(01:24:54) – Move 37 and entropy collapse
(01:24:54) – 第 37 手棋与熵坍塌
(01:28:32) – Rapid-fire timelines
(01:28:32) – 快速问答时间表
Transcript
逐字稿
00:00:00 – Steelmanning the case against RSI
00:00:00 – 为反对 RSI 的观点做最强辩护
Dwarkesh Patel
Dwarkesh Patel
Today, I’m chatting with three of my AI researcher friends from whom I learn a lot every time we talk. They also happen to be at somewhat open-ish labs and companies, so you guys can actually say things on the record. I’m joined by Beren Millidge, who is the CTO of Zyphra, which is developing open source models. John Schulman is the chief scientist at Thinking Machines, previously a co-founder of OpenAI, and led the RLHF work that led to ChatGPT. And Charlie O’Neill is head of model training at Baseten.
今天,我和三位 AI 研究员朋友聊天,每次交谈我都能从他们那里学到很多。他们恰好也在一些相对开放的实验室和公司工作,所以你们可以公开谈论这些内容。与我一起的有 Beren Millidge,他是 Zyphra 的首席技术官(CTO),该公司正在开发开源模型。John Schulman 是 Thinking Machines 的首席科学家,他曾是 OpenAI 的联合创始人,并领导了促成 ChatGPT 的 RLHF 工作。Charlie O’Neill 是 Baseten 的模型训练负责人。
The first question I have: If we’re in 2036 and we don’t have billions of crazy superintelligences running around that have radically transformed the world, what is the most likely reason that doesn’t end up being the case? Other than exogenous political shocks, or there’s a war, or they ban AI or something. What is the most likely technical reason that 2036 isn’t a crazy alien superintelligence world?
我的第一个问题是:如果我们身处 2036 年,却没有数十亿个疯狂运行的超级智能在四处活动并彻底改变世界,那么最可能的原因是什么?排除外部政治冲击、战争或禁止 AI 等情况。2036 年没有变成一个疯狂的异星超级智能世界,最可能的技术原因是什么?
Beren Millidge
Beren Millidge
There’s been a classic thing, almost like Moravec’s paradox, where we think of the AI as, “If it can do this, it’s going to be amazing.” If it can solve these hard maths problems, if it can win at chess, blah, blah, blah… Then it solves these things, and it’s not that impactful. Obviously, it’s somewhat impactful, but not everything.
这有一个经典现象,几乎类似于莫拉维克悖论(Moravec’s paradox),我们认为人工智能如果“能做到这一点,那就太棒了”。比如它能解决那些高深的数学问题,能在国际象棋中获胜,诸如此类……然后它确实解决了这些问题,但影响却没那么深远。显然,它有一定影响力,但并非无所不能。
If somehow that continues, and there’s never the true spark of generalization that occurs, I think that could lead to the AI just being extremely good at everything that people put into a benchmark or put into an environment. But there’s still some persistent sim-to-real gap which is somehow blocking everything. I think this is unlikely. We do actually see this kind of generalization even from RL in practice already. But if it is just ridiculously hard to generalize meta-learning, plus we don’t solve continual learning and it’s just super hard and impossible… This would be my default scenario in that case.
如果这种情况持续下去,而从未出现真正的泛化火花,我认为这可能导致 AI 只是在人们放入基准测试或环境中的任务上表现得极其出色。但仍存在某种持续的模拟到现实(sim-to-real)差距,阻碍着一切。我认为这种情况不太可能发生。事实上,我们已经在实践中看到这种来自强化学习(RL)的泛化现象。但如果元学习的泛化真的难如登天,加上我们无法解决持续学习问题,且这一切既超级困难又不可能实现……那这将是我在这种情况下默认的情景。
John Schulman
John Schulman
I agree with that. Humans have a lot of advantages over models now. Each time a new model comes out, it’ll catch up in some of these areas. But you end up getting bottlenecked by the places where the model is weaker and where it has worse judgment, or the models can’t check themselves well enough.
我同意这个观点。人类目前相比模型拥有许多优势。每出一款新模型,它都会在部分领域赶上人类。但最终你会受到模型较弱之处以及判断力较差环节的瓶颈制约,或者模型无法足够好地自我检查。
There’s this cycle that keeps repeating where a new model comes out and people are blown away and they’re like, “This is it. This is AGI.” But then they use it a bit, and it starts to feel dumb after a month or so. That cycle just might keep going. It’s hard to predict how many times it’s going to repeat.
存在着一个不断重复的循环:新模型问世,人们为之惊叹,认为“就是它了,这就是通用人工智能(AGI)”。但随后使用一段时间后,大约一个月左右,它开始显得笨拙。这个循环可能会一直持续下去。很难预测它会重复多少次。
Right now, you don’t get explosive growth in capabilities because you still get bottlenecked enough when you’re trying to do research and engineering. Even if the model can write way more code than a person, it doesn’t make you 100X more productive. So maybe there are just more of these cycles than we would expect.
目前,你无法获得能力的爆炸性增长,因为在进行研究和工程开发时,你仍然会受到足够的瓶颈限制。即使模型能比人类编写更多的代码,也不会让你提高100倍的生产力。所以,也许这类周期比我们预期的要多。
Charlie O’Neill
查理·奥尼尔(Charlie O’Neill)
For me, it’s a question of how far off the global optimum of “a learner you could have on a chip” is from the transformer + RL, basically the current recipe. People imagine that once you have an agent which is better than all humans at AI research, even if it’s 0.1% better than all humans, then the fact that you can run hundreds of thousands, if not millions, of these in parallel — and you can run them much faster as chips speed up — is going to outweigh every other bottleneck. You’re eventually going to hit this very fast takeoff with regards to self-improvement.
对我来说,这是一个关于“芯片上的学习者”这一全局最优解与 Transformer + RL(强化学习)——也就是当前的基本配方——之间差距有多大的问题。人们设想,一旦你拥有一个在 AI 研究方面优于所有人类的智能体,哪怕它只比所有人类好 0.1%,那么你可以并行运行数十万甚至数百万个这样的智能体——并且随着芯片速度提升,你可以运行得更快——这一事实将压倒其他所有瓶颈。最终,你将迎来自我改进方面的非常快速的起飞。
I could imagine that if we continue along the trajectory that we’re currently on with that paradigm, where it’s basically self-attention, RL, scaling up RL environments… Think about what happened with Moore’s law. We had this very nice straight line and that held for a really, really long time. But there were so many discrete discontinuities and innovations that had to happen to keep that scaling law going. The same thing has happened with LLMs. We had this pre-training scaling law, and then that was hitting diminishing returns. Then we came up with RL and solved that, and then we got this new diminishing returns curve to hit that made it keep looking like a straight line going up.
我可以想象,如果我们继续沿着当前基于该范式的轨迹前进,即基本上是自注意力机制、RL、扩展 RL 环境……想想摩尔定律发生了什么。我们曾有一条非常漂亮的直线,并且这条线维持了非常非常长的时间。但为了保持这种缩放定律的持续,必须发生许多离散的间断和创新。LLM 也发生了同样的事情。我们有了预训练缩放定律,然后那遇到了收益递减。接着我们引入了 RL 并解决了这个问题,然后又出现了新的收益递减曲线,使得整体看起来仍然像一条上升的直线。
So if it requires another one of those discontinuities to solve, I’m not sure that the current method of training LLMs with these RL environments, even RSI-targeted RL environments, would be able to discover that discontinuity. If not, we’re probably going to hit this asymptotic curve.
因此,如果需要另一次这样的间断来解决,我不确定当前使用这些 RL 环境(即使是针对 RSI 的 RL 环境)训练 LLM 的方法是否能够发现这种间断。如果不能,我们可能会遇到这条渐近曲线。
Dwarkesh Patel
达韦什·帕特尔(Dwarkesh Patel)
But do you think the discontinuity will be harder than anything that’s come since 2012?
但你认为这次间断会比 2012 年以来出现的任何情况都更难吗?
Charlie O’Neill
查理·奥尼尔(Charlie O’Neill)
If we had the answer to that, we’d kind of have the ability to implement it. But maybe we should distinguish between a discontinuity which adds to the current paradigm, which is cumulative — there’s something beyond the RL that we have to discover, and maybe they’re capable of connecting the dots in that straight line — or, again, how far off the global optimum are we? Do we have to go back and throw out gradient descent and neural nets in general? I don’t think, if you continue to scale up the current paradigm, an LLM, no matter how many LLMs you’re running, is necessarily capable of discovering that if it’s too far away.
如果我们能得出答案,我们就具备实施它的能力。但也许我们应该区分一种对当前范式有所补充的间断性,这种补充是累积性的——在强化学习(RL)之外还有我们必须发现的东西,他们或许能够沿着这条直线将各个点连接起来——或者再说一次,我们距离全局最优解还有多远?我们是否需要回过头来抛弃梯度下降法以及广义上的神经网络?我认为,如果你继续扩大当前范式的规模,无论运行多少个大型语言模型(LLM),如果差距过大,LLM 未必有能力发现这一点。
Dwarkesh Patel
达瓦克什·帕特尔
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力