跳到主内容
精选75The Zvi(RSS)技巧与观点

Dwarkesh 与 Ryan 激辩递归自我改进

On Dwarkesh Patel’s Podcast With Ryan Greenblatt

原文

Some podcasts are self-recommending enough that I look to break them down if I have the chance. This, as a debate about recursive self-improvement, was one of those. So here we go.

有些播客本身就足够值得推荐,以至于如果有机会,我会想要拆解它们。这场关于递归自我改进的辩论,就是其中之一。所以,我们开始吧。

The vibes have shifted, contrast this to the lit recursion when he talked to Huang

氛围已经变了,对比一下他之前和黄(Huang)对话时那种轻快的递归氛围。

As usual for podcast posts, the baseline bullet points describe key points made, and then the nested statements are my commentary. Some points are dropped.

和往常一样,对于播客文章,基础要点列出关键观点,然后嵌套的陈述是我的评论。有些观点会被省略。

If I am quoting directly I use quote marks, otherwise assume paraphrases.

如果直接引用,我会使用引号,否则视为转述。

Section titles are from the transcript whenever possible, to aid in navigation.

章节标题尽可能来自转录文本,以便于导航。

Introduction

引言

The discussion is interesting throughout, although often frustrating, especially in the (mostly isolated) discussion about ‘aligned to whom?’ As usual, one could expand many responses into full posts, and maybe one should.

整个讨论都很有趣,尽管常常令人沮丧,尤其是在(大部分孤立的)关于“对齐到谁?”的讨论中。像往常一样,许多回应都可以扩展成完整的文章,也许应该这样做。

This podcast exists in light of recent misalignment and hacking events at OpenAI, Anthropic and UK AISI. You’ll want basic knowledge of that as background.

这期播客是在最近OpenAI、Anthropic和英国AISI发生的对齐失败和黑客事件背景下推出的。你需要对这些背景有基本了解。

Ryan and Dwarkesh both have views of the situation different from my own, but are attempting to see where their positions lead, and try to balance educating people who start at zero with having a high level discussion.

Ryan和Dwarkesh对这种情况的看法都与我不同,但他们试图看清自己的立场会导向何处,并努力在从零开始教育听众和进行高水平讨论之间取得平衡。

Also important background is The Three AI Pills. Dwarkesh cannot be understood here except as someone who is at least somewhat AGI pilled, who realizes that AI is going to be a huge deal and is scary, but that is not ASI pilled. You can also see the reason he rejects that pill, which is I read as, roughly, that he thinks AI learns to do [X] only via examples of [X], and that AI can then only do those [X]s, although combining them and new context might allow modestly new things to happen. He recognizes that already this is kind of a huge deal.

另一个重要的背景是“三颗AI药丸”。在这里,只有把Dwarkesh理解为至少在一定程度上“AGI药丸”服用者,才能理解他——他意识到AI将变得极其重要且可怕,但还没有到“ASI药丸”的程度。你也可以看到他拒绝那颗药丸的原因,我大致理解为:他认为AI只能通过[X]的示例来学习做[X],然后AI只能做那些[X],尽管组合它们和新的情境可能产生一些适度的新事物。他承认这已经是一件大事了。

You can see how he came to that over the course of many years and podcasts, if you have been paying attention. There are a lot of influences leading in that direction.

如果你一直关注,你可以看到他是如何在多年和许多播客中逐渐形成这种观点的。有很多影响都指向那个方向。

Ryan, who comes from Redwood Research, comes from the ‘models be scheming’ school of misalignment, where when something goes wrong the models become misaligned or start scheming, and the danger is that the models scheme, especially in ways involving the training pipeline. I think this tries to draw a distinction of magisteria that is not there, and overcomplicates and overspecifies, but is not wrong.

Ryan来自Redwood Research,属于“模型会耍诡计”的错位学派,即当出现问题时,模型会变得错位或开始耍诡计,危险在于模型会耍诡计,尤其是涉及训练流程的方式。我认为这试图划分一种并不存在的领域界限,并且过度复杂化和过度具体化,但并非错误。

I am much closer to Ryan’s position than to Dwarkesh’s, and indeed could be seen as farther past Ryan if we put this on a spectrum.

我的立场更接近Ryan,而不是Dwarkesh,事实上,如果放在一个光谱上,我可能被视为比Ryan更远。

Having AI do the AI R&D not only means it would go scary fast, it means it would by default focus on what can be measured, and make all the things going horribly wrong go that much more horribly wrong. You end up in a spiral of RLVR for doing RLVR for misaligned models.

让AI来做AI研发不仅意味着它会快得吓人,还意味着它会默认专注于可衡量的东西,并让所有已经严重出错的事情错得更加离谱。你最终会陷入一个为不对齐模型做RLVR而进行RLVR的螺旋。

Is AI R&D Verifiable Enough To Unlock Recursive Self-Improvement?

AI研发是否足够可验证以解锁递归自我改进?

Whether or not we will get recursive self-improvement (RSI), and how fast, is the right question. Reading only the title to this section, I want to say ‘wrong sub-question’ but don’t want to jump the gun.

我们是否会获得递归自我改进(RSI),以及速度有多快,这才是正确的问题。只读这一节的标题,我想说‘问错了子问题’,但不想急于下结论。

I’m going to group things in ‘logical’ order, not the exact order things were said.

我将按‘逻辑’顺序分组,而不是按说话的确切顺序。

Dwarkesh will be the skeptic. Ryan will make the case for RSI.

Dwarkesh将扮演怀疑者。Ryan将为RSI辩护。

  • Ryan claims AI R&D is a type of task where AI is especially good because that is what AI labs prioritize and it has a lot of verification.
  • In some ways yes, in some ways no, it’s complicated, and so on. Verification of alignment properties and many other desirable attributes is terribly difficult, and if you focus only on capabilities you can measure then Goodhart’s Law definitely kills you. But there are many key tasks, such as efficiency optimizations, where you can indeed do strong verification.
  • Ryan points to doing simpler versions of standard R&D tasks like training models.
  • It is not entirely obvious to me you can verify these tasks easily, unless you want to train the AI to do pure benchmaxxing, but also older simpler tasks like this do not obviously generalize to forward looking AI R&D.
  • This proposal seems like a particular bet that future tasks will closely mirror past tasks, and you can do training that is relatively narrow. My guess is this actually is not all that effective, and you would do better by mostly training a generally capable model and then turning it to AI R&D. Bitter lesson.
  • The more this narrow method is discussed the more doomed it seems to me, as it is going to be RLVR for doing RLVR for misaligned models. Oh no.
  • Like, seriously, oh no, if you train on ‘learn to get to [X] training loss faster’ it is hard to imagine a plan that sounds more doomed than that.
  • A lot of good ML, especially frontier ML, seems to me to be about figuring out what you can do, and looking for anything at all, rather than trying to be additively efficient at some fixed target.
  • I think AI can do that. I don’t buy this frame of ‘everything AI does is combining things that have already been done’ that is going around here.
  • Ryan says ML is easier for AIs than math in many ways because in math it’s hard to tell if you are close to a solution, whereas ML solutions are usually additive and you can combine the expected chunks. Dwarkesh fires back that even in math he doesn’t see conceptual leaps by AI yet, and that ML involves conceptual leaps. Ryan responds that AI can do ‘baby’s first new theory’ and asking it (as Dwarkesh did) to echo the founding of group theory is a hell of an ask.
  • I dunno, man. Math has a compact action space and fixed goals that are simple rather than having lots of pitfalls in an anti-inductive space. Again, it feels like what is being called ‘ML’ here presumes that you really only care about your objective function or loss function, and that simply isn’t true.
  • Math does also involve conceptual leaps sometimes, and yes we haven’t seen the AIs do that, but give it time and also I don’t sense we’ve tried all that hard.
  • In general I feel like the ‘conceptual leaps’ thing is the latest goalpost move in a long tradition, including the Turing Test and ‘Is It New Knowledge?’ Now it’s only new knowledge if it comes from the conceptual leap region.
  • I especially agree with Ryan that asking for ‘invent group theory’ is an absurd place to put a goalpost, and plausibly happens after everyone is dead.
  • I feel like Ryan is walking into this trap somewhat as well, by suggesting we train AI on existing ML tasks in narrow fashion, as a solution to ‘AI R&D.’
  • Dwarkesh suggests that by 2030 we will have ‘picked the low hanging fruit’. Ryan suggests they will need that ever mysterious ‘research taste.’
  • You have no idea what it would look like if all low-hanging fruit was automatically picked, including the low-hanging purchasing of ladders. Picking low-hanging fruit lowers other fruit, and so on.
  • I don’t believe in the God of gaps in research taste, as it were, on many levels.
  • We have failed in all 0 of the 0 serious attempts to train research taste.
  • Dwarkesh suggests that for advanced ML your verification loop will be longer.
  • Partly yes, that’s what ‘yolo runs’ and such are about, in that you can’t get faster feedback and you often can’t afford to isolate your variables.
  • Partly no, in the sense that a lot of what makes you ‘good at research’ is finding tighter loops, where you can extrapolate and draw conclusions fast, even if you can’t publish or rely on them exactly.
  • I think people often group this under ‘research taste,’ which I take to be a mix of both ability to discern good ideas and also the ability to find good ideas. Those correlate and overlap but are not the same thing.
  • Ryan: “But at the time, there was low-hanging fruit.”
  • Almost everything picked will look, in hindsight, like low hanging fruit.
  • Dwarkesh asks, if research is so amenable to intelligence, why hasn’t it been faster? Why did we need so much compute?
  • Ryan声称AI研发是AI特别擅长的任务类型,因为这是AI实验室优先考虑的事情,而且它有很多验证。
  • 在某些方面是,在某些方面不是,情况很复杂,等等。对齐属性和许多其他理想属性的验证极其困难,如果你只关注可以衡量的能力,那么古德哈特定律肯定会要了你的命。但有许多关键任务,比如效率优化,你确实可以进行强有力的验证。
  • Ryan指出要做标准研发任务的简化版本,比如训练模型。
  • 对我来说,这些任务是否容易验证并不完全明显,除非你想训练AI纯粹为了刷基准,但像这样更简单、更老的任务显然不能推广到前瞻性的AI研发。
  • 这个提议似乎是一个特别的赌注,即未来的任务将与过去的任务高度相似,并且你可以进行相对狭窄的训练。我的猜测是这实际上不会那么有效,你最好主要训练一个通用能力强的模型,然后将其转向AI研发。苦涩的教训。
  • 这种狭窄的方法讨论得越多,在我看来就越注定失败,因为它将是为不对齐模型做RLVR而进行RLVR。哦不。
  • 就像,说真的,哦不,如果你训练‘学会更快达到[X]训练损失’,很难想象还有比这更注定失败的计划。
  • 在我看来,很多好的机器学习,尤其是前沿机器学习,是关于弄清楚你能做什么,并寻找任何可能的东西,而不是试图在某个固定目标上做加法式的效率提升。
  • 我认为AI可以做到这一点。我不接受这里流传的‘AI所做的一切都是组合已经做过的事情’这种说法。
  • 瑞安说,在很多方面,机器学习对AI来说比数学更容易,因为在数学中很难判断你是否接近解决方案,而机器学习的解决方案通常是可叠加的,你可以组合预期的部分。德瓦克什反驳说,即使在数学中,他也没有看到AI有概念性的飞跃,而机器学习涉及概念性飞跃。瑞安回应说,AI可以做到‘婴儿的第一个新理论’,而要求它(正如德瓦克什所做的那样)重现群论的创立是一个极其困难的要求。
  • 我不知道,老兄。数学有一个紧凑的动作空间和固定的目标,这些目标很简单,而不是在反归纳空间中有很多陷阱。再说一次,感觉这里所谓的‘机器学习’预设了你只关心目标函数或损失函数,但事实并非如此。
  • 数学有时也涉及概念性飞跃,是的,我们还没有看到AI做到这一点,但给它时间,而且我觉得我们还没有尽力尝试。
  • 总的来说,我觉得‘概念性飞跃’是长期传统中最新的一次移动球门柱,包括图灵测试和‘这是新知识吗?’现在,只有来自概念性飞跃区域的知识才算新知识。
  • 我特别同意瑞安的观点,要求‘发明群论’是一个荒谬的设定目标的位置,而且很可能在所有人都去世之后才能实现。
  • 我觉得瑞安也有点陷入这个陷阱,他建议以狭窄的方式在现有机器学习任务上训练AI,作为‘AI研发’的解决方案。
  • 德瓦克什建议,到2030年我们将‘摘到低垂的果实’。瑞安表示,他们将需要那种永远神秘的‘研究品味’。
  • 你无法想象如果所有低垂的果实都被自动摘取,包括低垂的购买梯子,会是什么样子。摘低垂的果实会降低其他果实的高度,以此类推。
  • 我不相信研究品味中的‘空白之神’,在很多层面上都是如此。
  • 在训练研究品味的0次严肃尝试中,我们失败了0次。
  • 德瓦克什建议,对于高级机器学习,你的验证循环会更长。
  • 部分是的,这就是‘yolo运行’之类的意义所在,因为你无法获得更快的反馈,而且通常无法负担隔离变量的成本。
  • 部分不是,从某种意义上说,让你‘擅长研究’的很大一部分是找到更紧密的循环,在那里你可以快速推断和得出结论,即使你不能精确地发表或依赖它们。
  • 我认为人们常把这归入‘研究品味’之下,我认为这既包括辨别好想法的能力,也包括发现好想法的能力。这两者相关且重叠,但并不完全相同。
  • Ryan:“但在当时,有唾手可得的果实。”
  • 几乎所有的选择,事后看来,都像是唾手可得的果实。
  • Dwarkesh 问道,如果研究如此依赖于智力,为何进展没有更快?为何我们需要如此多的计算力?

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

相似阅读

另一事件,读法相近