跳到主内容
@wquguru
精选85Dwarkesh Podcast(RSS)技巧与观点

Ryan Greenblatt:AI 自动化 AI 研究后会发生什么

Ryan Greenblatt – What happens once AI can automate AI research?

原文
发到 X
推荐理由

关注 AI 对齐与递归自我改进的从业者必听,Greenblatt 给出了具体时间线与论证框架,值得深入思考。

Had Ryan Greenblatt on to discuss/debate recursive self-improvement.

This might be the most important question in the world right now – whether within a year or so of achieving human-level intelligence, you slingshot towards having 10s of billions of superintelligences, each of which is dramatically more competent than human experts across all fields.

I’ve historically been skeptical of this possibility. My intuition has been that we will end up significantly bottlenecked by not only compute scaling but human expert data, which I think underlies most of the AI progress today.

If, because of RSI, we got a jump as big as GPT-3 to a Mythos (i.e. 6 years of AI progress) within a single year of achieving AGI, then the thing we get there at the end of that year is definitively and wildly superhuman.

We hashed it out, and I think Ryan made a pretty good case that this kind of speedup is plausible. FWIW, Ryan’s median for when we automate AI R&D is 2031.

We then discussed the alignment implications of this scenario. Who should these superintelligences be aligned to? In the future, our capacity to steward our votes and our capital, and to make sense of what’s happening in the world, will all be titrated by superintelligences. And I worry that specs like the Claude Constitution are not shaping these ASIs to truly be my personal advocates and guardian angels.

And can we get them aligned to anything in the first place? Ryan and I had a long debate about whether the kind of reward hacking we saw with the OAI/Hugging Face hack extrapolates to superintelligences that would team up to literally take over the world.

The first piece of advice you get when you’re learning to drive is that it will go much smoother if you look at the horizon instead of directly in front of your tires. And so it is with the trajectory of AI. Hope you enjoy!

Watch on YouTube; listen on Apple Podcasts or Spotify.

Sponsors

  • Antithesis is a software testing platform that finds the failures no human or AI could ever anticipate. It runs thousands of copies of your code inside a fully deterministic computer, injecting faults and steering each trajectory toward the most insidious bugs. This lets you find critical issues in minutes rather than waiting months for your users to uncover them. Learn more at antithesis.com/dwarkesh
  • Jane Street’s back with a new puzzle. They designed an ASIC and sent me the final masks… but they didn’t tell me what the chip actually does. So that’s the challenge: reverse engineer the circuit and figure out the chip’s purpose. Jane Street has a bunch of swag ready to send to the most creative solutions, and they’re also planning to feature the top write-ups in a blog post. Download the files and get started at janestreet.com/dwarkesh
  • Cursor and SpaceX recently released Grok 4.5, and I’ve been surprised by just how good the model is. For example, when I tested it against Fable and Sol on a bunch of AI governance questions, all three models gave substantially the same answers, but Grok was faster, more concise, and cheaper. Grok 4.6 is coming soon, but in the meantime, you can try 4.5 at cursor.com/dwarkesh

Timestamps

(00:00:00) – Is AI R&D verifiable enough to unlock recursive self-improvement?

(00:16:52) – Is AI progress bottlenecked by human expert data?

(00:34:02) – Flat token prices suggest scaling has been slow

(00:39:47) – Skills AI can’t train on: does it even need them?

(00:48:07) – Aligned to whom?

(01:09:18) – Recent incidents of AIs colluding and deceiving humans

(01:19:38) – What could possibly go wrong? A concrete scenario

(01:48:02) – From reward hacking to takeover

Transcript

00:00:00 – Is AI R&D verifiable enough to unlock recursive self-improvement?

Dwarkesh Patel

Today I’m chatting with Ryan Greenblatt, who is the chief scientist at Redwood Research, where he focuses on technical AI safety and security work.

I want to talk to you about recursive self-improvement. This is the idea that once we build human-level intelligences, they quickly slingshot towards tens of billions of superintelligences, which are each individually more competent than the top human experts across every field. Whether or not this turns out to be the case is probably the most important question in the world right now. And historically, I’ve been quite skeptical that this kind of thing happens, but, you seem to think that it might be plausible, and so I wanted to hear the case for it.

Ryan Greenblatt

Let’s talk about this. First, I think it’s worth noting that AI R&D is a type of task at which the AIs are especially good, because the companies are trying really hard to make their AIs good at AI R&D. It’s also the kind of domain that has a lot of nice properties from the perspective of how AI development works right now. It’s pretty verifiable. You can do a bunch of stuff iteratively, and it’ll hill climb on various metrics.

I think once you have AIs which are roughly matching the top human experts in AI R&D, that could kick off a feedback loop where the AIs are doing AI research. That produces smarter AIs. That feeds back in. That feedback loop could be strong enough that you end up with a lot of progress in a short period of time. Maybe my median expectation is something like four or five years of AI progress in a single year. This requires really overcoming a huge amount of diminishing returns in research and basically doing the equivalent of the progress we would have gotten after a really large compute scale-out. So this is a pretty impressive, big thing.

It’s worth keeping in mind that five years of AI progress, four years of AI progress, even three years of AI progress, is really a lot of fucking AI progress. A little over three years ago, GPT-4 had come out. Right now, of course, we have Mythos 5 or whatever, and maybe a somewhat better model that Anthropic has internally. That is just a huge amount of progress in a bit over three years. If we’re talking about five years, then maybe we’re talking more about a jump from GPT-3 to Mythos 5 or whatever.

Dwarkesh Patel

I think this argument has three different parts. Now I want to evaluate each one of them. First is the argument that AI R&D is very verifiable. Second is the argument that if you automate AI R&D, you could get four or five years of progress in a single year. Third is the argument that what comes out the other end of four or five years of AI progress at the current pace, starting at the point whenever AI R&D is automated, is an AI where you can drop it on the job at basically anything you can imagine.

You can drop it in Texas politics in the 1940s, and it outmaneuvers Lyndon Johnson. You can drop it in TSMC, and it learns how to do better process engineering at TSMC. It’s certainly a better video editor… My video editors are very excellent, but it is just, in general, better than humans at any given job that it finds itself trying to do.

So I want to evaluate all of these sub-arguments that lead to basically getting ASI pretty soon after this benchmark, which you’re expecting by 2030 or something, right?

Ryan Greenblatt

I would say that I expect full automation of AI R&D perhaps somewhere around 2031, 2030. Getting to the “beats all humans on the job” milestone, maybe my median expectation is around 2033. But if I see AIs fully automating AI R&D, I think I’m expecting that probably within a year. The way the forecasting works out, the difference between medians is bigger than the median difference between milestones. Anyway, whatever.

Dwarkesh Patel

By the way, there’s this meme on the internet. Every time I’m trying to ask about people’s timelines, when I’m asking Dario or somebody, I’m always like, “Okay, how long before you automate my video editors?” There’s this meme of my video editor editing the podcast every time I listen to this.

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

相似阅读

另一事件,读法相近