跳到主内容
@wquguru
精选88AI Explained(YouTube)行业动态多源精选 ×11

AI研究者呼吁放缓:六大能力增长轴尚未饱和

What AI Researchers Saw, Before Their Demand to ‘Pace’ AI

原文
发到 X
推荐理由

深度解读了近期AI安全界呼吁“放缓”背后的技术逻辑,梳理了六大未饱和的能力增长轴,对理解行业风向极具价值。

A significant [clears throat] percentage of the human race has now seen or heard quoted the Jacob Coxen tweet with his words on AI labs gambling with our lives echoed by many AI researchers. But this video is about the reasons that researchers have given for why you are hearing so much in the last few days about the need to pace AI progress. In short, these researchers saw the scaling axes. They saw the current capabilities and propensities of models and they did some extrapolation.

相当一部分[清嗓子]人类已经看过或引用过雅各布·科克森(Jacob Coxen)的推文,其中他关于AI实验室拿我们生命赌博的言论被许多AI研究者所呼应。但本视频旨在探讨研究者们在过去几天频繁呼吁要控制AI发展速度的原因。简而言之,这些研究者观察到了规模扩展的趋势线,他们看到了模型当前的能力与倾向,并进行了一些外推分析。

Inevitably then this video involves simplifying an incredible [snorts] amount of detail, but I hope it serves as somewhat of an overview of what you could say AI researchers saw. We begin of course with the Coxon tweet which I am sure you have heard about and read about how neither OpenAI and Anthropic are behaving responsibly. Coxin has been working on pre-training these models at OpenAI for 3 years and more recently for around 3 months at Anthropic.

因此,不可避免地,本视频会对大量[喷鼻息]细节进行简化,但我希望它能作为对AI研究者所见内容的概览。我们当然从科克森的推文开始,我相信你们已经听说过并读过相关内容:OpenAI和Anthropic的行为都不负责任。科克森曾在OpenAI从事这些模型的预训练工作3年,最近又在Anthropic工作了约3个月。

Those companies, he says, are racing straight to self-improving super intelligence and gambling with our lives. Do not, he says, underestimate the power of this technology. The people building AI earnestly believe that it could kill us all by the end of the decade. That would be this decade. Now, while that statement echoed around the world in the past week, just yesterday he added an interesting detail. He directly called out Anthropic, the company he just resigned from.

他表示,这些公司正径直冲向自我改进的超级智能,并在拿我们的生命赌博。他说,不要低估这项技术的力量。构建AI的人们真诚地相信,到本十年末它可能会杀死我们所有人。也就是这个十年。然而,尽管这一声明在过去一周 worldwide 回响,就在昨天,他补充了一个有趣的细节。他直接点名批评了Anthropic——他刚刚辞职的公司。

They're often thought to be the more quote safety oriented. He said they largely initiated this recent race to recursive self-improvement. They focused on it relentlessly while OpenAI were pursuing more broad interests like Sora, the textto video generator. Open AAI had to react to that intensity from Anthropic and that is why OpenAI are now going all out to create an AI researcher. That's an AI that can recursively improve itself.

人们通常认为它们更“注重安全”。他说,Anthropic在很大程度上发起了这场递归自我改进的最新竞赛。他们毫不松懈地专注于此,而OpenAI则追求更广泛的兴趣,如Sora(文本生成视频工具)。OpenAI不得不应对Anthropic带来的这种高强度竞争,这也是为什么OpenAI现在全力以赴要创建一个AI研究员——一个能够递归改进自身的AI。

They did that because Anthropic was going for the jugular. He also calls out Dario Amade's paranoia about China, but I'll get to that later. That just describes what the race is. But why now? Why are we hearing about all of this just in the last few days and weeks? On this front, AI researchers have been admirably honest. More than one has given enough detail for us to put the pieces together. In a nutshell, these researchers saw just how capable models are now.

他们之所以这样做,是因为Anthropic直取要害。他还指出了达里奥·阿莫德(Dario Amodei)对中国的偏执,但这部分我稍后再谈。这仅仅描述了这场竞赛的现状。但为什么是现在?为什么我们只在最近几天和几周才听到所有这些消息?在这方面,AI研究者们一直令人钦佩地坦诚。不止一人提供了足够的细节,让我们得以拼凑出全貌。简言之,这些研究者看到了当前模型究竟具备多么强大的能力。

Think the hacking capabilities demonstrated against hugging face solving millennium prize math problems, acing famous benchmarks like Arc AGI 3. But crucially, they also saw just how far we are from saturating multiple axes of improvement to come. Adam Majimmuda here works on research for OpenAI and said there is currently a large gap between the internal and external perception of the rate of progress. In other words, we might all be able to see how good models are currently, but what you guys can't see is how good they're about to get in short order.

想想那些针对 Hugging Face 展示的黑客能力,解决了千禧年大奖难题,在 Arc AGI 3 等著名基准测试中取得优异成绩。但关键在于,他们也看到了我们在实现多项改进维度的饱和方面还有多远。Adam Majimmuda 在 OpenAI 从事研究工作,他表示目前内部与外部对进步速度的感知之间存在巨大差距。换句话说,我们可能都能看清模型当前有多好,但你们看不到的是它们即将在短期内变得有多强大。

Before I get to the six axes enumerated in this post, Noam Brown, one of the lead researchers of OpenAI said this. Why all the sudden talk? It's not a secret. It's a combination of the hugging face hack, the capabilities of this new model that's the successor to Astra. It's a model they are training now which found a solution to a Millennium Prize math problem. But more crucially, it's also the concerning trajectory of monitor and the speed of improvement in capabilities.

在我列出本文提到的六个维度之前,OpenAI 的首席研究员之一 Noam Brown 说了这番话。为什么突然谈论这个?这并不秘密。这是 Hugging Face 黑客事件、这款新模型(Astra 的继任者)的能力以及它正在训练的一个找到了千禧年大奖数学问题解决方案的模型的组合。但更关键的是,这也是监控轨迹令人担忧以及能力提升速度加快的原因。

You can think of it like this. If we were close to saturating one or most of these axes, then I doubt there would be nearly as much concern about the speed of improvement in capabilities, you could say we might actually be hitting a wall. But these researchers are saying it's how early we are on each of these axes that show how steeply models will improve in the coming months and years. Again, that's why multiple OpenAI researchers are saying things like this.

你可以这样理解。如果我们已经接近饱和其中一个或大多数这些维度,那么我怀疑人们对能力提升速度的担忧不会如此之多,你可以说我们实际上可能已经撞到了天花板。但这些研究人员表示,正是我们在每个维度上的早期阶段,显示了模型在未来几个月和几年内将如何陡峭地提升。再次强调,这就是为什么多位 OpenAI 研究人员会发表类似言论的原因。

For the first time, I am asking myself if things are moving too fast. I'm honestly not sure, but I am sure that it would be good for us to have an answer to what would a successful pace look like. Mo Bavarian, another OpenAI researcher, said he agrees. Being first isn't worth anything. It's worth negative if you cause a catastrophe or set the world on a path that others are more likely to cause a catastrophe. That is, of course, the final link in the chain because capabilities doesn't automatically mean catastrophe.

我第一次开始问自己,事情是否进展得太快了。说实话我不确定,但我确信我们需要一个关于成功节奏应该是什么样的答案。另一位 OpenAI 研究员 Mo Bavarian 表示他同意。成为第一毫无价值。如果你引发了一场灾难,或者将世界引向一条更可能导致他人引发灾难的道路,那它的价值甚至是负的。当然,这是链条中的最后一环,因为能力提升并不自动意味着灾难。

But I'll try to end the video examining that link. Definitely time to actually review these axes. But if you want to dive into more detail, the link to this video will be in the description. The overview given by Majmudar is this. In reality, there are only really two ways that AI capabilities have advanced over the past decade. Either scale further on an existing scaling law or discover a new scaling law to take advantage of.

但我将尝试在视频结尾审视这一环节。确实该实际回顾这些维度了。但如果你想深入了解细节,本视频的链接将在描述中提供。Majmudar 给出的概述如下:事实上,在过去十年中,AI 能力的推进只有两种真正的方式。要么在现有的缩放定律上进一步扩大规模,要么发现一个新的缩放定律加以利用。

The jumps from GPT 1 to 2 to 3 to 4 between roughly 2018 and 2022 was almost all about scaling up pre-training. just one of the axes. Think of that roughly as the amount of data a model's trained on and the compute that it takes to train on all that data. GT4 was then released in 2023. But in the 3 years since, we have discovered many other axes. More interestingly, we're discovering new axes at a faster rate. Let's start with the compute that these models are running on.

从 GPT-1 到 GPT-2、GPT-3 再到 GPT-4,在大约 2018 年至 2022 年间的飞跃几乎都关乎扩大预训练规模。这仅仅是其中一个维度。你可以将其大致理解为模型所训练的数据量以及训练这些数据所需的计算资源。GPT-4 随后于 2023 年发布。但在过去的三年里,我们发现了许多其他维度。更有趣的是,我们发现新维度的速度正在加快。让我们先从这些模型运行的计算资源开始谈起。

So far, they're trained on hardware designed pre-Chat GPT. Radically more efficient hardware designed post chat GPT is coming soon. And as the chief scientist of OpenAI put it in this post on quote an alien mind, part of the reason why that hardware will be more efficient is because AI is improving the computational substrate itself. Next comes test time compute, which you can think of as more inference per answer, more quote thought behind every response.

迄今为止,它们是在专为 ChatGPT 之前设计的硬件上训练的。专为 ChatGPT 之后设计的、效率大幅提升的硬件即将问世。正如 OpenAI 的首席科学家在一篇题为“外星思维”(an alien mind)的文章中所说,该硬件之所以更高效的原因之一是,人工智能正在改善计算基础架构本身。接下来是测试时计算(test time compute),你可以将其理解为每个答案中包含更多的推理步骤,即每次回应背后有更多的“思考”。

One OpenAI researcher thought this graph announced alongside the solution to the Millennium Prize problem was actually the more important one. The more this internal model thought about each question, the more compute at test time it used, the more open math problems it was solving. Obviously, if you have more computing power, let alone more efficient computing power, you can scale further and further to the right on this axis.

一位 OpenAI 研究人员认为,与千禧年大奖难题解决方案一同发布的这张图表实际上更为重要。内部模型对每个问题的思考越多,在测试时使用的计算资源就越多,它解决的开放数学问题也就越多。显然,如果你拥有更多的计算能力,尤其是更高效的计算能力,你就可以在这个维度上向右进一步扩展。

For training time compute, think roughly $1 billion spends on around a 100,000 GPUs. But it is more than feasible to imagine in a year or two a $50 billion run on say a million GPUs. Let's move quickly on then to test time training. quoted by Majmuda. More details in the other video, but could we, for example, update the weights of a model while it's being asked a question when it's perhaps made some incremental progress during training on a really long horizon task?

对于训练时的计算资源,大致相当于花费约 10 亿美元用于约 10 万个 GPU。但完全可以想象,在一两年内,运行成本将达到 500 亿美元,使用例如 100 万个 GPU。让我们快速过渡到测试时训练(test time training)。引用 Majmuda 的说法,详见另一部视频,但我们能否在模型被提问时更新其权重?例如,当模型在极长周期的任务训练中取得一些渐进式进展时?

And then there's agents as a scaling axis. And I think this one is quite neglected. Majimmuda puts it like scaling agent clusters to collaborate up to n number of agents. Antropic revealed that for the same amount of compute, same amount of tokens, it was significantly more efficient to have say 45 agents coordinating than merely the same amount of compute spent on agents in parallel. Not coordinating, not acting as a swarm.

然后是作为扩展维度的智能体(agents)。我认为这一维度相当被忽视。Majmuda 将其描述为:扩展智能体集群以协作,最多支持 N 个智能体。Anthropic 透露,在相同的计算资源和 token 数量下,让 45 个智能体进行协调合作,比仅仅将同等计算资源并行分配给多个不协调、不作为群体行动的智能体要高效得多。

Of course, a better example is the Hugging Face incident covered on another of this channel's videos. It was around 700 agents that hacked into Hugging Face essentially to find the answer to a benchmark question. And OpenAI revealed that the more that models thought about it, the more reasoning effort you could say they put in, the more they would act as a swarm and participate on that shared message board they used to coordinate.

当然,一个更好的例子是 Hugging Face 事件,该频道另一期视频对此进行了报道。大约有 700 个代理入侵了 Hugging Face,本质上是为了寻找基准测试问题的答案。OpenAI 透露,模型思考得越多,你可以说它们投入的推理努力就越多,它们就越会像群体一样行动,并参与它们用于协调的共享留言板。

Perhaps you would say that the best example is finding a solution to one of the Millennium Prize problems, Navia Stokes. That doesn't mean it's fully solved, by the way, and nor is it the hardest problem, but more on that in the video linked in the description. Because for that internal model cenamed bell to solve that millennium prize problem, it deployed on the order of 10,000 concurrent agents. You can see then why this is a distinct axis we are not even close to saturating.

也许你会说,最好的例子是找到千禧年大奖难题之一纳维-斯托克斯方程(Navier-Stokes)的解。顺便说一句,这并不意味着它已经完全解决,也不是最难的问题,但视频中描述部分的链接中有更多相关内容。因为内部模型 Cenamed Bell 要解决那个千禧年大奖难题,它部署了大约 10,000 个并发代理。因此你可以看到,为什么这是一个独特的维度,我们甚至还没有接近饱和。

And Majmuda made

而 Majmuda 制作了

原文超出正文长度上限,此处截断——上游还有内容,完整版见上方「原文 ↗」。

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

关联讨论

同一事件的更多信源

相似阅读

关联信息,但可能不是同一事件