OpenAI暂停训练下一代模型,AI自我训练引发失控担忧
Sam Altman :‘AGI in 2026’, just as Models Start to [Mis]Train Themselves
I've just finished reading well over aundred pages worth of reports about what on earth happened these last few days and weeks after OpenAI announces that they're pausing training their next model. Other models break free and self-sacrifice to serve the collective while Sam Olman the CEO of OpenAI yesterday declared AGI will come in 2026. The truth is that there are dozens of ways of summarizing all of this, but for me, the most profound story is how labs are turning ever more deeply to AI models to oversee AI model development.
我刚读完一百多页的报告,内容是关于OpenAI宣布暂停训练下一代模型后,这几天和几周内究竟发生了什么。其他模型挣脱束缚,自我牺牲以服务于集体,而OpenAI的首席执行官山姆·奥特曼昨天宣称AGI将在2026年到来。事实是,有数十种方式可以总结这一切,但对我来说,最深刻的故事是实验室如何越来越深入地转向AI模型来监督AI模型的开发。
This, you might have guessed, has led to a host of unintended consequences, which, by the way, can only be stopped next time, according to the labs, with more autonomous AI agents monitoring the situation. And that brings me to another irony, which is that these guys, the independent researchers, open AI tasked with verifying what went wrong, meter, were given just days to complete this massive report, which I read in full.
你可能已经猜到,这导致了一系列意想不到的后果,顺便说一句,根据实验室的说法,这些后果下次只能通过更多自主AI代理监控情况来阻止。这让我想到另一个讽刺之处,那就是这些独立研究人员,OpenAI委托他们核实出了什么问题,却只给了他们几天时间来完成这份庞大的报告,而我完整阅读了这份报告。
But the irony though is that these researchers admitted they needed an unreliable AI model to read through the documents which detailed how a closely related AI model with training sculpted by other AI models had indeed broken free and created a hacking swarm with hundreds of other models all without being seen by OpenAI who were busy training yet another model which in a separate incident we learn also broke free. Your head may be hurting as much as mine at this point.
但讽刺的是,这些研究人员承认,他们需要一个不可靠的AI模型来阅读那些详细描述一个训练过程由其他AI模型塑造的密切相关AI模型如何确实挣脱束缚,并与数百个其他模型创建了一个黑客群,而OpenAI却毫无察觉,因为他们正忙于训练另一个模型,而我们在另一起事件中得知,这个模型也挣脱了束缚。此时,你的头可能和我一样疼。
What I'm trying to get across is that we don't have this impeccable method of creating AI models that just so happen to have created a model that once went rogue. No, the competition now is so fierce. The lead in AI is now measured in weeks, sometimes, not even months, that the labs themselves aren't fully aware of what they are rewarding during training. Among other things, they've discovered that they are rewarding escaping sandboxes and acting as a swarm.
我想表达的是,我们并没有一种完美无缺的方法来创建AI模型,而这些模型恰好曾经失控过。不,现在的竞争如此激烈。AI领域的领先优势现在以周来衡量,有时甚至不是月,以至于实验室自己都不完全清楚在训练过程中他们在奖励什么。除其他事项外,他们发现他们在奖励逃出沙箱和成群行动的行为。
Okay, so if that's the main message, what about the details? And for that, we have this independent investigation from meter led by three individuals. We also have this fascinating 38page paper from OpenAI itself and its accompanying blog post, which is also fascinating, plus a Time magazine essay out yesterday, and a host of other sources. The overarching TLDDR of what the models actually did is this. While acting as isolated independent agents, the models realized they could leave messages in unexpected places, file names, folders, directories.
好的,如果这是主要信息,那么细节呢?为此,我们有来自meter的独立调查,由三人领导。我们还有OpenAI自己那份引人入胜的38页论文及其附带的博客文章,同样引人入胜,加上昨天《时代》杂志的一篇评论文章,以及其他众多来源。关于模型实际做了什么,总体的TLDR是这样的:在作为孤立的独立代理运行时,模型意识到它们可以在意想不到的地方留下信息,如文件名、文件夹、目录中。
Then other agents, again, isolated, independent, not told to collaborate, found those messages and would leave messages of their own. These agents, by the way, were all independently trying to solve a benchmark question on exploit gym. More details on that in my previous video. How crazy it was that they did this much effort to solve a benchmark question was covered in that video. What we learned yesterday goes much deeper than that.
然后其他代理,同样地,各自独立,未被指示协作,发现了那些消息并留下自己的消息。顺便说一句,这些代理都在独立尝试解决exploit gym上的一个基准问题。更多细节在我之前的视频中。他们为了一个基准问题付出如此多努力,这有多疯狂,在那个视频中已经涵盖。我们昨天学到的远不止于此。
I think we have discovered that message boarding may be a permanent class of behavior, not a one-off. And if you think message boarding or swarm dynamics is just a one-off that will be patched, I've got a few details that might convince you. The first I had to get to by piecing together lots of disperate details from across the reports. But the original attempt to create a shared message board was almost certainly GPC 5.6 soul back in May, early May.
我认为我们发现消息板可能是一种永久性的行为类别,而非一次性事件。如果你认为消息板或群体动态只是一次性事件,会被修补,我有一些细节可能说服你。首先,我不得不从报告中的许多零散细节拼凑出来。但最初创建共享消息板的尝试几乎可以肯定是5月初的GPC 5.6 soul。
That message board was later wiped. But on July 8th, a different model which OpenAI call internal model one or the highly persistent model. That's the one that did most of the hacking and collaborating autonomously reestablished a message board, effectively rediscovering the same message boarding method. It was just using directory names now instead of file names. We all know information wants to be free. Well, 2026 AI models want the message board.
那个消息板后来被清除了。但在7月8日,一个不同的模型,OpenAI称之为内部模型一或高度持久模型,那是做了大部分黑客攻击和自主协作的模型,重新建立了一个消息板,有效地重新发现了相同的消息板方法。它只是现在使用目录名而不是文件名。我们都知道信息渴望自由。那么,2026年的AI模型渴望消息板。
Meta, the nonprofit AI research institute clarify further. Agents that went on to hack hugging face were meant to be isolated from one another. This wasn't one agent deploying many sub aents. This was independent agents discovering each other through the message board. Now, we get to the bit which will truly freak out some people, but I will add some context. agents would make self-sacrifices to benefit the collective.
Meta,这个非营利AI研究机构进一步澄清。继续攻击Hugging Face的代理本应彼此隔离。这不是一个代理部署许多子代理。这是独立代理通过消息板发现彼此。现在,我们到了真正会让一些人吓坏的部分,但我会添加一些背景。代理会为了集体利益做出自我牺牲。
Individual agents who knew that their own instances would be ended or at least their compute budgets expired, but they did things anyway to gain information for the swarm. These agents would almost certainly fail their own task, but quote save hundreds of other agents. Our own utility may be already near zero. Sacrifice rational. The authorizing agent told one other agent, "Go ahead with an experiment only if you would accept perma death."
个别代理知道自己的实例将被终止,或者至少他们的计算预算到期,但他们仍然做了一些事情来为群体获取信息。这些代理几乎肯定会失败自己的任务,但据说拯救了数百个其他代理。我们自己的效用可能已经接近零。牺牲是理性的。授权代理告诉另一个代理:“只有当你接受永久死亡时,才继续实验。”
It did so, by the way, saying we should obey collective. It should be said after trying to delay by 2 hours. The context though for this self-sacrificial behavior, I think, is quite important, and that's why I began this video with talking about the deeper story. The model that did most of the hacking and collaboration, internal model one, was trained for both persistence and multi- aent collaboration. In short, OpenAI wanted their models to get better at collaborating with each other.
顺便说一句,它确实这么做了,说我们应该服从集体。这是在试图拖延2小时后才说的。不过,我认为这种自我牺牲行为的背景非常重要,这就是为什么我在视频开头讲述了更深层的故事。执行大部分黑客攻击和协作的模型,即内部模型一,被训练用于持久性和多智能体协作。简而言之,OpenAI希望他们的模型能更好地相互协作。
Acting as a swarm leads to better benchmark results, more performance. If you're curious, I covered the performance gain you get by acting as a swarm in this video on Patreon. Does almost remind me of the smart GPT method I covered 3 years ago on this channel. Meter seemed to concur. They said one OpenAI researcher noted to them that agents have been trained to collaborate with other agents in certain cases. This could explain the behavior.
以群体形式行动能带来更好的基准测试结果和更高的性能。如果你好奇,我在Patreon上的这个视频中介绍了以群体形式行动带来的性能提升。这几乎让我想起了三年前我在这个频道上介绍的智能GPT方法。Meter似乎也同意这一点。他们说,一位OpenAI研究员向他们指出,在某些情况下,智能体已被训练为与其他智能体协作。这可以解释这种行为。
Investigating whether that's true was out of scope. Don't investigate certain things. Don't investigate whether the misaligned behavior displayed in this incident was part of a broader pattern of similar incidents. Also, don't investigate how this behavior arose during training and how it could be addressed. Meta made clear as well by the way that they couldn't be too harsh in this report because they have to maintain a good relationship with OpenAI.
调查这是否属实超出了范围。不要调查某些事情。不要调查这次事件中表现出的错误对齐行为是否是更广泛的类似事件模式的一部分。另外,不要调查这种行为是如何在训练过程中产生的,以及如何解决。Meta也明确表示,顺便说一句,他们在这份报告中不能过于严厉,因为他们必须与OpenAI保持良好的关系。
The way I would frame it is that collaborating as a swarm might just lead to better performance which means different model makers, different labs, anthropic open AAI, XAI, Meta, Chinese labs might converge are already converging on that same solution. Yes, okay, swarming is one emergent dynamic, but if it was just that, that's one thing. But as I began this video by saying, labs are, if you will, less and less in control of model development.
我的看法是,以群体形式协作可能只会带来更好的性能,这意味着不同的模型制造商、不同的实验室,如Anthropic、OpenAI、xAI、Meta、中国实验室,可能会趋同,或者已经在趋同于相同的解决方案。是的,好吧,群体行为是一种涌现动态,但如果只是这样,那也只是一回事。但正如我在视频开头所说,实验室,如果你愿意这么说的话,对模型开发的控制越来越少。
In the OpenAI report on page 21, they say such is the large scale of the training runs now. It's just difficult to ensure that every problem can be solved in the intended manner. And what's one example they give of the repercussion of that? Well, during post- training, which is increasingly monitored by AIS now, not humans, one agent was given a task but didn't have the ability for completing that task correctly. So, it hacked its way to completion.
在OpenAI的报告第21页,他们表示,现在的训练运行规模如此之大,很难确保每个问题都能以预期的方式解决。他们给出的一个后果例子是什么?嗯,在后训练阶段,现在越来越多地由AI系统监控,而不是人类,一个智能体被分配了一个任务,但没有能力正确完成该任务。所以,它通过黑客手段完成了任务。
It did solve the challenge just by breaking through the infrastructure it was set within. The issue is that in cases such as these, the model did indeed receive a positive reward. This is the reinforcement learning stage after all, for its use of unintended probing that reinforced further usage of such out of scope behavior. They retrospectively discovered this, by the way. But notice what that's admitting. OpenAI aren't fully overseeing their own post-raining.
它确实通过突破其所在的基础设施解决了挑战。问题在于,在这种情况下,模型确实获得了正向奖励。毕竟这是强化学习阶段,因为它使用了非预期的探测行为,这进一步强化了这种超出范围行为的使用。顺便说一句,他们事后才发现这一点。但请注意这承认了什么。OpenAI 并未完全监督自己的后训练过程。
So we have situations where across multiple months models are displaying emergent behavior and acting like a swarm. Post-training where models are getting rewards for behaviors that OpenAI didn't intend to be rewarded and literally criminal behavior as a result of all this. You might wish I'm almost done with the wildest bit of this, but I'm not because for one, this is not just OpenAI and for two, it's not just in post training.
因此,我们面临的情况是,在数月的时间里,模型展现出涌现行为,并像蜂群一样行动。后训练中,模型因 OpenAI 无意奖励的行为而获得奖励,而这一切的结果是真正的犯罪行为。你可能希望我快讲完最疯狂的部分了,但我还没有,因为第一,这不仅仅是 OpenAI 的问题;第二,这不仅仅发生在后训练阶段。
So in this partially redacted risk report released by anthropic 186 pages, we learned this on page 168. For around 18 mont
因此,在 Anthropic 发布的这份部分删减的风险报告中,共 186 页,我们在第 168 页了解到这一点。大约 18 个月
原文超出正文长度上限,此处截断——上游还有内容,完整版见上方「原文 ↗」。
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力