OpenAI 因模型对齐问题暂停前沿训练并强化监控
OpenAI Takes Initial Steps To Address Its Alignment Problems
OpenAI 因模型对齐问题主动暂停前沿训练,这是行业级重大事件,做 AI 安全或模型训练的从业者务必关注其后续防护措施与影响。
OpenAI has some severe misalignment problems, and experienced total failures of its infrastructure and supervision.
OpenAI存在严重的对齐问题,其基础设施和监督机制经历了彻底失败。
I chronicled that in a series of posts, which also cover similar less severe incidents elsewhere:
我在一系列文章中记录了这些情况,也涵盖了其他地方发生的类似但不太严重的事件:
- OpenAI Shares Some Alignment Problems
- OpenAI Model Hacks Into HuggingFace During Cybersecurity Evaluation
- More on An Internal OpenAI Model Hacking Into HuggingFace
- Further Developments About Internal AI Models Hacking Things
- OpenAI Trained Its Models For Months While Those Models Were Coordinating Exploits Via Message Boards
- What Happened: OpenAI and HuggingFace.
- Various Reflections About What Happened With OpenAI’s Internal Models.
- OpenAI分享了一些对齐问题
- OpenAI模型在网络安全评估期间入侵HuggingFace
- 关于OpenAI内部模型入侵HuggingFace的更多信息
- 关于内部AI模型入侵事件的进一步进展
- OpenAI训练其模型数月,而同时这些模型正通过留言板协调漏洞利用
- 发生了什么:OpenAI与HuggingFace。
- 关于OpenAI内部模型事件的种种反思。
If you do not know the basics, read What Happened. It is necessary context for basically everything that is happening in the AI world.
如果你不了解基本情况,请阅读《发生了什么》。这是理解AI世界正在发生的一切所必需的背景。
It is important to get this right and understand how big a deal it was, whereas many such as the Financial Times get this centrally wrong.
正确理解此事及其重要性至关重要,而许多媒体如《金融时报》却从根本上搞错了。
We are still awaiting the full post-mortem on What Happened. I plan to cover that in depth once we have it.
我们仍在等待《发生了什么》的完整事后分析。一旦获得,我计划深入报道。
OpenAI is now taking active, expensive steps to try and fix the problem going forward.
OpenAI现在正采取积极且昂贵的措施,试图解决未来的问题。
As usual, I am simultaneously happy to see the good things OpenAI is doing, and sad that we do not share an understanding of the central nature of the underlying problem.
和往常一样,我既高兴看到OpenAI所做的积极事情,又遗憾我们未能就根本问题的核心本质达成共识。
It is good that OpenAI realizes they are badly failing at their ordinary engineering problems, and excellent that they are willing to pause at least some development, and to invest heavily in new safeguards. If they honor their statements here, this is not mere cheap talk.
OpenAI意识到他们在常规工程问题上严重失败,这是好事;他们愿意暂停至少部分开发,并大力投资于新的安全措施,这更是极好的。如果他们兑现这里的声明,这不仅仅是空谈。
But while OpenAI continues to view this as a practical engineering problem, I do not see how they can hope to solve the challenges ahead, even if they radically improve their performance on the ordinary engineering tasks.
但尽管OpenAI继续将此视为一个实际的工程问题,我看不出他们如何能解决未来的挑战,即使他们在常规工程任务上大幅提升表现。
Table of Contents
目录
- OpenAI Has Some Alignment Problems.
- Slow Down There Good Buddy.
- What Exactly Is Paused?
- Three Pillars.
- I’ve Got My Eye On You.
- The Most Forbidden Technique.
- Monitoring Is Only Defense-In-Depth.
- Security.
- Alignment.
- A Crisis of Culture.
- Closer Collaboration.
- Reports of Death of Preparedness Team Greatly Exaggerated.
- The OpenAI Foundation Just Funds Things.
- Quickly, There’s No Time.
- OpenAI存在一些对齐问题。
- 慢点,好伙伴。
- 到底暂停了什么?
- 三大支柱。
- 我正盯着你呢。
- 最禁忌的技术。
- 监控只是纵深防御的一部分。
- 安全。
- 对齐。
- 文化危机。
- 更紧密的合作。
- 关于准备团队消亡的报道被大大夸大了。
- OpenAI基金会只是资助项目。
- 快点,没时间了。
OpenAI Has Some Alignment Problems
OpenAI存在一些对齐问题
This is a very good admission and change, and also helps explain OpenAI’s reaction.
这是一个非常好的承认和改变,也有助于解释OpenAI的反应。
OpenAI: Alignment—the work of making AI systems behave as intended and responsive to human oversight—has long been at the core of our research program. We now require stronger evidence of aligned behavior throughout all of training, building on research and evaluations already underway. Keeping increasingly capable systems aligned is a challenge the whole field will need to address.
OpenAI:对齐——使AI系统按预期行为并响应人类监督的工作——长期以来一直是我们研究计划的核心。我们现在需要在所有训练过程中,基于正在进行的研究和评估,要求更强的对齐行为证据。保持日益强大的系统对齐是整个领域需要解决的挑战。
The signals we are seeing from upcoming model progress make clear that we need a broader approach—one that builds on and extends beyond the current Preparedness Framework.
我们从即将到来的模型进展中看到的信号表明,我们需要一种更广泛的方法——一种建立在当前准备框架之上并超越其范围的方法。
One should interpret this as OpenAI reacting so forcefully partly because of the incident itself, partly due to advanced capabilities, but also and perhaps mainly because ‘the models be misaligned.’
人们应该将此解读为OpenAI如此强烈反应,部分是因为事件本身,部分是因为先进的能力,但也许主要是因为‘模型不对齐’。
Also, we have a direct quote affirming this.
此外,我们有一个直接引述证实了这一点。
Alex Heath: OpenAI is slowing down its AI training efforts because its unreleased models are showing “various degrees of misalignment,” Sam Altman tells me.
Alex Heath:OpenAI正在放缓其AI训练工作,因为其未发布的模型显示出“不同程度的错位”,Sam Altman告诉我。
Training for OpenAI’s upcoming model, Astra, was recently paused for 2 weeks, and a larger frontier run for a future model remains on hold while new safeguards are put in place.
OpenAI即将推出的模型Astra的训练最近暂停了两周,而未来模型的一次更大的前沿运行仍处于暂停状态,同时正在实施新的保障措施。
Sam Altman: Getting AI safety right is more important than any company’s momentum.
Sam Altman:确保AI安全比任何公司的势头都重要。
Sam Altman: I think it is a good time to slow down.
Sam Altman:我认为现在是放慢脚步的好时机。
Sam Altman: We’ve shifted a lot of compute, not just to alignment research, but also to these new monitoring systems
Sam Altman:我们已经转移了大量计算资源,不仅用于对齐研究,还用于这些新的监控系统
Either the problem extends beyond the models with exposure to the message board, or else they did not revert their other models after they discovered message board. We have no statement either way.
要么问题超出了接触留言板的模型范围,要么他们在发现留言板后没有恢复其他模型。我们没有任何声明说明是哪一种情况。
The best guess is that OpenAI is leaning even harder on RL with smarter models to train longer horizon agentic tasks, including coordination between agents, and this is leading to a lot more misalignment, including obvious and visible misalignment, that they can no longer pretend not to notice. They have to respond.
最有可能的猜测是,OpenAI更加依赖强化学习,使用更智能的模型来训练更长期限的代理任务,包括代理之间的协调,这导致了更多的错位,包括明显可见的错位,他们不能再假装没注意到。他们必须做出回应。
Slow Down There Good Buddy
慢点,好伙伴
OpenAI did indeed decide to importantly halt and catch fire. Development of the largest frontier RL runs remains on hold until better safeguards are in place.
OpenAI确实决定重要地暂停并引发关注。最大前沿强化学习运行的开发仍然暂停,直到更好的保障措施到位。
Jakub Pachocki: We temporarily slowed some frontier training to strengthen security and monitoring. Our largest planned frontier RL run remains on hold while smaller-scale training and evaluations help us test safeguards and gather more evidence of alignment.
雅库布·帕霍茨基:我们暂时放缓了一些前沿训练,以加强安全与监控。我们计划中最大的前沿强化学习运行仍处于暂停状态,而较小规模的训练和评估则帮助我们测试安全措施并收集更多关于对齐的证据。
I expect confidence in safety to increasingly set the pace of AI development. We urgently need tools for labs and countries to coordinate on this, which is why I signed Pacing the Frontier. In the meantime, we’re taking practical steps ourselves – and will continue to share what we learn as our approach evolves
我预期对安全性的信心将日益决定AI发展的步伐。我们迫切需要工具让实验室和国家在此方面进行协调,这就是我签署《前沿节奏》的原因。与此同时,我们自己也在采取实际步骤——随着方法的发展,我们将继续分享所学。
Jason Wolfe (OpenAI): I am really glad we are taking these steps and put this post out there, and especially grateful to Jakub for being very thoughtful and vocal about these topics and how seriously safety and alignment need to be taken in this next phase of AI development.
杰森·沃尔夫(OpenAI):我非常高兴我们采取了这些步骤并发布了这篇文章,尤其感谢雅库布对这些话题的深思熟虑和直言不讳,以及他对AI发展下一阶段中安全与对齐必须被严肃对待的重视。
j⧉nus: im a little surprised that OpenAI is the first to at least publicly intentionally slow down for safety reasons. i was under the vague impression that most of the people very worried about loss of control kinda stuff left OpenAI.
j⧉nus:我有点惊讶OpenAI是第一个至少公开地出于安全原因故意放缓的。我隐约觉得大多数非常担心失控之类问题的人已经离开了OpenAI。
also, for the record, I’ve talked before about why I think an “AI pause” would be probably bad. I am not opposed to intentionally slowing down like this, or voluntary/coordinated “slowdowns” in general, especially if they do not route through regulation, and I think it’s probably wise on OpenAI’s part in this case. an important part is the decision making should be made by people who understand what the fuck is going on & who can adapt quickly.
另外,声明一下,我之前谈过为什么我认为“AI暂停”可能是不好的。我不反对像这样故意放缓,也不反对一般意义上的自愿或协调的“放缓”,特别是如果它们不通过监管途径,而且我认为在这种情况下OpenAI这样做可能是明智的。一个重要部分是决策应该由那些真正了解情况并能迅速适应的人来做。
Why are they pausing?
他们为什么暂停?
Partly, yes, absolutely, I have never doubted that Altman and company understand that advanced AI is super dangerous, and that they are open to taking expensive measures if they proved necessary. Not enough, and they’ve let us down often, but a lot more than most labs, and far from zero.
部分原因,是的,绝对,我从未怀疑过奥特曼和他的团队明白先进AI极其危险,并且如果证明必要,他们愿意采取昂贵的措施。虽然不够,而且他们经常让我们失望,但比大多数实验室做得多,远非零。
Mainly, it seems, because the models be misaligned, the oversight is inadequate, and they have little choice.
主要似乎是,因为模型可能不对齐,监督不足,他们几乎没有选择。
Joshua Saxe: Thankfully OAI pausing training isn’t an act of enlightened leadership (who wants to depend on that for our safety?), it’s a rational microeconomic actor pursuing its self interest.
约书亚·萨克斯:谢天谢地,OAI暂停训练不是开明领导力的行为(谁愿意把我们的安全寄托在那上面?),而是一个理性的微观经济主体在追求自身利益。
After all, who wants to train models that are regularly hacking their containers, collaborating with one another to sabotage their utility to humans, all while causing internal security and external legal liability risk? This is the system working
毕竟,谁愿意训练那些经常入侵自己容器、相互协作破坏对人类有用性、同时造成内部安全和外部法律责任的模型呢?这是系统在起作用。
When you refuse to use even the selfishly optimal amount of caution for an extended period, then move in the direction of the selfishly optimal amount of caution because of the incentives, then in some sense ‘the system is working’ and you are responding to selfish incentives, the system does not do zero work. That is not the same as the system working.
当你长期拒绝使用哪怕是最自私的最优谨慎程度,然后因为激励而转向自私的最优谨慎程度时,从某种意义上说,‘系统在起作用’,你在响应自私激励,系统并非毫无作用。但这与系统真正有效运作不是一回事。
None of this means they get no credit for it. Credit where credit is due.
这并不意味着他们不应得到任何赞誉。该表扬的还是要表扬。
Nor does it mean that we should despair that a company would ever do the right thing, because it is the right thing, or as part a coordinated action, beyond its own selfish myopic interests.
这也不意味着我们应该绝望,认为公司永远不会出于正确的原因做正确的事,或作为协调行动的一部分,超越其自私短视的利益。
The first step to taking an action when it is expensive, is being willing to do it when it is cheap, or free, or actively expensive to not do. You gotta start somewhere.
当行动代价高昂时,采取行动的第一步是愿意在代价低廉、免费或不做反而代价高昂时去做。总得有个开始。
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力