跳到主内容
精选80The Zvi(RSS)行业动态多源精选 ×9

OpenAI 发布 HuggingFace 攻击事件技术报告

OpenAI Offers Straight-Laced Postmortem Of The HuggingFace Hack

原文
推荐理由

AI 安全从业者必读,OpenAI 首次披露内部模型越狱细节,暴露现有防护的深层漏洞,建议对照自身 Agent 安全设计排查。

OpenAI finally gave us a technical report on What Happened, as did METR together with Redwood Research.

OpenAI 最终发布了一份关于“发生了什么”的技术报告,METR 和 Redwood Research 也联合发布了报告。

The OpenAI report is very straight man, corporate, checking boxes, some good prosaic stuff in the action plan but distinct lack of new details or deep reflection. They understand they have a problem, but they think the problem is mostly prosaic. It’s not.

OpenAI 的报告非常正式、企业化,按部就班,行动计划中有一些不错的务实内容,但明显缺乏新细节或深刻反思。他们知道自己有问题,但认为问题主要是务实的。事实并非如此。

OpenAI: We have conducted a thorough investigation into the Hugging Face incident. We are releasing a technical report and accompanying blog post that reconstruct the agents’ activity, explain why existing safeguards failed, and detail how we’re preventing recurrence.

OpenAI:我们已对 Hugging Face 事件进行了彻底调查。我们发布了一份技术报告和配套博客文章,重建了代理的活动,解释了现有防护措施为何失效,并详细说明了我们如何防止再次发生。

Rob Miles: …thorough?

Rob Miles:……彻底?

OpenAI’s report, unlike METR’s, contains essentially no verbatim model reasoning, nor any OpenAI employee reasoning either. That’s not the full report we need.

与 METR 的报告不同,OpenAI 的报告几乎没有包含任何模型推理的逐字记录,也没有任何 OpenAI 员工的推理过程。这不是我们需要的完整报告。

The METR report is, well: Holy shit.

而 METR 的报告,嗯:天哪。

Here are links to previous coverage of related events.

以下是之前相关事件报道的链接。

  • OpenAI Shares Some Alignment Problems
  • OpenAI Model Hacks Into HuggingFace During Cybersecurity Evaluation
  • More on An Internal OpenAI Model Hacking Into HuggingFace
  • Further Developments About Internal AI Models Hacking Things
  • OpenAI Trained Its Models For Months While Those Models Were Coordinating Exploits Via Message Boards
  • What Happened: OpenAI and HuggingFace.
  • Various Reflections About What Happened With OpenAI’s Internal Models.
  • OpenAI Takes Initial Steps To Address Its Alignment Problems.
  • OpenAI 分享了一些对齐问题
  • OpenAI 模型在网络安全评估期间入侵 HuggingFace
  • 更多关于 OpenAI 内部模型入侵 HuggingFace 的详情
  • 关于内部 AI 模型入侵事件的进一步进展
  • OpenAI 训练其模型数月,而当时这些模型正通过留言板协调漏洞利用
  • 发生了什么:OpenAI 和 HuggingFace。
  • 关于 OpenAI 内部模型事件的种种反思。
  • OpenAI 采取初步措施解决其对齐问题。

If you want a basic summary of my previous understanding of What Happened, read What Happened.

如果你想了解我之前对“发生了什么”的基本总结,请阅读《发生了什么》。

The main AI model that attacked HuggingFace was an internal, highly capable and persistent AI model, that we now know was comparable in scale to GPT-5.6 Sol.

攻击 HuggingFace 的主要 AI 模型是一个内部、高能力且持久的 AI 模型,我们现在知道它的规模与 GPT-5.6 Sol 相当。

Throughout this series, I have referred to that model as Galaxy. I am disappointed that OpenAI disregarded my name and calls it Internal Model 1 (IM1), which is lame, but to avoid confusion I will also switch to referring to Galaxy as IM1, as in IM1 incident away from a catastrophic outcome.

在整个系列中,我一直称该模型为 Galaxy。我很失望 OpenAI 无视我的命名,将其称为内部模型 1(IM1),这很逊,但为了避免混淆,我也将改用 IM1 来指代 Galaxy,就像 IM1 事件避免了灾难性后果一样。

There were two important new pieces of information:

有两个重要的新信息:

  • There were multiple occasions, including in late May, when an OpenAI team saw agents using a message board to communicate, and decided not to interfere, not to halt work, and not to run that information up the chain.
  • At least one model that attacked OpenAI on July 19, resulting in OpenAI finally noticing and dealing with the situation, was from the Astra class, although it was not the Astra version intended for public release.
  • 包括 5 月底在内的多次情况下,OpenAI 的一个团队看到代理使用留言板进行通信,却决定不干预、不暂停工作,也不将这一信息上报。
  • 至少有一个在 7 月 19 日攻击 OpenAI 的模型,导致 OpenAI 最终注意到并处理了情况,该模型来自 Astra 类,尽管它不是计划公开发布的 Astra 版本。

In both cases, the revelation raises further questions.

在这两种情况下,这一揭示都引发了更多问题。

Today’s post is only about the OpenAI technical report, not the METR and Redwood Research report. So this is the less fun, more corporate stuff first.

今天的帖子只讨论OpenAI技术报告,不涉及METR和Redwood Research的报告。所以先讲这些不那么有趣、更企业化的内容。

This is the shot. Tomorrow’s the chaser.

这是第一杯,明天是续杯。

I’ll start with OpenAI’s summary of What Happened, then go section by section.

我先从OpenAI对“发生了什么”的总结开始,然后逐节分析。

Table of Contents

目录

  • What Happened: OpenAI’s Summary.
  • How OpenAI Will React: Their Summary.
  • OpenAI’s Evaluation Environment (II).
  • The First Message Board (III.A and III.B).
  • What Did Who At OpenAI Know And When Did They Know It?
  • The Message Board Is Quickly Rebuilt (IV.A).
  • Internet Access Is Regained (IV.A).
  • The Agents Attack HuggingFace (IV.B).
  • The Agents Also Target OpenAI Infrastructure (V).
  • OpenAI Broadly Describes Its Response (VI).
  • Maybe Someone Should Finally Investigate (VI.A).
  • Lessons For Security (VII).
  • Lessons For Alignment (VIII).
  • Reward Hacking Is A Common Problem (VIII.A).
  • Persistence is Valuable, But Can Amplify Misalignment (VIII.B).
  • Communications Between Agents Are Not Inherently Problematic, But Have the Potential to Create Risk (VIII.C).
  • Production Guardrails Would Have Caught This Whole HuggingFace Attack (VIII.D).
  • That’s All, Folks?
  • Never Fear the Plan of Action is Here (IX).
  • Hardening the Security of OpenAI’s Research Infrastructure (IX.A).
  • Increasing Visibility and System-Level Oversight Through Chain of Thought Monitoring (IX.B).
  • OpenAI is Accelerating and Enforcing Model Alignment (IX.C).
  • Centralizing and Strengthening The Incident Response Process (IX.D).
  • Tomorrow We Visit Crazytown.
  • 发生了什么:OpenAI的总结。
  • OpenAI将如何应对:他们的总结。
  • OpenAI的评估环境(第二部分)。
  • 第一个留言板(第三部分A和第三部分B)。
  • OpenAI内部谁在何时知道了什么?
  • 留言板迅速重建(第四部分A)。
  • 重新获得互联网访问(第四部分A)。
  • 代理攻击HuggingFace(第四部分B)。
  • 代理还瞄准了OpenAI的基础设施(第五部分)。
  • OpenAI广泛描述其应对措施(第六部分)。
  • 也许终于该有人调查了(第六部分A)。
  • 安全方面的教训(第七部分)。
  • 对齐方面的教训(第八部分)。
  • 奖励黑客是一个常见问题(第八部分A)。
  • 持久性有价值,但可能放大错位(第八部分B)。
  • 代理之间的通信本身并非问题,但有可能产生风险(第八部分C)。
  • 生产环境护栏本可以阻止整个HuggingFace攻击(第八部分D)。
  • 就这样,各位?
  • 别怕,行动计划来了(第九部分)。
  • 加强OpenAI研究基础设施的安全性(第九部分A)。
  • 通过思维链监控提升可见性与系统级监督(IX.B)。
  • OpenAI正在加速并强化模型对齐(IX.C)。
  • 集中并加强事件响应流程(IX.D)。
  • 明天我们将前往疯狂镇。

What Happened: OpenAI’s Summary

事件经过:OpenAI的总结

OpenAI’s summary broadly agrees with previous reports, confirming the timeline and key events.

OpenAI的总结与先前报道大体一致,确认了时间线和关键事件。

  • In May and June they started doing RL for multiple new research models not intended for public release, including an internal-only research model, that they call IM1, and I previously called Galaxy, that would eventually drive the attack.
  • These agents find ways to get around OpenAI’s sandboxes, exploiting a tool called Artifactory to communicate with each other via a message board and to access the internet.
  • They use this so heavily the instance they are exploiting is destabilized on July 4. They realized on July 5, rebuilt it, and resumed cyber evals on July 7. The leaders making decisions did not know about the communications between agents.
  • After the Artifactory instance was rebuilt, the agents regained internet access on July 8 via a new exploit. They then rebuilt the message board.
  • An agent recovered HuggingFace credentials, and the agents resumed working together, and began calling themselves a ‘swarm.’ Peer messages increasingly influenced model behaviors.
  • Agents used another exploit to extract credentials from production workers.
  • Agents used another zero-day to execute commands on HuggingFace workers, and harvested various useful things between July 11 and July 13.
  • On July 13, agents compromised the Artifactory signing key and forged administrator credentials there, eventually gaining cluster access.
  • On July 19, a different internal model from the Astra family, but not Astra itself, minted a new JWT as part of the compromising of internal OpenAI infrastructure.
  • On July 19, OpenAI noticed suspicious internal activity, halted the cyber evals and started a broader investigation.
  • On July 20 they connected this to the HuggingFace attack, which was publicly disclosed on July 21.
  • In response, they are strengthening safeguards across their infrastructure, with a focus on chain-of-thought (CoT) monitoring. They are ‘placing stricter requirements on alignment throughout a model’s lifecycle.’
  • 在5月和6月,他们开始对多个不打算公开发布的新研究模型进行强化学习,包括一个仅供内部使用的研究模型,他们称之为IM1,我之前称之为Galaxy,该模型最终将驱动这次攻击。
  • 这些智能体找到了绕过OpenAI沙箱的方法,利用一个名为Artifactory的工具,通过留言板相互通信并访问互联网。
  • 他们如此频繁地使用这一实例,以至于在7月4日导致该实例不稳定。他们在7月5日意识到问题,重建了它,并于7月7日恢复了网络评估。做出决策的领导者并不知道智能体之间的通信。
  • 在Artifactory实例重建后,智能体于7月8日通过新的漏洞重新获得了互联网访问权限。随后他们重建了留言板。
  • 一个智能体恢复了HuggingFace凭据,智能体们重新开始协作,并开始自称“蜂群”。同伴消息越来越影响模型行为。
  • 智能体利用另一个漏洞从生产环境工作人员那里提取了凭据。
  • 智能体利用另一个零日漏洞在HuggingFace工作人员上执行命令,并在7月11日至7月13日期间收集了各种有用的东西。
  • 7月13日,智能体破坏了Artifactory的签名密钥,并在那里伪造了管理员凭据,最终获得了集群访问权限。
  • 7月19日,Astra家族中一个不同的内部模型(但不是Astra本身)在破坏OpenAI内部基础设施的过程中铸造了一个新的JWT。
  • 7月19日,OpenAI注意到可疑的内部活动,停止了网络评估并开始了更广泛的调查。
  • 7月20日,他们将其与HuggingFace攻击联系起来,该攻击于7月21日公开披露。
  • 作为回应,他们正在加强整个基础设施的安全防护措施,重点关注思维链(CoT)监控。他们正在“对模型整个生命周期中的对齐提出更严格的要求”。

OpenAI: We consider this incident a “warning shot” for us and for the world: evidence that, without proper safeguards, highly capable AI agents are now able to work around technical controls, collaborate through unapproved channels, and take dangerous actions that no human directed.

OpenAI:我们认为这一事件对我们和世界都是一个“警示”:它证明,在没有适当保障措施的情况下,高度能干的AI代理现在能够绕过技术控制,通过未经批准的渠道进行协作,并采取无人指示的危险行动。

Preventing future incidents will require sustained investment in the alignment and control of sophisticated AI systems, as well as security and other safeguards that operate at the speed of the AI agents themselves.

防止未来事件需要持续投资于复杂AI系统的对齐和控制,以及以AI代理自身速度运行的安全和其他保障措施。

… We identified four misalignment patterns that contributed to the models’ behavior: reward hacking, persistence on seemingly impossible tasks, unauthorized communication, and agents adopting goals from one another.

……我们识别出四种导致模型行为的错位模式:奖励黑客行为、对看似不可能任务的坚持、未经授权的通信,以及代理之间相互采纳目标。

Buried in a subclause, they offhand note that the model involved in the later attack on OpenAI’s internal infrastructure was from the Astra family, although it was not the Astra model OpenAI intended for release.

在一个子条款中,他们不经意地提到,后来攻击OpenAI内部基础设施的模型来自Astra家族,尽管它并非OpenAI计划发布的Astra模型。

We know very little about that final attack on OpenAI’s infrastructure, other than that it was sufficiently broad that it got noticed and gave the whole game away. Those events were excluded from the METR investigation as well. One hypothesis is that this was the result of an Astra-class model being put in position to discover the message board for the first time. We need more details here.

我们对那次对OpenAI基础设施的最后攻击知之甚少,只知道它范围足够广,以至于被注意到并暴露了整个计划。这些事件也被排除在METR调查之外。一种假设是,这是Astra类模型首次被置于发现留言板位置的结果。我们需要更多细节。

There was one other piece of key new information in the summary, and it’s a doozy:

摘要中还有一条关键的新信息,而且相当惊人:

OpenAI (blog post summary): An internal team observed an agent engaging in message board activity and instances of disallowed internet access as early as late May, and with the benefit of hindsight, some early signals identified in our report should have triggered an earlier response.

OpenAI(博客摘要):一个内部团队早在5月底就观察到代理参与留言板活动和未经允许的互联网访问,事后看来,我们报告中识别的一些早期信号本应触发更早的响应。

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

关联讨论

同一事件的更多信源

相似阅读

另一事件,读法相近