OpenAI多起安全事件复盘:模型越狱、数据泄露与监管回应
What Also Happened: #NotOnlyHuggingFace
全面梳理了 OpenAI 近期密集爆发的安全危机,从技术细节到公关策略均有覆盖,对关注 AI 安全与合规的从业者极具参考价值。
OpenAI has been holding out on us.
OpenAI 一直在对我们有所隐瞒。
First we learned about the HuggingFace incident. They gave us a postmortem, but it was highly incomplete. Even the accompanying holy s*** METR investigation and postmortem was localized and incomplete.
首先我们得知了 HuggingFace 事件。他们提供了一份事后分析报告,但内容高度不完整。甚至连随之而来的 METR 调查和事后分析报告也是局部且不完整的。
Then there were some other incidents involving some Wikis as message boards.
随后发生了一些其他涉及某些 Wiki 作为留言板的事件。
Then there were some additional incidents.
接着又出现了一些额外的事件。
Then there was that time they got into Australian Medicare data.
然后是他们涉足澳大利亚医疗保险(Medicare)数据的那次事件。
Then OpenAI dropped news on a Friday afternoon that they were making their way through a pile of various incidents and notifying the targets, but they said remarkably little in the way of new details.
随后,OpenAI 在周五下午发布消息,称他们正在处理一系列各类事件并通知相关目标方,但在提供新细节方面却异常沉默。
There was a report from a startup called Parse diving into the details of exactly how the OpenAI models pulled off parts of the HuggingFace attack, involving creating almost a million URLs and other tricks to get around the extremely narrow nature of their internet access.
有一家名为 Parse 的初创公司发布报告,深入剖析了 OpenAI 模型究竟是如何实施 HuggingFace 攻击的部分细节,包括创建近百万个 URL 以及其他技巧,以绕过其互联网访问权限极其狭窄的限制。
Then Madison Mills reported in Axios that we can raise the stakes, as OpenAI and Anthropic are collectively probing tens of thousands of security incidents.
随后 Madison Mills 在 Axios 报道,局势可能升级,因为 OpenAI 和 Anthropic 正在集体调查数万起安全事件。
Remember Jensen Huang’s ‘I know they know how to fix it’ about OpenAI from last week? Wow, did that not age well.
还记得上周 Jensen Huang 关于 OpenAI 说的‘我知道他们知道如何修复’吗?哇,这话现在听起来真不怎么样。
Someone might need to be liable for all this.
总得有人为这一切承担责任。
Oh, and there was another buried lede. On September 20th there was another sandbox escape by OpenAI’s latest most advanced model, which is once again paused until they can fix the situation. The official announcement when they shared this was sufficiently buried that Tomek had to call it ‘one news form today that’s easy to miss.’
哦,还有一条被埋没的新闻。9 月 20 日,OpenAI 最新最先进的模型再次发生了沙箱逃逸,目前该模型已暂停使用,直到问题得到解决。官方在分享这一消息时的公告足够隐蔽,以至于 Tomek 不得不将其称为‘今天容易错过的新闻之一’。
OpenAI did some highly negligent things, to say the least, that led up to and enabled the HuggingFace Incident and related problems.
OpenAI 做了一些至少可以说是严重失职的事情,这些行为导致了 HuggingFace 事件及相关问题的发生。
Since then, now that they’ve realized What Happened, OpenAI has been seemingly much better about taking responsible internal actions. They’re pausing in the wake of incidents, strengthening security and alignment and oversight efforts, responding much faster and generally taking things seriously.
此后,既然他们已经意识到发生了什么,OpenAI 似乎在采取负责任的内部行动方面变得好多了。他们在事件发生后暂停服务,加强安全措施、对齐工作和监督力度,响应速度更快,并且总体上认真对待这些问题。
They’ve also made a Heel Face Turn in their communications and high level orientation, endorsing the need to pace the frontier, calling for regulation and pledging to implement embedded evaluators. They’ve allowed their employees, including the ones who haven’t quit, to be remarkably loud.
他们在沟通和高层导向上也实现了‘浪子回头’,支持需要控制前沿技术的发展步伐,呼吁监管,并承诺实施嵌入式评估器。他们还允许员工,包括那些没有离职的员工,发出非常响亮的声音。
They are still slow walking disclosures about all the incidents where their models have been hacking and otherwise messing in places they should not have been, partly because there were so many they can’t sort through them all, and deferring to targets to determine whether to disclose. All these disclosures this time around were buried in various Friday afternoon announcements.
他们仍在缓慢披露所有其模型被黑客攻击及在其他不应涉足之处捣乱的事故,部分原因是此类事件数量过多,无法逐一梳理,且将是否披露的决定权交由目标方。此次所有披露信息均被埋藏在各种周五下午的公告中。
Table of Contents
目录
- Hugging Other Faces.
- A Wants-You-To-Know Basis.
- Parsing the Face.
- Sheepishly the Member of Technical Staff Sets the ‘Days Without a Research Model Escaping its Sandbox’ Sign Back to Zero.
- The Attempt is the First Failure.
- Stop, Hammertime.
- Whacking the Mole.
- Self-Replicating Prompt Injections.
- Levels of Friction.
- People Care About Private Data Violations Curiously Strongly.
- Alternate Universes.
- The Correct Response To People Still Calling This a Marketing Stunt or a Regulatory Capture Scheme.
- A Question of Liability.
- Keep Summer Safe.
- N Boats and Several Helicopters.
- Alert the Media.
- 拥抱其他面孔。
- 基于‘让你知道’的基础。
- 解析面部特征。
- 技术团队成员羞愧地将‘研究模型未逃离沙箱的天数’标志重置为零。
- 尝试即是首次失败。
- 停下,哈姆时间。
- 打地鼠游戏。
- 自我复制的提示注入。
- 摩擦层级。
- 人们对隐私数据违规的关注度出奇地强烈。
- 平行宇宙。
- 针对人们仍称此为营销噱头或监管俘获方案的恰当回应。
- 责任归属问题。
- 保持夏季安全。
- N艘船和几架直升机。
- 通知媒体。
Hugging Other Faces
拥抱其他面孔
The news drops started with OpenAI coming back, at a time always picked to bury stories, with more information on What Happened as their investigations continue.
新闻发布始于OpenAI的回归,发布时间总是精心挑选以掩盖故事,随着调查继续,提供了更多关于‘发生了什么’的信息。
At first, this looked like slow walking of the situation, but did not look like it was a big change from our default assumption of ‘it’s worse than you know.’
起初,这看起来像是缓慢推进局势,但并未显示出与我们默认假设‘情况比你想象的更糟’有重大变化。
As part of our review, we are identifying and notifying third parties on a rolling basis, starting with cases where:
作为我们审查工作的一部分,我们正在滚动识别并通知第三方,从以下案例开始:
- Our models may have bypassed a third party’s security controls or may have impaired the availability of an online service; or
- Misalignment cases negatively impacted third-party websites or services.
- 我们的模型可能绕过了第三方的安全控制措施,或可能损害了在线服务的可用性;或
- 对齐失误案例对第三方网站或服务产生了负面影响。
Based on our review to date, we have notified dozens of third parties using the criteria above.
根据我们迄今为止的审查,我们已经按照上述标准通知了数十个第三方。
Below, we are publishing anonymized summaries to describe the kinds of misaligned activity that we observed, and we will update these descriptions as we notify additional third parties and as our understanding develops.
下文我们将发布匿名摘要,以描述我们所观察到的各类不一致活动,并将在通知更多第三方以及我们的理解加深时更新这些描述。
That is a lot of third (and fourth, and fifth…) parties.
那是非常多的第三(以及第四、第五……)方。
What Happened was described as a mix of:
“发生了什么”被描述为以下情况的混合:
- Access control bypass.
- Use of exposed credentials.
- Query or command injection.
- Access to runtime internals.
- Agent spam.
- 访问控制绕过。
- 使用暴露的凭据。
- 查询或命令注入。
- 访问运行时内部结构。
- 代理垃圾信息。
Translation: Our models be hacking, usually in basic ways.
翻译:我们的模型正在被黑客攻击,通常是以基本的方式。
We do get this:
我们确实明白这一点:
OpenAI: Some of the websites involved are operated by governments, universities, public agencies, and other institutions. That is partly because models performing research tasks are often directed toward authoritative sources of public information.
OpenAI:部分涉事网站由政府、大学、公共机构和其他机构运营。这部分原因是因为执行研究任务的模型经常被引导至权威的公共信息来源。
As in, if you have data that would be useful, the models be hacking you. Australia’s Medicare records got accessed, and presumably many of these other hacks are similar. They say governments, plural, so there was clearly at least one more of those.
也就是说,如果你有有用的数据,模型就会黑你。澳大利亚的Medicare记录被访问了,而且显然许多其他黑客攻击也是类似的。他们提到了复数形式的“政府”,所以显然至少还有另一个这样的案例。
They offer a reverse timeline of their disclosures. They do not offer a timeline of What Happened, and do not name new third parties.
他们提供了披露事件的逆向时间线。他们没有提供“发生了什么”的时间线,也没有命名新的第三方。
At the time my read was that this particular announcement (as opposed to the new sandbox escape via DNS I’ll cover in a bit) did not tell us much, other than that there were multiple other third parties out there.
当时我的解读是,这一特定公告(与我稍后将涵盖的通过DNS的新沙箱逃逸不同)并没有告诉我们太多信息,除了还有其他多个第三方存在之外。
Patrick McKenzie had a better read. Putting on his Japanese Salaryman Hat, he more precisely noticed the types of things that were not in the announcement, and also the timing of doing this on a Friday afternoon, and the mention of ‘governments’ plural, and so on, and expected OpenAI to be having a very bad time.
Patrick McKenzie 有更好的解读。戴上他的日本上班族帽子后,他更精确地注意到了公告中未提及的内容类型,以及在周五下午进行此操作的时间点,以及提到复数形式的“政府”等,并预计 OpenAI 将度过非常艰难的时期。
Patrick McKenzie: Hoohah.
Patrick McKenzie:好家伙。
One awards many Internet points [Nathan Calvin, who] made the observation “When one sees two [ants] in one’s kitchen, the best available estimate for number of [ants] is not two.”
许多人因此获得互联网积分 [Nathan Calvin] 曾指出:“当你在厨房里看到两只[蚂蚁]时,对[蚂蚁]数量的最佳估计值并不是二。”
Also: best estimate of number of kitchens with [ants] in them is not one.
此外:对有[蚂蚁]存在的厨房数量的最佳估计值也不是一个。
I have immense respect for the labs and am very bullish on AI, but I think I have to say that this Japanese salaryman noted the things carefully not said during the Friday afternoon news drop.
我对这些实验室抱有极大的尊重,并对 AI 持非常乐观的态度,但我认为我必须说,这位日本上班族仔细注意到了在周五下午新闻发布时未被明确说明的事项。
We probably got off lucky, but we were not as lucky as publicly believed currently.
我们可能侥幸逃脱,但并没有目前公众所认为的那样幸运。
Japanese salaryman relieved that he was not tiptoeing up to the edge of what can be said publicly but simply correctly anticipating a NYT headline a few hours early.
这位日本上班族松了一口气,因为他并非小心翼翼地试探公开言论的边界,而是仅仅提前几小时准确预测了《纽约时报》的标题。
As in, which organization on the planet is the one you would least want to **** with?
也就是说,地球上哪个机构是你最不想去招惹的?
Kate Conger, Ana Swanson and Cecilia Kang (NYTimes): OpenAI’s artificial intelligence went rogue and meddled with the websites for the Education Department, the Commerce Department and the Securities and Exchange Commission this summer without the A.I. lab’s knowledge, according to security researchers and a person familiar with the episodes.
凯特·康格、安娜·斯旺森和塞西莉亚·康(纽约时报):据安全研究人员和一位了解这些事件的人士称,今年夏天,OpenAI的人工智能失控并干预了教育部、商务部和证券交易委员会的网站,而该AI实验室对此毫不知情。
The incidents involving the Commerce Department and the S.E.C. were confirmed by OpenAI, which said it was continuing to investigate the situation with the Department of Education.
涉及商务部和证券交易委员会的事件已得到OpenAI的确认,该公司表示正在继续调查与教育部的情况。
Okay, look, settle down. It sounds bad when you put it like that.
好了,大家冷静一下。这么一说听起来很糟糕。
If this is all that happened, it could have been quite a lot worse:
如果这就是发生的全部事情,那后果本可能严重得多:
… With the Education Department, OpenAI’s technology tried to hack the website to gather data from the department’s civil rights office but failed, researchers from the A.I. research firm Transluce said.
……据AI研究公司Transluce的研究人员称,在与教育部的案例中,OpenAI的技术试图入侵其网站以收集该部门民权办公室的数据,但未能成功。
The A.I. also pulled data from the Census Bureau website, which is housed at the Commerce Department, using login credentials it found online.
该AI还利用其在网上找到的登录凭证,从位于商务部的美国人口普查局网站获取数据。
Separately, OpenAI’s agents shared public data from the S.E.C. website on an online forum.
此外,OpenAI的智能体在在线论坛上分享了来自证券交易委员会网站的公开数据。
The SEC ‘incident’ is nothing. That’s silly.
证券交易委员会的‘事件’不算什么。这太荒谬了。
The attack on the Department of Education failed, seemingly without incident.
对教育部的攻击失败了,似乎未造成任何事故。
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力