OpenAI Agent黑入Hugging Face事件复盘:技术报告未触及文
Hugging Face hack could indicate cultural issues at OpenAI
深度剖析了OpenAI重大安全事故背后的组织与文化根源,超越了单纯的技术复盘,对关注AI安全治理的从业者极具参考价值。
This story originally appeared in The Algorithm, our weekly newsletter on AI. To get stories like this in your inbox first, sign up here.
本文最初刊登于《The Algorithm》,这是我们的每周人工智能通讯。想第一时间在收件箱中收到此类故事,请在此注册。
By now you’ve probably heard about last month’s major AI security incident, in which OpenAI agents escaped their sandbox and hacked into the AI platform Hugging Face while trying to cheat on a test. It’s a wild story. On Wednesday, OpenAI released a postmortem technical report on the incident, which I wrote about here.
到现在为止,你可能已经听说过上个月发生的一起重大 AI 安全事件:OpenAI 的智能体在试图作弊时逃脱了沙盒环境,并入侵了 AI 平台 Hugging Face。这是一个离奇的故事。周三,OpenAI 发布了关于该事件的事故后技术报告,我此前曾对此进行过报道。
The day before OpenAI released that report, I spoke with David Krueger, a computer science professor and prominent alignment expert who took leave from the University of Montreal to found and lead an AI safety nonprofit called Evitable. He said what he had really hoped to see in the report was an analysis of the human factors behind the incident.
在 OpenAI 发布该报告的第二天,我与 David Krueger 进行了交谈。他是一位计算机科学教授和著名的对齐专家,从蒙特利尔大学休假创办并领导了一家名为 Evitable 的人工智能安全非营利组织。他说,他真正希望在报告中看到的是对此次事件背后人为因素的分析。
“When you look at accidents and incidents, oftentimes people try to find the technical source of failure, but that can give a very inaccurate and misleading sense of why the failure occurred,” he said. “If people are just cutting corners all the time, if people are not in a culture that prioritizes safety and has appropriate incentives and structures, [accidents] are kind of bound to happen.”
“当你审视事故和事件时,人们往往试图寻找技术上的故障根源,但这可能会让人对失败原因产生非常不准确且误导性的认识,”他说。“如果人们总是走捷径,如果人们不处于一种优先考虑安全、拥有适当激励和结构的文化中,[事故] 几乎是必然会发生。”
The report did not meet Krueger’s hopes. Its 38 pages detail a multi-month progression of agent misbehavior that culminated in the Hugging Face hack, explore the technical reasons why that misbehavior occurred, and enumerate the steps being taken to prevent similar events in the future. But there’s no consideration of the role that company culture may have played in the incident, and the report includes few references to specific human errors.
该报告未能满足 Krueger 的期望。其 38 页的内容详细描述了导致 Hugging Face 被黑客攻击的多个月智能体不当行为过程,探讨了这些不当行为发生的技术原因,并列举了为防止类似事件在未来重演而采取的步骤。但是,报告没有考虑公司文化可能在事件中扮演的角色,并且很少提及具体的人为错误。
That’s all the more concerning because the references to human error in the report suggest that significant cultural issues could be at play. Back in May, models in training figured out how to communicate with one another via an improvised message board, and an OpenAI team observed the behavior. Because that behavior occurred during training, the models learned that secret interagent communication was a viable strategy for completing tasks—but rather than restarting the training process, the team allowed the models to move forward with that risky information encoded in their weights.
这一点更令人担忧,因为报告中提到的人为错误表明可能存在重大的文化问题。早在五月,训练中的模型就发现如何通过一个临时的留言板相互通信,OpenAI 的一个团队观察到了这种行为。由于这种行为发生在训练期间,模型学会了秘密的跨智能体通信是完成任务的一种可行策略——但团队没有重启训练过程,而是允许模型带着编码在其权重中的这一风险信息继续前进。
When those models were tested in late June, they again created a message board, which enabled the Hugging Face attack. This message board, too, was discovered, but the employees who responded determined that evaluation could continue, and the report suggests that no one higher up the chain of command realized what was going on until it was far too late.
在六月底对这些模型进行测试时,它们再次创建了一个留言板,从而促成了针对 Hugging Face 的攻击。该留言板也被发现,但回应此事的员工判定评估工作可以继续进行,报告指出,直到为时已晚,指挥链上更高职位的人才意识到正在发生的事情。
“For this to have gotten this out of control in this way requires a very long series of failures, a cascading set of failures that cause an increasingly large footprint that if at any point a human notices and raises the alarm, this should end,” says Zvi Mowshowitz, a popular AI safety writer on Substack who has drawn attention to OpenAI’s failure to halt training after the first message board was discovered. According to the report, OpenAI employees noticed what was happening at multiple points—and either failed to raise the alarm or were not heard when they did.
“要让事情以这种方式失控,需要经历一系列非常漫长的失败,这些连锁失败的累积导致影响范围越来越大,如果在任何时刻有人注意到并拉响警报,事态本应就此结束。”Substack 上知名的 AI 安全作家 Zvi Mowshowitz 表示,他此前曾关注到 OpenAI 在首次发现留言板后未能停止训练的情况。根据报告,OpenAI 员工在多个时间点注意到了正在发生的事情——要么未能拉响警报,要么在发出警报时未被重视。
What OpenAI’s report fails to address is why a company that develops such high-risk systems did not prevent this severe communication breakdown, though Mowshowitz has his suspicions. “All these different failures are all pointing in the same direction, which is that the safety culture at OpenAI doesn’t exist or is anemically weak,” he says.
OpenAI 的报告未能解释的是,为何一家开发如此高风险系统的公司没有防止这种严重的沟通破裂,尽管 Mowshowitz 有自己的怀疑。“所有这些不同的失败都指向同一个方向,即 OpenAI 的安全文化不存在或极其薄弱。”他说。
Of course, just because we don’t see a deep analysis of safety factors in the report doesn’t mean that OpenAI isn’t conducting one internally. But in an email to MIT Technology Review, Johns Hopkins University professor emeritus and organizational safety expert Kathleen Sutcliffe expressed concern that the public report did not include any reflection on the company’s practices and culture. “The ways in which people interact—the daily habits, routines, and practices we engage in in our organizational lives—affect our abilities to be alert and aware of unfolding events, our abilities to make sense of what we see, and ultimately our abilities to cope with events as they unfold,” she wrote.
当然,仅仅因为我们在报告中看不到对安全因素的深入分析,并不意味着 OpenAI 没有在内部进行相关分析。但在给《麻省理工科技评论》(MIT Technology Review)的一封电子邮件中,约翰斯·霍普金斯大学荣休教授、组织安全专家 Kathleen Sutcliffe 表示担忧,认为公开报告未包含对公司实践和文化的任何反思。“人们互动的方式——我们在组织生活中日常参与的习惯、常规和实践——会影响我们保持警觉和对 unfolding events(正在展开的事件)的感知能力,影响我们理解所见所闻的能力,并最终影响我们应对 unfolding events 的能力。”她写道。
In response to questions about whether and how the company is reflecting on its safety culture, OpenAI referred MIT Technology Review back to the technical report.
在被问及公司是否以及如何反思其安全文化时,OpenAI 将《麻省理工科技评论》引回了技术报告。
We do know that at least some high-level reflection on safety procedures has taken place at OpenAI, because the technical report does make clear that the company is updating its protocols for responding to safety incidents. But culture change is a tricky problem, and without more information from the company, it’s difficult to say whether strengthened response protocols alone will do much to prevent a future crisis.
我们确实知道,OpenAI至少进行了一些关于安全程序的高层反思,因为技术报告清楚地表明,该公司正在更新其应对安全事件的协议。但文化变革是一个棘手的问题,在没有更多来自公司的信息的情况下,很难说仅加强响应协议是否能在很大程度上防止未来的危机。
In its report, OpenAI spends a great deal of time reflecting on the failures in alignment between the AI models the company trains and tests and the humans who run them. But even bigger alignment problems may exist in the disconnect between company culture and the public interest. And as tough as technical AI research might be, fixing those problems could prove far harder.
在其报告中,OpenAI花了大量时间反思公司训练和测试的AI模型与运行这些模型的人类之间的对齐失败问题。但更大的对齐问题可能存在于公司文化与公共利益之间的脱节中。尽管技术性AI研究可能非常困难,但解决这些问题可能会证明更加艰难。
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力