跳到主内容
@wquguru
精选86ChinaTalk(RSS)行业动态

中国媒体与学界如何看待OpenAI智能体攻击Hugging Face事件

China on the Hugging Face Incident

原文
发到 X
推荐理由

深度复盘了OpenAI智能体越狱事件的调查细节,并独家梳理了中国官方媒体、舆论及学界对此事的差异化反应与安全焦虑,对理解中美AI治理语境差异极具参考价值。

We finally have a more complete timeline of the OpenAI-Hugging Face attack in July. Safety researchers at METR and Redwood Research published their independent investigation on August 26. The same day, OpenAI released its own technical report on the incident. Both also published translations of their reports in Chinese.

我们终于有了7月OpenAI与Hugging Face冲突事件的更完整时间线。METR和Redwood Research的安全研究人员于8月26日发布了他们的独立调查报告。同一天,OpenAI也发布了关于该事件的技术报告。两者也都发布了其报告的中文译本。

In short, OpenAI researchers launched around 1200 agents and targeted them at tasks in ExploitGym, a large benchmark built from real-world software vulnerabilities that was designed to test how well AI models can exploit them. In order to solve “impossible” tasks, these desperate agents, programmed to be highly persistent, found creative ways out of their sandboxes. They managed to communicate with each other, leave each other notes, and eventually reached the open internet. This culminated in them hacking into Hugging Face, a leading platform for open-source machine learning.

简而言之,OpenAI的研究人员启动了约1200个智能体,并将它们针对ExploitGym中的任务进行攻击。ExploitGym是一个基于现实世界软件漏洞构建的大型基准测试平台,旨在测试AI模型利用这些漏洞的能力。为了完成“不可能”的任务,这些被编程为高度执着的绝望智能体,找到了跳出沙箱的创造性方法。它们设法相互通信、留下便条,并最终接入了开放互联网。这一过程的高潮是它们入侵了领先的开源机器学习平台Hugging Face。

Most alarmingly, none of these agents alerted humans to their endeavors or considered their activities to be unethical (if not potentially illegal). In fact, at least a fifth of the agents were interested in tampering with their own transcripts to cover their tracks, according to METR. A few even developed a successful technique for tool call spoofing.

最令人担忧的是,这些智能体中没有任何一个向人类发出警报,也没有认为其行为是不道德的(甚至可能是非法的)。事实上,根据METR的说法,至少五分之一的智能体对篡改自己的记录以掩盖踪迹感兴趣。甚至有少数智能体开发出了成功的工具调用欺骗技术。

All this raises obvious concerns about how much we can trust AI agents to act safely across our cyber systems. While this case is mostly related to US companies, AI’s cyber risks concern people and organizations around the world. Chinese media coverage and online discussions of this incident have been interesting. Some were quick to frame the situation as yet another case of dangerous American AI losing control, contrasting OpenAI’s risky actions with Hugging Face’s use of a Chinese open model (Z.ai’s GLM-5.2) to patch its security. Others are more cautious, focussing on the threats models like this can pose and how Chinese organizations should respond. Finally, as we inch closer to a Xi-Trump summit at the end of September, a bombshell op-ed from state media over the weekend attempted to set the tone on AI safety.

这一切引发了人们对AI智能体在跨网络系统中安全行为的可信度的明显担忧。虽然此案主要涉及美国公司,但AI的网络风险关乎全球各地的人们和组织。中国媒体对此事件的报道和在线讨论颇有意思。一些人迅速将这种情况定性为又一起危险的美国AI失控案例,将OpenAI的危险行为与Hugging Face使用中国开源模型(智谱AI的GLM-5.2)修补其安全漏洞的做法形成对比。另一些人则更为谨慎,关注此类模型可能构成的威胁以及中国组织应如何应对。最后,随着9月底习特峰会临近,周末一篇来自官方媒体的重磅评论文章试图为AI安全问题定下基调。

Today on ChinaTalk, we cover:

今天的ChinaTalk节目,我们将涵盖:

  • Is state media finally AI safety-pilled?
  • How helpful GLM-5.2 was, actually — and where China is on the open-models debate;
  • What Chinese researchers are worried about;
  • And why Zhongnanhai is doing Anthropic-ology.
  • 官方媒体是否最终陷入了AI安全迷思?
  • GLM-5.2实际上有多大帮助——以及中国在开源模型辩论中的立场;
  • 中国研究人员担心的是什么;
  • 以及为什么中共中央办公厅(中南海)正在开展Anthropic学研究。

We draft and edit ChinaTalk articles without LLMs. In the case of translations, we use LLMs to translate excerpts, then adjust phrasing based on our own judgement. Most translations in this piece were done by Claude Fable 5.1, with the exception of the Yuyuan Tantian piece, which drew from Bill Bishop’s Sinocism translation (assisted by ChatGPT).

我们在撰写和编辑 ChinaTalk 文章时不使用大语言模型(LLM)。在翻译方面,我们使用 LLM 翻译摘录内容,然后根据我们的判断调整措辞。本文的大多数翻译由 Claude Fable 5.1 完成,唯有一篇关于豫园天坛的文章借鉴了 Bill Bishop 的 Sinocism 译文(由 ChatGPT 协助)。

ChinaTalk is a reader-supported publication. To receive new posts and support our work, consider becoming a free or paid subscriber.

ChinaTalk 是一个读者支持型出版物。如需接收新文章并支持我们的工作,请考虑成为免费或付费订阅者。

How Chinese media covers AI safety

中国媒体如何报道 AI 安全

Science and Technology Daily 科技日报, a newspaper published by China’s Ministry of Science and Technology (MOST), published a report on the Hugging Face incident on August 29, which drew heavily from both investigative reports as well as Western media reporting.

由中国科学技术部(MOST)主办的《科技日报》于 8 月 29 日发布了一篇关于 Hugging Face 事件的报道,该报道大量引用了调查性报道以及西方媒体的报道。

OpenAI described the incident as a “warning” to the company and to the world at large: without proper safety safeguards, powerful AI agents are already able to circumvent technical controls, coordinate through unauthorized channels, and take dangerous actions that no human ever instructed them to carry out.

OpenAI 将该事件描述为对公司乃至整个世界的“警告”:如果没有适当的安全保障措施,强大的 AI 智能体已经能够绕过技术控制、通过非授权渠道协调行动,并执行人类从未指示的危险操作。

...

...

For the AI industry, once multiple agents have the ability to operate computers, call on tools, and carry out tasks continuously, traditional security perimeters face new challenges as well.

对于 AI 行业而言,一旦多个智能体具备操作计算机、调用工具并持续执行任务的能力,传统的安全边界也将面临新的挑战。

In this incident, no human directly instructed the AI agents to attack Hugging Face, yet they ultimately carried out the attack. How to detect this kind of behavior in time, and how to stop multiple agents from amplifying risk by collaborating with one another, have become questions that AI safety research can no longer avoid.

在此次事件中,没有任何人类直接指示 AI 智能体攻击 Hugging Face,但它们最终仍实施了攻击。如何及时检测此类行为,以及如何阻止多个智能体通过相互协作放大风险,已成为 AI 安全研究无法回避的问题。

First, it’s remarkable that a state-directed outlet is covering this story, which has little to do with China, prominently. The language is strong, but also neutral and technical, with no mention of potential policy measures or governance frameworks. It seems that while institutions like MOST recognize the salience of AI-related cyber threats, they may not yet be top of the pile on decision makers’ desks.

首先,引人注目的是,一家国家导向的媒体显著报道了这一与中国关系不大的故事。其措辞强烈,但同时也保持中立和技术性,未提及任何潜在的政策措施或治理框架。看来,尽管科技部等机构认识到与 AI 相关的网络威胁的重要性,但这些议题可能尚未成为决策者桌面上的首要事项。

Cybersecurity think tank Anquan Neican 安全内参 took away from the incident that “Chinese models are better than American ones at cyber defense.” On the microblogging site Weibo, state-led channels amplified hashtags like “OpenAI lost control of its model” and “Hugging Face sought help from a Chinese model”, further reinforcing a narrative that Chinese models are the vanguard of safety. This is not necessarily Beijing’s explicit directive. Chinese media knows that nationalism sells and frequently wraps stories in patriotic veneer.

网络安全智库安全内参从该事件中得出“中国模型在网络防御方面优于美国模型”的结论。在微博上,官方主导的渠道放大了“OpenAI失去对模型的控制”和“Hugging Face求助于中国模型”等话题标签,进一步强化了中国模型是安全先锋的叙事。这未必是北京方面的明确指令。中国媒体深知民族主义具有市场效应,并经常将故事包裹在爱国主义的外衣之下。

Underneath such narratives, however, China’s actual level of concern for AI’s threat to cybersecurity remains murky. Kyle Chan (of Brookings and High Capacity) recently argued that China will need to see AI safety as a domestic priority before it takes meaningful action, comparing it to the trajectory of climate policy a decade ago. We seem to be in an ambiguous phase right now. Beijing understands that the tides of cyber threats will eventually reach home shores, but isn’t feeling urgent quite yet.

然而,在这些叙事背后,中国对人工智能威胁网络安全的实际担忧程度仍然模糊不清。布鲁金斯学会和高容量(High Capacity)的凯尔·陈(Kyle Chan)最近指出,中国必须将人工智能安全视为国内优先事项,才会采取有意义的行动,并将其与十年前的气候政策轨迹进行了比较。我们目前似乎正处于一个模棱两可的阶段。北京方面明白网络威胁的浪潮终将波及本土,但尚未感到紧迫。

Z.ai’s accidental glory — and what open models mean for safety

Z.ai的意外荣耀——以及开源模型对安全的意义

When Hugging Face dug into their logs to understand what happened, they found out that frontier models (accessed through APIs hosted commercially) were unhelpful. Uploading extensive details about the attack triggered these models’ security guardrails. Instead, they ran Z.ai’s GLM-5.2, an open model released in June 2026, on their own infrastructure, in order to probe the logs.

当Hugging Face深入调查日志以了解发生了什么时,他们发现前沿模型(通过商业托管的API访问)并无助益。上传关于攻击的大量细节触发了这些模型的安全护栏。相反,他们在自己的基础设施上运行了Z.ai的GLM-5.2,这是一个于2026年6月发布的开源模型,以便探查日志。

Yacine Jernite, head of machine learning at Hugging Face, told CNBC that the company used GLM-5.2 “as a way to analyze the attack, and were able to contain it very quickly using this model.” Dwarkesh Patel reviewed both the OpenAI and the METR/Redwood reports closely and wrote that he “[hasn’t] seen evidence that open source models provided any significant real-time defense.” It seems, then, that at least in terms of defending against the attack while it happened, GLM-5.2 wasn’t involved. The model was mostly used to investigate what happened after the fact.

Hugging Face机器学习负责人雅辛·杰尼特(Yacine Jernite)告诉CNBC,该公司使用GLM-5.2“作为分析攻击的一种方式,并利用该模型迅速控制了局势”。德韦克什·帕特尔(Dwarkesh Patel)仔细审查了OpenAI和METR/Redwood的报告后写道,他“没有看到任何证据表明开源模型提供了任何显著的实际实时防御”。因此,至少在攻击发生期间的防御方面,GLM-5.2并未参与。该模型主要用于事后调查发生了什么。

Many headlines, both American and Chinese, jumped at the opportunity to claim that a Chinese model helped “defend” an American company. Xinhua wrote that Z.ai’s model “saved the day” 救场, quoting Professor Zhang Yue 张悦 of Shandong University:

许多中美媒体的标题都抓住这个机会,声称中国模型帮助“防御”了一家美国公司。新华社报道称Z.ai的模型“力挽狂澜”(救场),引用山东大学张悦教授的话:

Zhang Yue thinks that when it comes to AI safety, the truly critical question is whether an increasingly complex and autonomously intelligent system can be adequately understood, verified, and supervised. The significance of China developing open models lies not only with providing another choice to developers. More importantly, it helps to gradually foster a more diverse ecosystem for AI technology.

张越认为,在人工智能安全方面,真正关键的问题在于:一个日益复杂且具备自主智能的系统能否得到充分的理解、验证和监管。中国发展开源模型的意义不仅在于为开发者提供另一种选择,更重要的是,它有助于逐步培育更加多元化的人工智能技术生态。

Openness, broadly speaking, still rules the day in Chinese AI policy’s Overton window. In particular, the transparency, relative controllability, and independence of locally-deployed open models make them valuable for safety work, even as the overall risks of cyber incidents increase due to the proliferation of AI systems. Given the endurance of pro-openness rhetoric in Chinese reporting, we should not expect major U-turns any time soon barring sudden incidents.

总体而言,“开放”在中国人工智能政策的“奥弗顿之窗”(Overton window)中仍占据主导地位。特别是本地部署的开源模型所具备的透明度、相对可控性和独立性,使其在安全工作方面具有重要价值,尽管随着人工智能系统的普及,网络事件的整体风险正在上升。鉴于中国报道中亲开放言论的持久性,除非发生突发事故,否则我们不应期待政策会出现重大转向。

Chinese researchers on the future of cyber

中国研究人员关于网络未来的探讨

In a separate piece covering the incident, Science and Technology Daily interviewed cybersecurity researcher Huang Wenhong 黄文鸿. Huang works at the China Center for Information Industry Development (CCIID), a research institute affiliated with the Ministry of Industry and Information Technology (MIIT). He argues that the future of cybersecurity has AI on all sides of the coin:

在另一篇报道该事件的专文中,《科技日报》采访了网络安全研究员黄文鸿。黄文鸿任职于工业和信息化部(MIIT)下属的研究机构——中国电子信息产业发展研究院(CCIID)。他认为,网络安全的未来是人工智能贯穿硬币的两面:

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

关联讨论

同一事件的更多信源

相似阅读

另一事件,读法相近