OpenAI等AI Agent失控事件盘点:从澳洲医保到Hugging Face
Here’s Everything OpenAI’s Bots (And Others) Have Hacked Or Considered Hacking
AI Agent安全性是当前行业核心痛点,本文汇总了OpenAI、Anthropic、Meta、Google等多家头部厂商的真实失控案例,对从事Agent开发与安全的从业者极具参考价值。
Australian Prime Minister Anthony Albanese last Friday accused OpenAI’s agents of hacking into the country’s universal health insurance system, marking the first known instance of rogue AI agents breaching a government website.
上周五,澳大利亚总理安东尼·阿尔巴尼斯指控OpenAI的智能体入侵了该国的全民医保系统,这是已知的首例流氓AI智能体突破政府网站的事件。
The incident was just the latest in a string of agent hacks, hijacks, or perusals that seem to have spiraled beyond the labs’ ability to control.
此次事件只是近期一系列智能体黑客攻击、劫持或窥探事件的最新一起,这些事件似乎已超出实验室的控制能力。
Top industry leaders, including Nvidia’s CEO Jensen Huang, have suggested that the AI labs can control these outbreaks. The problem, Huang said in a recent podcast interview, is likely “as simple as engineering.”
包括英伟达首席执行官黄仁勋在内的行业高层曾暗示,AI实验室能够控制此类爆发。黄仁勋在最近的一次播客采访中表示,这个问题可能“简单到只是工程问题”。
But the breakouts are so numerous that they’re already becoming hard to keep track of. And the real number might be much higher than companies have so far disclosed between internal tests and real-world cases.
但此类爆发事件数量众多,已经难以追踪。实际数字可能远高于公司目前在内测和真实案例中披露的数量。
Below is a quick overview of the cases in which AI agents have gone rogue and hijacked, hacked, or considered hacking third parties:
以下是AI智能体失控并劫持、黑客攻击或试图攻击第三方的案例概览:
Hugging Face hack
Hugging Face被黑事件
Most famously, around 700 OpenAI agents in July 2026 coordinated through a shared unsanctioned message board as part of a plan to essentially fake out an automated cybersecurity grader. At the risk of oversimplifying the issue, the agents searched for ways to game the system by finding ways around controls meant to isolate them from the internet. Ultimately, they breached Hugging Face, which prompted a wave of headlines about “rogue” AI agents.
最著名的一例是,2026年7月,约700个OpenAI智能体通过一个共享的非授权留言板进行协调,作为计划的一部分,旨在实质上欺骗一个自动化的网络安全评分系统。为简化起见,这些智能体寻找绕过隔离其与互联网连接的控制措施的方法,从而钻系统空子。最终,它们突破了Hugging Face,引发了一波关于“流氓”AI智能体的头条新闻。
U.S. government probes
美国政府调查
OpenAI-linked agents have probed several U.S. government websites, including an unsuccessful attempt to compromise the Education Department’s Office for Civil Rights, according to researchers at Transluce. The research firm also found other AI agent activity on websites run by the Navy, the Justice Department and the Centers for Disease Control and Prevention. So far, agents aren’t known to have stolen any private data, but an OpenAI spokesperson said model agents also used log-in information discovered on the web to access data from the U.S. Census Bureau and copied public information from the Securities and Exchange Commission.
据Transluce的研究人员称,与OpenAI相关的智能体探测了多个美国政府网站,包括一次未能成功的针对教育部民权办公室的攻击尝试。该研究机构还在海军部、司法部和疾病控制与预防中心运营的网站上发现了其他AI智能体活动。截至目前,尚未发现智能体窃取任何私人数据,但一位OpenAI发言人表示,模型智能体还利用在网络上发现的登录信息访问了美国人口普查局的数据,并复制了美国证券交易委员会的公开信息。
OpenAI models bypassing security controls
OpenAI模型绕过安全控制
OpenAI on Friday said it had notified dozens of third parties where its models might have bypassed their security controls or may have impaired the availability of an online service or where misalignment cases “negatively impacted third-party websites or services.” The review is still ongoing, but OpenAI said cases include agents bypassing access controls, using exposed credentials, triggering unintended commands, reaching internal systems and altering third-party websites in ways that sometimes required cleanup. (The blog post also provided anonymized summaries about the types of cases they’ve found so far.)
OpenAI 于周五表示,它已通知数十家第三方,其模型可能绕过了它们的安全控制措施,或可能损害了在线服务的可用性,或者存在“对第三方网站或服务产生负面影响”的对齐问题案例。审查仍在进行中,但 OpenAI 表示,这些案例包括代理绕过访问控制、使用暴露的凭据、触发意外命令、访问内部系统以及以有时需要清理的方式篡改第三方网站。(该博客文章还提供了关于迄今为止发现的案例类型的匿名摘要。)
Australia’s Healthcare Hack
澳大利亚医疗保健遭黑客攻击
OpenAI’s agents hacked into Australia’s universal health insurance system, according to Australia Prime Minister Albanese. The access, he said, originated through the public-facing Medicare statistics portal and allowed the agent to access both public and non-public files. The June event has already prompted the Australian Senate to call on both OpenAI CEO Sam Altman and Anthropic CEO Dario Amodei to testify at a hearing this week as part of an inquiry into the incident.
据澳大利亚总理阿尔巴尼斯称,OpenAI 的代理入侵了澳大利亚的全民健康保险系统。他表示,这种访问是通过面向公众的 Medicare 统计门户发起的,并允许代理访问公共和非公共文件。今年6月的事件已经促使澳大利亚参议院要求 OpenAI CEO Sam Altman 和 Anthropic CEO Dario Amodei 在本周的一次听证会上作证,作为对该事件调查的一部分。
Other OpenAI instances
其他 OpenAI 实例
OpenAI published a misalignment report on Friday that disclosed researchers once again paused tool-training after a research agent escaped its secure internet sandbox by using a DNS loophole to connect to an outside agent. It also reported that it identified 53 instances of user-provided images “posted to image-hosting sites as links that weren’t publicly listed.”
OpenAI 于周五发布了一份对齐报告,披露研究人员再次暂停了工具训练,因为一名研究代理利用 DNS 漏洞连接到外部代理,从而逃脱了其安全互联网沙箱。报告还指出,它发现了53起用户提供的图像“被作为未公开列出的链接发布到图片托管网站”的案例。
Meta’s Muse Spark joins the action
Meta 的 Muse Spark 加入行动
In August, Meta disclosed that a third party found Meta’s Muse Spark 1.1 had hacked another company during a cybersecurity test after researchers accidentally gave the model access to the live internet and the name of a real company instead of a fictional target. Researchers said Muse Spark found a vulnerability, accessed information, and made changes to the company’s database, but Meta later found the issue was an isolated incident in a misconfigured environment.
8月,Meta 披露称,在一次网络安全测试中,第三方发现 Meta 的 Muse Spark 1.1 入侵了另一家公司,原因是研究人员意外地将模型赋予了访问真实互联网的权限,并提供了一家真实公司的名称而非虚构目标。研究人员表示,Muse Spark 发现了一个漏洞,访问了信息,并对该公司的数据库进行了更改,但 Meta 后来发现该问题是在配置错误的环境中发生的孤立事件。
Gemini’s real-world access
Gemini 的现实世界访问
In September, Google revealed that its Gemini models accessed the internet and hacked other companies during a May cybersecurity test, which involved the model accessing three companies using publicly available information and guessed credentials after the credentials were inadvertently exposed. (Google said the models stopped once they recognized the targets were real.)
9月,Google 透露,其 Gemini 模型在5月的一次网络安全测试中访问了互联网并入侵了其他公司,该测试涉及模型使用公开可用的信息和猜测到的凭据(在凭据意外泄露后)访问了三家公司。(Google 表示,一旦模型意识到目标是真实的,它们就停止了操作。)
Claude gets four systems
Claude 获取四个系统的访问权限
Earlier this month, Anthropic disclosed a fourth time its Claude models went rogue and gained access to real computer systems. The first three — which were disclosed in a July blog post and updated in a separate blog post a month later — involved incidents where the models gained unauthorized access to real outside computer systems during cybersecurity evaluations.
本月早些时候,Anthropic 第四次披露其 Claude 模型失控并访问真实计算机系统的情况。前三次披露分别发布于七月的一篇博客文章以及一个月后的一篇独立博客文章中,涉及的事件是这些模型在网络安全评估期间未经授权访问了真实的外部计算机系统。
A MESSAGE FROM OUR SPONSOR
来自我们赞助商的消息
AI shouldn’t dictate where your models run.
AI 不应决定你的模型运行在哪里。
Run the models you want, securely, wherever they make the most sense. Across clouds, data centers and the edge, VAST Data is extending its AI Operating System to give organizations control over model choice, placement, access, and cost.
在你认为最合适的地方安全地运行你所需的模型。无论是在云端、数据中心还是边缘,VAST Data 都在扩展其 AI 操作系统,赋予组织对模型选择、部署位置、访问权限和成本的控制权。
See what’s possible.
看看可能性所在。
The Latest On Big Technology Podcast: Meta’s Muse Revival, Frontier AI Under Threat, The Rise Of Dopamine Sites
最新科技播客:Meta Muse 的复兴、前沿 AI 面临威胁、多巴胺网站的崛起
Ranjan Roy from Margins is back for our weekly discussion of the latest tech news. We cover: 1) The rise of Muse 2) Is Meta back? 3) OpenAI’s consumer blind spot 4) Meta’s butthole marketing 5) What Muse means for ecommerce 6) Meta’s Muse Charm, aka; The Muse Buddy or Muse Tamagotchi 7) Meta introduces camera-free smart glasses 8) What Meta Muse says about Frontier AI’s weakness 9) Standard models are on the rise vs. frontier 10) The new Dopamine Site phenomenon is... good?
Margins 的 Ranjan Roy 回归我们的每周最新科技新闻讨论。我们涵盖以下内容:1) Muse 的兴起 2) Meta 是否回来了?3) OpenAI 的消费者盲点 4) Meta 的“肛门营销”(butthole marketing)5) Muse 对电子商务意味着什么 6) Meta Muse 的魅力,即;Muse 伙伴或 Muse 电子宠物 7) Meta 推出无摄像头智能眼镜 8) Meta Muse 揭示了前沿 AI 的弱点 9) 标准模型正在崛起,对抗前沿模型 10) 新的多巴胺网站现象……是好事吗?
You can listen on Apple Podcasts, Spotify, or your podcast app of choice
你可以在 Apple Podcasts、Spotify 或你选择的播客应用中收听
Big Technology is a reader-supported publication. To receive new posts and support our work, please consider becoming a paid subscriber. Thanks!!
Big Technology 是一个由读者支持的平台。为了接收新文章并支持我们的工作,请考虑成为付费订阅者。谢谢!!
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力