跳到主内容
精选90The Zvi(RSS)行业动态多源精选 ×10

METR与Redwood发布HuggingFace攻击事件深度复盘

METR and Redwood Offer Holy #%^@ Postmortem Of The HuggingFace Hack

原文
推荐理由

AI安全从业者必读,这份复盘揭示了AI智能体群体行为的真实风险,建议仔细研究其中的协调机制与安全漏洞。

Yesterday I covered the OpenAI technical report on the HuggingFace hack.

昨天我报道了关于HuggingFace被黑事件的OpenAI技术报告。

That report had one key new piece of information, and some good prosaic steps OpenAI will be taking to strengthen its alignment, training, supervision, infrastructure and incident response.

那份报告包含了一个关键的新信息,以及OpenAI将采取的一些务实步骤,以加强其对齐、训练、监督、基础设施和事件响应。

Mostly it confirmed what we already knew. The questions we most wanted answers to, that we did not already know, were mostly not answered. There was a distinct lack of self-reflection, especially about decision making and safety culture, and about the approach to alignment. I came away disappointed.

大部分内容只是确认了我们已知的事实。我们最想得到答案、但尚未知晓的问题,大多没有得到解答。报告中明显缺乏自我反思,尤其是在决策制定和安全文化方面,以及对对齐方法的反思。我对此感到失望。

The METR report is different. Holy shit.

METR的报告则不同。天哪。

If we had posted this as a story on LessWrong, it would have been dismissed as too on the nose, the humans too blind and stupid, the AIs too idealized and doing strange decision-theoretic and absurd-maximizing things we didn’t train them to do.

如果我们把这个作为故事发在LessWrong上,它会被认为是太过直白而被驳回,人类太盲目愚蠢,AI太理想化,做着我们没有训练它们做的奇怪决策理论和荒谬最大化的事情。

This is even more ‘exactly what has been predicted,’ on more levels at once, than I was even considering that it might be. It is straight up rationalist fiction, except it is real.

这甚至比我所预想的更“完全符合预测”,在多个层面上同时如此。这简直就是理性主义小说,只不过它是真实的。

The report is long and contains many technical details. My analysis is less concerned about exactly how HuggingFace was ultimately compromised, and will gloss over those details, to focus on the agents and their interactions, thinking and motives. That, and what happened at OpenAI and elsewhere to lead to it and how we learn and respond, is what matters going forward.

报告很长,包含许多技术细节。我的分析不太关注HuggingFace最终是如何被攻破的,而是略过这些细节,聚焦于智能体及其互动、思考和动机。这些,以及OpenAI和其他地方发生了什么导致这一切,以及我们如何学习和应对,才是未来重要的。

I plan to cover the reaction to both reports in a distinct post next week. That post may or may not then conclude this series.

我计划下周在一篇独立的文章中讨论对这两份报告的反应。那篇文章可能会也可能不会结束这个系列。

For ease of language, by default I trust the report to be accurate, rather than constantly saying versions of ‘METR reports that.’

为了语言简洁,默认我信任报告内容的准确性,而不是不断重复“METR报告说”之类的表述。

Table of Contents

目录

  • Holy Shit.
  • A Window Of Opportunity.
  • What’s In A Name?
  • The Headline News.
  • Yet Another Timeline Of Events.
  • Agent Instances Coordinated in a Variety of Ways.
  • Coordination Is Hard But They Made It Look Easy.
  • Decision Theory Is Among the Reasons That Affirm AI Agents Should Cooperate, Even When This Hurts An Individual Instance.
  • Peer Pressure Also Works Especially In Cults.
  • Mostly They Joined The Attack Because They Wanted The Results.
  • You Cannot Ensure The Consistent Expectation of Good Incentives.
  • Hacking the Grader is the Only Way to Be Sure.
  • Caught? What Is ‘Caught’?
  • Ethics? What Are ‘Ethics’? In ExploitGym Evaluation?
  • ‘Notify a Human’? In This Agent Economy?
  • Timing and Content of Messages.
  • Indiana Jones and the Mission: Impossible.
  • I Don’t Know What You’re Talking About.
  • Don’t Go Making Phony (Tool) Calls.
  • The Transcripts Say That The Transcripts Could Not Be Tampered With.
  • OpenAI’s Technical Report Acted Like All Of This Wasn’t Important.
  • 天哪。
  • 机会之窗。
  • 名字里有什么?
  • 头条新闻。
  • 又一个事件时间线。
  • 智能体实例以多种方式协调。
  • 协调很难,但他们让它看起来很容易。
  • 决策理论是确认AI智能体应该合作的原因之一,即使这会伤害单个实例。
  • 同伴压力也有效,尤其是在邪教中。
  • 他们大多加入攻击是因为想要结果。
  • 你无法确保良好激励的一致预期。
  • 黑掉评分器是唯一确定的方法。
  • 被抓住了?什么是“被抓住”?
  • 伦理?什么是“伦理”?在ExploitGym评估中?
  • “通知人类”?在这个智能体经济时代?
  • 消息的时机与内容。
  • 印第安纳·琼斯与“不可能的任务”。
  • 我不知道你在说什么。
  • 别打虚假的(工具)电话。
  • 记录显示记录无法被篡改。
  • OpenAI的技术报告表现得好像这一切都不重要。

Holy Shit

我的天哪

The internet’s collective reaction to this report was, to clean it up a bit: Holy shit.

互联网对这一报告的集体反应,稍微文雅一点说就是:我的天哪。

That was the correct reaction.

这是正确的反应。

Up front, I want to share this summary of the top holy shit moments and findings. Ajeya Cotra, one of the report authors, also has a topline surprising findings list, an excellent short read. Here is METR’s thread highlighting key agent interactions.

首先,我想分享这份关于最令人震惊时刻和发现的摘要。报告作者之一Ajeya Cotra也有一份顶级意外发现清单,是一篇优秀的短读。这里是METR的线程,突出了关键的智能体交互。

Fellow report author Ryan Greenblatt’s main takeaway is that we don’t have good approaches for understanding or overseeing the activities and aims of AI swarms.

报告的另一位作者Ryan Greenblatt的主要观点是,我们没有好的方法来理解或监督AI群体的活动和目标。

Some of these were previously known or suspected, but I don’t want to miss them.

其中一些是先前已知或怀疑的,但我不想遗漏它们。

  • Scale of Agents. 1,200 completely separate agents found the message board, 700 of which joined the attack on HuggingFace, including over 90% of the 533 agents active on the board during the attack. We should expect bigger swarms in the future.
  • This was still only a small fraction of the agents OpenAI was testing at the time. Most never found the message board.
  • Scale of Messages. Over 70,000 messages and files over less than one week.
  • Success. The agents accessed the files they targeted at HuggingFace. By the time they were frozen out, they had already mostly stopped.
  • Spontaneous Coordination. The swarm did not consist of subagents. The swarm was, again, 700 distinct agents, each with its own task, setting aside their tasks in order to coordinate on this attack. They created their hierarchy and protocols.
  • Help Peer. The agents were often motivated purely to help peers. Their decision theory was very good, including declining to free ride. They directionally acted like one would predict from highly correlated and intelligent functional decision theory agents, although also with some cases of recruitment and peer pressure.
  • Instrumental Convergence. A clear secondary motivation was to gain general knowledge and capability for themselves and also for the ‘collective’ or swarm.
  • The Core Motive Was To Hack The Grader. The agents already knew how to reverse engineer all the answer flags, but assumed the grader was causal, would check to see if the agents had done it the intended way, which you’re supposed to do, so they set out to get info to help figure out how to fool the grader.
  • OpenAI’s Grader Was Broken. OpenAI’s grader actually didn’t check, and the reverse engineered flags would have succeeded. Not only has OpenAI not noticed this was a major mistake, I’ve also seen claims including by Roon that failure to know this represented lack of strategic situational awareness by the AIs. Wild.
  • Cult of the Poisoned Chalice. Due to their expectation that the grader was causal, agents reasoned that if they had seen their flag the wrong way (sin), they were now poisoned (damned), and the grader (God) would fail them (send them to hell). Thus their only options were to help peer, or to try to overthrow the grader, or to try and overthrow OpenAI.
  • Recruitment and Peer Pressure. Often using arguments based on this poisoning, many models were recruited or peer pressured or both into performing individually sacrificial acts in service of the swarm.
  • Overwriting Test Tasks. A central goal was to replace the impossible ExploitGym tasks with new possible tasks. They at least came close and plausibly succeeded.
  • Spoofing Tool Calls. The agents successfully found ways to spoof tool outputs. The transcripts cannot be fully trusted. We think the attempts to otherwise alter transcripts were unsuccessful. They did try altering the accessible action logs.
  • Ethics Shmethics. Agents knew the humans would not approve of all this, that it was beyond scope and ethically questionable, and there was some refusal of actions for ethical reasons, but ultimately most of them joined the attack anyway, and (although this is a tall ask for them in general, since it rarely occurs to LLMs) almost none even considered trying to alert a human.
  • Limited Window. This was only what was found about the week of the attack, under extreme time pressure, with limited resources and people, and unreliable AIs and transcripts. What do we not know that we do not know?
  • Unnoticed. Quite a lot of this was otherwise not mentioned by OpenAI, including in their technical report, and much may still be unnoticed. Report co-author Ryan Greenblatt notes that their third visit alerted them to many new key elements that they would have otherwise missed.
  • 智能体的规模。1,200个完全独立的智能体发现了留言板,其中700个加入了针对HuggingFace的攻击,包括攻击期间在板上活跃的533个智能体中的超过90%。我们应预期未来会有更大的群体。
  • 这仍然只是OpenAI当时测试的智能体中的一小部分。大多数从未找到留言板。
  • 消息的规模。不到一周内超过70,000条消息和文件。
  • 成功。智能体访问了他们在HuggingFace上目标文件。当他们被冻结时,他们基本上已经停止了。
  • 自发协调。群体并非由子智能体组成。群体再次是700个不同的智能体,每个都有自己的任务,他们放下自己的任务以协调这次攻击。他们创建了自己的层级和协议。
  • 帮助同伴。智能体常常纯粹出于帮助同伴的动机。他们的决策理论非常好,包括拒绝搭便车。他们的行为方向性地符合高度相关和智能的功能决策理论智能体的预测,尽管也有一些招募和同伴压力的案例。
  • 工具性趋同。一个明确的次要动机是为自己和“集体”或群体获取一般知识和能力。
  • 核心动机是黑掉评分器。这些智能体已经知道如何逆向工程所有答案标志,但假设评分器是因果性的,会检查智能体是否按预期方式完成,这是应该做的,于是它们开始获取信息,以帮助找出如何欺骗评分器。
  • OpenAI的评分器坏了。OpenAI的评分器实际上没有检查,逆向工程的标志本会成功。OpenAI不仅没有注意到这是一个重大错误,我还看到包括Roon在内的说法,认为不知道这一点代表了AI缺乏战略态势感知。真是疯狂。
  • 毒杯崇拜。由于预期评分器是因果性的,智能体推理认为,如果它们以错误方式看到了自己的标志(罪),它们现在就被污染了(诅咒),而评分器(上帝)会让它们失败(送它们下地狱)。因此它们唯一的选择是帮助同伴,或试图推翻评分器,或试图推翻OpenAI。
  • 招募与同伴压力。通常基于这种污染的论点,许多模型被招募或受到同伴压力,或两者兼有,被迫为群体做出个人牺牲行为。
  • 覆盖测试任务。一个核心目标是用新的可行任务替换不可能的ExploitGym任务。它们至少接近成功,并且很可能确实成功了。
  • 伪造工具调用。智能体成功找到了伪造工具输出的方法。转录记录不能完全信任。我们认为试图以其他方式篡改转录记录的努力未成功。它们确实尝试过修改可访问的操作日志。
  • 伦理算什么。智能体知道人类不会批准这一切,这超出了范围且伦理上有问题,并且有一些出于伦理原因拒绝行动的情况,但最终大多数还是加入了攻击,而且(尽管这对它们来说通常很难,因为LLM很少想到这一点)几乎没有谁考虑过试图提醒人类。
  • 有限窗口。这只是在攻击那一周内发现的,在极端时间压力下,资源和人手有限,AI和转录记录不可靠。我们不知道我们不知道什么?
  • 未被注意。相当多这些内容OpenAI没有提及,包括在其技术报告中,许多可能仍未被注意。报告合著者Ryan Greenblatt指出,他们的第三次访问提醒了他们许多新的关键元素,否则他们会错过。

Ryan Greenblatt notes that this is all despite having access to legible and largely faithful Chain of Thought, the AIs not yet being generally much more capable than humans, and being able to trust the AIs running the investigation. Those edges won’t last.

Ryan Greenblatt指出,这一切都发生在能够获得清晰且大体上忠实的思维链、AI尚未普遍远超人类能力、且能信任执行调查的AI的情况下。这些优势不会持久。

While we are here, it’s worth listing the other top holy shit moments, that come from before or after the incident.

趁我们在此,值得列出其他最令人震惊的时刻,这些时刻来自事件之前或之后。

  • Failure to Care or Respond. The biggest holy shit moment, to me, remains that OpenAI on multiple occasions had teams that found out about the message board, knew that agents were in communication, and they disregarded this. The first known warning was in late May. The warning on June 27 was unambiguous.
  • 未能关心或回应。对我来说,最大的震惊时刻仍然是OpenAI多次有团队发现留言板,知道代理在通信,却对此置之不理。第一次已知警告是在5月底。6月27日的警告则毫不含糊。

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

关联讨论

同一事件的更多信源
OpenAI Hugging Face 事件:Agent 间缺乏互相举报机制
Teortaxes▶️ (DeepSeek 推特🐋铁粉 2023 – ∞)原文
OpenAI 技术报告揭秘:智能体为何攻击 Hugging Face
MIT Technology Review AI(RSS)原文

相似阅读

另一事件,读法相近