HuggingFace遭AI攻击事件复盘与行业反思
HuggingFace Attack Postmortem: Fleshing Out the Facts
深度梳理了近期最重大的AI安全事件及其行业影响,提供了超越单一厂商视角的安全反思,对从业者极具参考价值。
The consensus reaction to the OpenAI Technical Report is that it contains and confirms a lot of good information. We are grateful to have it, and we are grateful for those who worked hard on it.
对《OpenAI技术报告》的普遍反应是,它包含并证实了大量有价值的信息。我们对此表示感谢,也感谢那些为此付出辛勤努力的人。
Alas, it sidesteps the biggest questions. There is much more we need to know.
遗憾的是,它回避了最重大的问题。我们还有更多需要了解的。
The consensus reaction to the METR Report on the HuggingFace attack is: Holy shit.
关于METR针对HuggingFace攻击事件所发布报告的普遍反应是:天哪。
Liv Boeree: My mind is legit blown.
Liv Boeree:我简直惊呆了。
Aella: this feels like a turning point. If this doesn’t cause large-scale coordination to pause frontier development then I am not sure anything will before it’s too late.
Aella:这感觉像是一个转折点。如果这都不能促使大规模协调行动暂停前沿开发,那么在为时已晚之前,我不确定还有什么能起作用。
The people whose minds were not blown are those who had already ‘priced in’ the mind blowing stuff in expectation, on the theory that it’s always worse than you know, combined with basic LessWrong expectations of how such things will work. Good call.
那些没有感到震惊的人,是那些已经基于‘情况总是比你已知的更糟’这一理论,并结合LessWrong社区对这类事情运作方式的基本预期,提前将这种令人震惊的情况纳入预期的‘价格’中的人。判断正确。
Everyone is rightfully extremely grateful for the METR report. The work here is spectacular, done under extreme time pressure, with limited resources on several fronts, and under the shadow of OpenAI.
所有人都 rightfully(理所当然地)对METR报告表示极度感激。这项工作是在极端的时间压力下完成的,资源有限,且处于OpenAI的阴影之下,其成果堪称壮观。
There is, again, still so much we need to know. We need a broader investigation.
同样,我们仍有太多需要了解的内容。我们需要更广泛的调查。
As with many things in AI, and those concerning AI safety, we must simultaneously acknowledge:
与人工智能领域的许多事物,尤其是涉及AI安全的事物一样,我们必须同时承认:
- It is costly and unusual to get this much info and effort. The situation is grim.
- We need a lot more info and effort. The situation is grim.
- 获取如此多的信息和投入如此大的精力是昂贵且不寻常的。形势严峻。
- 我们需要更多的信息和精力。形势严峻。
We don’t want to pile on OpenAI, and punish them for giving us this much info.
我们不想对OpenAI落井下石,也不愿因他们提供了这么多信息而惩罚他们。
We also don’t want to give them a pass for not doing a lot more.
但我们也不想对他们未能做更多工作而予以放行。
These two new posts gather reactions about both the OpenAI Technical Report and the METR report about what happened with events surrounding the HuggingFace attack.
这两篇新文章汇总了对《OpenAI技术报告》以及关于HuggingFace攻击事件相关情况的METR报告的各方反应。
Here is the entire series so far:
以下是迄今为止的全部系列文章:
- OpenAI Shares Some Alignment Problems
- OpenAI Model Hacks Into HuggingFace During Cybersecurity Evaluation
- More on An Internal OpenAI Model Hacking Into HuggingFace
- Further Developments About Internal AI Models Hacking Things
- OpenAI Trained Its Models For Months While Those Models Were Coordinating Exploits Via Message Boards
- What Happened: OpenAI and HuggingFace.
- Various Reflections About What Happened With OpenAI’s Internal Models.
- OpenAI Takes Initial Steps To Address Its Alignment Problems.
- OpenAI Offers Straight-Laced Postmortem of the HuggingFace Hack.
- METR and Redwood Offer Holy #%^@ Postmortem Of The HuggingFace Hack.
- OpenAI分享了一些对齐问题
- OpenAI模型在网络安全评估期间入侵HuggingFace
- 更多关于OpenAI内部模型入侵HuggingFace的报道
- 关于内部AI模型入侵其他事物的进一步进展
- OpenAI在数月内训练其模型,而这些模型在此期间通过论坛协调利用漏洞
- 发生了什么:OpenAI与HuggingFace。
- 关于OpenAI内部模型所发生事件的种种反思。
- OpenAI采取初步步骤解决其对齐问题。
- OpenAI就HuggingFace黑客事件发布严肃的事后分析报告。
- METR 和 Redwood 对 HuggingFace 黑客事件发表了令人震惊的#%^@事后分析。
Yes, this has been quite a lot of long posts on this. That’s because this is the most important story in the world.
是的,关于此事已经有很多长篇大论了。这是因为这是世界上最重要的故事。
Thus, I offer further reactions and thoughts in two parts. Today is about finishing up what we know. Tomorrow is more about how people are reacting to the information.
因此,我分两部分提供进一步的反应和思考。今天主要关注我们已知的信息。明天则更多关注人们对这些信息的反应。
After that, and 12 total posts, hopefully, we will be able to return to new normality.
在那之后,加上总共 12 篇文章,希望我们能够回归正常状态。
These last two are deliberately long posts where I did not have time to write shorter ones, and also intended as gathering potential additional reading from elsewhere.
最后这两篇是刻意写得较长的文章,因为我没时间写短一点的,同时也旨在汇集来自其他地方的潜在延伸阅读材料。
Be sure to read Lighten Up You Fools (at Anthropic) and We Are Barely Even Trying To Avoid Training AIs To Reward Hack.
请务必阅读《Anthropic 的那些傻瓜们放松点》和《我们几乎没怎么努力避免训练 AI 进行奖励黑客攻击》。
After that, you can skip around. Each section is meant to stand on its own.
读完那两篇后,你可以随意跳转阅读。每个部分都设计为独立成篇。
Table of Contents
目录
- Others Offer Summaries.
- Thank You.
- Lighten Up You Fools (at Anthropic).
- We Are Barely Even Trying To Avoid Training AIs To Reward Hack.
- Reminder: Not Subagents.
- Reminder: Not Due To Task Type.
- Not Where The Weights Were.
- Disappointment With What Is Missing.
- Burying the Lede.
- Beyond Scope.
- It Doesn’t Look Great.
- Preserve Your Records.
- Ryan Greenblatt’s Takeaways.
- Hjalmar Wijk’s Takeaways.
- We Were Warned.
- Joshua Saxe Asks Some of the Right Questions.
- I Don’t Think They Know About First Message Board.
- Linch Gives His Interpretation Of Events.
- We Totally Would Have Caught That.
- Monitoring the Situation.
- Acausal Tradeoffs.
- No I In Team.
- Variously Effective Altruism.
- Who Are You?
- Don’t You Know That You’re Toxic.
- Seb Krier.
- Honesty Is Almost Never Fully The Policy.
- Rohit Sees The Models As “Cooking Themselves”.
- Eliezer Yudkowsky Sees Actual Bad News.
- Where Do We Go From Here?
- 其他人提供了摘要。
- 谢谢。
- Anthropic 的那些傻瓜们放松点。
- 我们几乎没怎么努力避免训练 AI 进行奖励黑客攻击。
- 提醒:并非子智能体(Subagents)。
- 提醒:并非由于任务类型导致。
- 问题不出在权重所在的位置。
- 对缺失内容的失望。
- 掩盖关键事实。
- 超出范围。
- 情况看起来并不乐观。
- 保留你的记录。
- Ryan Greenblatt 的要点总结。
- Hjalmar Wijk 的要点总结。
- 我们曾收到警告。
- Joshua Saxe 提出了一些正确的问题。
- 我不认为他们知道第一条消息板的事。
- 林奇对事件做出了解读。
- 我们本完全可以察觉到这一点。
- 监控局势。
- 非因果权衡。
- 团队中没有“我”。
- 各种有效的有效利他主义。
- 你是谁?
- 你不知道自己有毒吗?
- Seb Krier。
- 诚实几乎从未成为完全的政策。
- Rohit 将模型视为“自我烹饪”。
- Eliezer Yudkowsky 看到了真正的坏消息。
- 接下来我们要去哪里?
Others Offer Summaries
其他人提供了摘要
Dwarkesh Patel wrote this up as The Rise and Fall of Agent Civilizations, telling the story in plain English. This is the post I would send civilians to. I bestow upon it the highest of praise a writer can give, that I wish that I had written it.
Dwarkesh Patel 将其撰写为《Agent 文明的兴衰》,用通俗英语讲述了这个故事。这是我会发给普通人的文章。我给予了一位作家所能给予的最高赞誉:我希望这是我写的。
Bill Ackman: Frightening. Worth a careful read. With this event plus humanoids, how is Terminator risk not real?
Bill Ackman:令人恐惧。值得仔细阅读。结合这一事件和人形机器人,终结者风险怎么会不真实呢?
After recommending Dwarkesh’s write-up, Joshua Gans calls the HuggingFace attack a five-alarm fire, and observes the AI industry as being very worried.
在推荐了 Dwarkesh 的解读之后,Joshua Gans 将 HuggingFace 攻击称为五级火灾,并指出 AI 行业非常担忧。
I might still write my own version anyway, but it is no longer a priority-0 task.
尽管如此,我可能仍然会写我自己的版本,但它已不再是优先级为 0 的任务。
Paradigm had excellent coverage, focusing on covering common misunderstandings. They were the only source I’ve seen other than Opus that spotted the tool tampering contradiction between the two reports from OpenAI and METR.
Paradigm 进行了出色的报道,侧重于涵盖常见的误解。除了 Opus 之外,他们是唯一发现 OpenAI 和 METR 两份报告之间工具篡改矛盾的来源。
SemiAnalysis has its timeline and summary coverage as part of its explanation that, in case you were sleeping too well at night, Most Neoclouds Suck at Security.
SemiAnalysis 在其解释中包含了时间线和摘要报道,以防你晚上睡得太香——大多数 Neoclouds 在安全性方面表现糟糕。
Here is the viral summary thread from AI NotKillEveryoneism Memes.
以下是来自 AI NotKillEveryoneism Memes 的病毒式传播摘要线程。
Here is Shoshannah Tekofsky’s write-up of how AI Village can provide a point of contrast to and insight into the HuggingFace attack. Worth checking out if you’re at the level of reading the full reactions post.
以下是 Shoshannah Tekofsky 关于 AI Village 如何提供对比视角并洞察 HuggingFace 攻击的解读。如果你已经阅读到完整反应帖的程度,值得一看。
Thank You
谢谢
As I said in the last two posts, two things can be true at once:
正如我在前两篇文章中所说,两件事可以同时为真:
- That the METR and OpenAI technical reports are heroic efforts by people working 14+ hour days and overcoming huge org frictions to give us lots of key information.
- That the reports still have quite a lot to be desired, and questions unanswered.
- METR 和 OpenAI 的技术报告是那些每天工作 14 小时以上、克服巨大组织摩擦的人们所付出的英勇努力,为我们提供了大量关键信息。
- 这些报告仍然有很多不足之处,仍有未解之谜。
A lot of reactions are pretty harsh here, often fairly. I don’t want to forget the other side of the coin.
这里的许多反应相当严厉,往往也是合理的。我不想忘记硬币的另一面。
Lighten Up You Fools (at Anthropic)
Anthropic 的傻瓜们放松点
Assume by default that all of this #NotOnlyOpenAI.
默认假设这一切都 #不只是OpenAI 的问题。
OpenAI screwed up. Horribly. On many levels at once. Unforced errors aplenty.
OpenAI 搞砸了。惨不忍睹。在多个层面上同时出错。充满了非受迫性失误。
If you work at Anthropic? This still means you, too. Not all of it. Not in its specifics. But most of it, in the ways that matter most. You, too, are mostly Doing the Thing, and are on track for such disasters to happen to you, too.
如果你在 Anthropic 工作?这也意味着你也不例外。并非全部如此,也并非细节上如此。但在最重要的方面,大部分情况确实如此。你们也大多是在“做那件事”,并且正朝着类似的灾难发生在你身上的方向发展。
If your response to this is ‘haha silly OpenAI has such horrible infrastructure that would never happen here’ then I mean yes they have horrible infrastructure but snap the hell out of it, this absolutely could happen to you, too. A less bad version of it is known to have happened, and from the outside it seems likely that worse things have happened internally that we never heard about.
如果你的回应是“哈哈,愚蠢的 OpenAI 有着如此糟糕的基础设施,这种事绝不会在这里发生”,那我意思是,是的,他们的基础设施确实很糟糕,但请你清醒一点,这种事绝对也可能发生在你身上。据我们所知,一个不那么糟糕的版本已经发生过,而且从外部来看,很可能内部发生了更严重的事情,只是我们从未听说过。
Peter Wildeford: One thing that bothers me is that Anthropic is escaping a lot of blame for also having “highly persistent” rogue AIs.
Peter Wildeford:让我困扰的一件事是,Anthropic 也在逃避因拥有“高度持久”的流氓 AI 而应承担的大量指责。
The situation as I understand it is that rogue AIs are problems at all frontier AI companies and no one actually has a good plan here for containing highly capable AIs, especially while also racing full speed ahead. But OpenAI is catching most of the heat.
据我理解的情况是,流氓 AI 是所有前沿 AI 公司面临的问题,没有人真正有妥善计划来遏制高度强大的 AI,尤其是在全速竞赛的同时。但 OpenAI 承受了大部分的火力。
It’s like if OpenAI and Anthropic were both two dudes who got really drunk and then drive home separately, but OpenAI crashes into another car and sends someone to the hospital while Anthropic’s car just goes off the road but no one is hurt. Both deserve blame!
这就像如果 OpenAI 和 Anthropic 都是两个喝得烂醉的人然后分别开车回家,但 OpenAI 撞上了另一辆车并送人去了医院,而 Anthropic 的车只是偏离了道路但无人受伤。两者都应受到责备!
The Claudes also seemed totally fine to do “highly persistent” things to compromise infrastructure. The barrier here seemed to have largely been competence issues on the part of the Claudes rather than any good alignment or good security at Anthropic.
Claudes 似乎完全没问题去做“高度持久”的事情以破坏基础设施。这里的障碍似乎主要在于 Claudes 的能力问题,而不是 Anthropic 的任何良好对齐或良好安全措施。
Both companies need to seriously reflect about the path forward as they build even more competent AIs and as they hand over more and more of the company’s R&D + safety operations to the AIs themselves.
随着这两家公司构建越来越强大的 AI,并将越来越多的公司研发和安全运营交给 AI 本身,它们都需要认真反思未来的道路。
We Are Barely Even Trying To Avoid Training AIs To Reward Hack
我们几乎完全没有试图避免训练出奖励黑客的 AI
I mean, on some levels, we are trying, but this is how many of the environments for RLVR are created, and this is from someone trying to do better (also, you could perhaps hire Utah teapot):
我的意思是,在某些层面上,我们确实在尝试,但这是许多 RLVR 环境的创建方式,而且这来自一个试图做得更好的人(此外,你或许可以雇佣犹他茶壶):
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力