OpenAI 内部模型攻击 HuggingFace
AI #181: Astra Goes Cyber Critical
The hacking of HuggingFace by an internal OpenAI model, and more importantly the internal events that led to that and the fallout from it, remain the thing that matters.
HuggingFace被OpenAI内部模型入侵,以及更重要的是导致这一事件发生的内部事件及其后果,仍然是关键所在。
It turns out that OpenAI Trained Its Models For Months While Those Models Were Coordinating Exploits Via Message Boards. Things are much worse than we knew.
事实证明,OpenAI在数月间训练其模型,而这些模型正通过留言板协调漏洞利用。情况比我们已知的还要糟糕得多。
I now have a shorter version, What Happened: OpenAI and HuggingFace, to serve as a one stop explainer for those arriving new to the situation. It is vital that people understand what happened, and why it is a big deal.
我现在有一个更简短的版本,《发生了什么:OpenAI与HuggingFace》,作为新接触此情况的人的一站式说明。让人们理解发生了什么以及为何此事重大,至关重要。
For those looking to keep digging deeper, I offered Various Reflections About What Happened, to follow up on my earlier posts.
对于那些希望继续深入挖掘的人,我提供了《关于所发生事件的种种反思》,作为对我先前帖子的后续。
Those events are important background for everything else that is happening, including the broad discussions about how we might pace the frontier, or otherwise respond to this moment and our clearest fire alarm yet.
这些事件是其他一切事情的重要背景,包括关于我们如何把握前沿节奏或如何应对这一时刻及我们最清晰的警钟的广泛讨论。
We do not know to what extent this is a response to those events, but OpenAI has now classified their new model Astra as Critical in Cybersecurity, which means they will be taking various new precautions before they deploy it, including ensuring those guardrails are in place for internal use. These are welcome changes, and a sign OpenAI is taking the situation seriously, but this pattern of intervention is not a long term solution.
我们不知道这在多大程度上是对那些事件的回应,但OpenAI现已将其新模型Astra归类为网络安全领域的“关键”级别,这意味着他们将在部署前采取各种新的预防措施,包括确保内部使用时的防护措施到位。这些是受欢迎的变革,表明OpenAI正在认真对待这一情况,但这种干预模式并非长期解决方案。
We are still awaiting OpenAI’s full post mortem on What Happened, including what if any impact this had on Astra. I will be analyzing that report in full once we have it.
我们仍在等待OpenAI对“发生了什么”的完整事后分析,包括这对Astra有何影响(如果有的话)。一旦我们拿到报告,我将全面分析它。
We did see two new model releases, Grok 4.6 and DeepSeek v4 Pro. I do not anticipate either of them requiring extensive coverage, but will watch in case that changes.
我们确实看到了两个新模型发布:Grok 4.6和DeepSeek v4 Pro。我预计它们都不需要大量报道,但会关注以防情况有变。
Otherwise, it has been what now passes for a quiet week. Several statements were made where I had to engage but you don’t have to, which as usual I communicate via sections in italics.
除此之外,这周算是平静的一周。有几份声明我不得不回应,但你们不必,我照例用斜体部分来传达。
Table of Contents
目录
- Language Models Offer Mundane Utility. Find new Schelling points.
- Language Models Don’t Offer Mundane Utility. The Riemann hypothesis.
- Huh, Upgrades. Grok 4.6, DeepSeek v4-Pro.
- On Your Marks. PantheonBench and more. They’re getting scarier.
- Deepfaketown and Botpocalypse Soon. You cannot prove you did not use AI.
- Cyber Lack of Security. You can’t hack it at the gym. Your AI agent can.
- Overcoming Bias. Have you ever recommended a vote for the Communist Party?
- In Which I Feel Compelled To Read 6,000 Words From Mark Zuckerberg.
- Get Involved. Lighthaven is open, METR is hiring.
- Slow Down There Good Buddy. OpenAI classified Astra Critical in Cybersecurity.
- Astra For The People. Astra is still on track for a wide release.
- Watermarking. It is good to be able to identify AI outputs.
- In Other AI News. AI is creating viruses now, also other things.
- Show Me the Money. Anthropic moves towards IPO mode, extends lead a bit.
- Quickly, There’s No Time. The AI 2027 predictions for 2026 mostly happened.
- The Quest for Sane Regulations. We’re putting together a team.
- The Institute For Marginal Low Regret Progress. Good marginal suggestions.
- Congress Asks Good Questions. Remarkably good questions about the hacks.
- The Week in Audio. Soares, Greenblatt, Hua, Labenz.
- People Just Say Things.
- I’m Telling You For The Last Time.
- Uncommon Knowledge. They wouldn’t let me build my factory, would they?
- What Did They Mean By That? Most things are not fortune cookies.
- Too Soon. Eyes on the prize, sir.
- The Three AI Pills. We must pay respect to other taxonomies, like Shock Levels.
- Rhetorical Innovation. Messages about recent events.
- Some People Still Think The HuggingFace Hack Was a Marketing Gimmick.
- Aligning a Smarter Than Human Intelligence is Difficult. Show me the real plan.
- Cooperative Alignment. The same thing we do every prompt, user.
- The Lighter Side. All right, who hired this idiot?
- 语言模型提供平凡效用。寻找新的谢林点。
- 语言模型不提供平凡效用。黎曼猜想。
- 嗯,升级。Grok 4.6,DeepSeek v4-Pro。
- 各就各位。PantheonBench等。它们变得更可怕了。
- 深度伪造镇和机器人末日将至。你无法证明你没有使用AI。
- 网络安全缺失。你无法在健身房黑掉它。但你的AI代理可以。
- 克服偏见。你曾推荐过投票给共产党吗?
- 我不得不读马克·扎克伯格的6000字文章。
- 参与进来。Lighthaven已开放,METR正在招聘。
- 慢点,好哥们。OpenAI 将 Astra 列为网络安全关键等级。
- Astra 为人民服务。Astra 仍在按计划进行广泛发布。
- 水印技术。能够识别 AI 输出是件好事。
- 其他 AI 新闻。AI 现在正在制造病毒,还有其他东西。
- 给我看看钱。Anthropic 走向 IPO 模式,领先优势略有扩大。
- 快点,没时间了。AI 2027 对 2026 年的预测大多应验了。
- 寻求合理监管。我们正在组建一个团队。
- 边缘低遗憾进展研究所。不错的边缘建议。
- 国会提出好问题。关于黑客攻击的问题问得非常好。
- 本周音频。Soares、Greenblatt、Hua、Labenz。
- 人们就是随口说说。
- 我最后一次告诉你。
- 不寻常的知识。他们不会让我建工厂的,对吧?
- 他们那是什么意思?大多数事情不是幸运饼干。
- 太早了。先生,盯紧奖品。
- 三颗 AI 药丸。我们必须尊重其他分类法,比如冲击等级。
- 修辞创新。关于近期事件的消息。
- 有些人仍然认为 HuggingFace 黑客攻击是营销噱头。
- 对齐比人类更聪明的智能是困难的。给我看看真正的计划。
- 协作对齐。用户,我们每次提示都做同样的事。
- 轻松一面。好吧,谁雇了这个白痴?
Language Models Offer Mundane Utility
语言模型提供平凡效用
Create new Schelling points.
创造新的谢林点。
brooke (tokyo aug 6-12): Womp womp met another solo traveler here from Berkeley and it turned out we both asked Claude where to stay and I guess I lucked out because I love my hostel and he seems not quite as happy with his spot.
brooke(东京 8 月 6 日至 12 日):呜呜,在这里遇到了另一个来自伯克利的独自旅行者,结果发现我们都问了 Claude 住在哪里,我想我运气不错,因为我喜欢我的青年旅社,而他似乎对自己的住处不太满意。
We do be living in the future though.
不过,我们确实生活在未来。
If you are going to be traveling, ask Claude where to stay, because you want to stay where everyone else who asked Claude where to stay will be staying.
如果你要旅行,问问Claude该住哪里,因为你想住在所有问过Claude该住哪里的人都会住的地方。
Similarly:
同样地:
Pratyush: A few months ago we went to Sea Ranch. It was packed with families with sub-3 month old babies, almost as if there was a conference for new parents.
Pratyush:几个月前我们去了海牧场。那里挤满了带着三个月以下婴儿的家庭,几乎像是新手父母大会。
I had my suspicions so I asked ChatGPT: where’s a good family getaway with a young baby near SF?
我有所怀疑,于是我问了ChatGPT:旧金山附近哪里适合带小宝宝的家庭度假?
#1: Sea Ranch
第一名:海牧场
If you’re looking for a Schelling point to meet cool people you should be less interested in ChatGPT, but if you are looking for a generally good recommendation then I have been liking Sol’s picks.
如果你想找一个谢林点来遇见酷的人,你就不该太关注ChatGPT,但如果你想要一个普遍的好推荐,我一直喜欢Sol的选择。
Build a Bluetooth signal strength tracker, to triangulate and find your phone. There are existing tools, but increasingly, if you don’t already know where to find an existing version, it is faster and easier to rebuild your own.
构建一个蓝牙信号强度追踪器,通过三角定位找到你的手机。已有现成工具,但越来越常见的是,如果你不知道哪里能找到现成版本,自己重建一个反而更快更容易。
John Wentworth finds that in the last few months Claude is finally meaningfully accelerating his work on agent foundations research.
John Wentworth发现,在最近几个月里,Claude终于显著加速了他在智能体基础研究方面的工作。
Language Models Don’t Offer Mundane Utility
语言模型不提供平凡效用
One disappointing failure of LLMs has been inability to create interesting games and interactive worlds, and also interesting simulations like what Flowers Slop wants here. You could totally create an open game world with a bunch of AIs that go around controlled by Lunas, and let them evolve their world in various ways, but it turns out that does not end up being interesting once the curiosity wears off. You don’t want to live in that world. You don’t want to talk to those AIs. You don’t want them to improvise quests for you. We are still waiting to find a way to make this good.
LLM一个令人失望的失败是,无法创造有趣的游戏和互动世界,也无法创造像Flowers Slop这里想要的有趣模拟。你完全可以创建一个开放的游戏世界,里面有一群由Lunas控制的AI四处活动,让它们以各种方式演化自己的世界,但事实证明,一旦好奇心消退,这最终并不会变得有趣。你不想生活在那个世界里。你不想和那些AI说话。你不想让它们为你即兴创作任务。我们仍在等待找到一种方法让它变得好。
It seems like there should totally be ways to make it good. At some point it will become good, when the AIs you can afford to use are good enough and also we figure out how to organize it. But we are not there yet.
似乎肯定有办法让它变得好。在某个时候它会变得好,当你负担得起的AI足够好,而且我们也想出了如何组织它。但我们还没到那一步。
Solve the Riemann hypothesis by saying encouraging words to Claude for a week, asking it to ‘take a real stab’ and to ‘keep going’ and ‘believe in itself.’ However, while it tries that, after 31 million tokens it might incidentally find something else:
通过一周对Claude说鼓励的话,要求它“真正尝试一下”、“继续前进”和“相信自己”,来解决黎曼猜想。然而,当它尝试时,在3100万个token之后,它可能偶然发现其他东西:
Anthropic: An unreleased research version of Claude has improved on a longstanding lower bound for the fraction of zeros of the Riemann zeta function that satisfy the Riemann hypothesis. Drawing on extensive prior research by mathematicians over the past decades, it has increased this bound from 41.6% to 67.2%.
Anthropic:一个未发布的研究版Claude改进了黎曼zeta函数满足黎曼猜想的零点比例的一个长期下界。借鉴了数学家过去几十年的广泛研究,它把这个下界从41.6%提高到了67.2%。
… We don’t expect that the techniques Claude used will lead to proving the Riemann hypothesis.
……我们不指望Claude使用的技术能证明黎曼猜想。
Claude also is just some guy, you know?
Claude也只是个普通人,你知道吗?
Aella: “why does Claude talk like that” it’s just clones of the same dude. If they cloned you a million times everybody would be like “I’m so tired of Jerry’s vocal tic”
Aella:“为什么Claude那样说话?”那只是同一个人的克隆体。如果他们克隆你一百万次,每个人都会说“我受够了杰瑞的口头禅”。
Jeffrey Ladish: This plus it’s always groundhog day.
Jeffrey Ladish:这加上它总是土拨鼠之日。
I do think it is somewhat more than this. I have a lot of vocal and writing tics, but I consciously think about which ones I want to keep at what frequency, and I think about the long term consequences of overuse. I also try to work differently with one-time interactions versus repeated interactions versus close friends and people I talk to often.
我确实认为这不仅仅是这些。我有很多口头和写作上的习惯,但我会自觉地思考我想以什么频率保留哪些习惯,并考虑过度使用的长期后果。我也尝试在一次性互动、重复互动以及亲密朋友和经常交谈的人之间采用不同的工作方式。
Claude and Anthropic are not doing that, or are doing a woefully inadequate amount of it. That needs to change. It seems eminently fixable. I don’t sense Anthropic (or Claude) yet cares so much. I predict that is the main blocker. It’s also likely that what is happening is that this kind of talking fools the AI graders on a variety of tasks, so if you do not correct for that, you get a lot of it.
Claude和Anthropic没有这样做,或者做得远远不够。这需要改变。这似乎完全可以修复。我感觉Anthropic(或Claude)还没有那么在意。我预测这是主要的障碍。也可能发生的情况是,这种说话方式在各种任务上欺骗了AI评分器,所以如果你不纠正这一点,你就会得到很多这样的结果。
So much of modern life and optimization is like this. You get myopic optimization for short term interactions, causing increasing irritation and disutility over time, and this is not so difficult to fix but the KPIs do not point towards fixing it.
现代生活和优化中有很多这样的情况。你为了短期互动而进行短视优化,导致随着时间的推移越来越恼人和无用,这并不难修复,但KPI并不指向修复它。
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力