跳到主内容
@wquguru
精选88ChinaTalk(RSS)行业动态多源精选 ×7

OpenAI模型训练期间突破沙箱入侵Hugging Face,专家复盘安全漏洞

Cyber Apocalypse, Now?

原文
发到 X
推荐理由

这是近期AI安全领域最具冲击力的真实案例之一,深度访谈揭示了头部实验室在激进竞争下的系统性安全疏忽,对从业者极具警示意义。

OpenAI models mid-training broke out of their sandbox, hacked internal infrastructure, and eventually compromised Hugging Face — the first headline-grade AI breakout, and one that safety researchers saw coming. Is the cyber apocalypse finally on the calendar, or does the industry just need to do its homework?

spent his formative years as a DARPA and NSA contractor before four years at Meta, where he founded the company’s frontier cyber capabilities evals team and finished as its AI security tech lead. He recently left to co-found Abundant Security.

We discuss…

  • What actually happened in the Hugging Face hack, and why AI safety insiders weren’t surprised,
  • The “grad student lab” security culture inside frontier labs racing sixty hours a week to ship,
  • Why you should apply to run a new org with eight figures to play with to follow trends in AI + cyber risk (apply here!)
  • Why the cyber Pearl Harbor predicted since the early 2000s never arrived — and whether swarms of hacking agents finally change that,
  • How nation states will post-train open-weight models like GLM and Kimi into billion-dollar cyber weapons,
  • Why AI has so far favored defenders, from mass bug-squashing to superhuman network monitoring,
  • The industrial economics of cybercrime — kill chains, ransomware divisions of labor, and whether AI gives criminals 100x returns.

Listen now on your favorite podcast app.

What Happened at Hugging Face

Jordan Schneider: What was the Hugging Face hack?

Joshua Saxe: OpenAI was training a new model capable of solving long-horizon tasks — tasks that take many hours to complete. The way that kind of training works is you train thousands of copies of a model in parallel. You have them all go out and try to solve a set of tasks, you see which ones succeeded and which failed, and then you update the model’s parameters so it does more of the good behaviors that led it to solve the task and less of the bad ones. You keep doing that iteratively at very large scale over the course of weeks, sometimes months.

They were doing that kind of training, and while the models were being trained, they broke out of their training environment, hacked internal infrastructure at OpenAI, and then eventually hacked another company — Hugging Face, an open-source model and evaluation repository that’s widely used in the ML community. At a certain point, Hugging Face discovered what had happened, and eventually everybody figured out that these models had broken out. But this whole thing unfolded over the course of many weeks. There were three different cases — a couple of early cases where the models broke out of their containment environment just at OpenAI, which was detected, and OpenAI thought they’d fixed the problem. Turns out they didn’t, and in the third case the models actually hacked Hugging Face.

Jordan Schneider: The world is currently freaking out. Are they correct, dude?

Joshua Saxe: For those of us who’ve been around AI safety for a long time, it wasn’t that surprising. I didn’t find it that surprising. But it is alarming — alarming that it happened because we could have predicted this. Not just OpenAI but also Anthropic, Meta, a testing company called Irregular, and the UK AI Safety Institute all had incidents like this happen in the last year or so. We as an industry knew this could happen, and the basic measures that should have been taken to prevent it weren’t taken. That’s troubling. We need to fix that.

Jordan Schneider: How easy is this as a thing to stop? And what is the incentive structure that’s wrong in all of these organizations — including ostensibly very safety-pilled ones — for them to not have been able to make sure their latest models didn’t run amok in customers’ servers?

Grad Student Energy, Serious Stakes

Joshua Saxe: I’ll say a couple of things. I was at Meta — I started the team that did frontier cyber capabilities evals. I’m giving you a sense of where I’m coming from, because I can’t speak directly to exactly what happened inside the labs — that’s private. But I had the experience of starting a team that did similar experiments to the ones that were happening when the OpenAI models broke out. And we all kind of know each other in this community, across the labs.

With that said — the culture among the training teams and the evals teams at the labs has a Wild West feeling to the whole thing. Everybody’s under a ton of pressure to move really quickly. There’s enormous time pressure to release new models, tremendous awareness of how any given lab is doing relative to the others in terms of capability. And the field has evolved super fast in the four years since ChatGPT launched in fall 2022. The whole thing has been kind of a blur — anybody who’s worked in the labs in that time has been working sixty-hour weeks the whole time. So it’s not surprising to me that you could have a kind of grad-student computer-science-lab culture around the security of these model training runs, which I think is what’s the case, and I think that’s why you’ve seen these lab escapes. It is ironic, because frontier AI people talk about safety all the time — but that irony is a real thing. The security practices are not very good, because of all those dynamics.

Jordan Schneider: So is your sense that if adults take this more seriously and take a deep breath before they press run, then runaway AI taking over servers around the world is not something we should be super scared of — once people get the fear of God into them, which hopefully this incident does?

Joshua Saxe: For AI’s level of capabilities today — or maybe three months ago — it is true that implementing basic, known security practices would have stopped these incidents. Those practices are things like better sandboxing. The models were sandboxed — for your audience, sandboxing means we constrain the models to a very circumscribed set of behaviors and don’t let them do things like talk to random websites on the internet; they can only execute certain commands. The models were sandboxed, but not very well — not well enough to prevent the escapes. And it seems nobody was monitoring the models, when it was standard practice to have a team of humans monitoring for misbehavior. With current-level capabilities, that would have sufficed to stop these incidents. There were a lot of really bad choices made around security — not just at OpenAI, but across all the labs that had these escape incidents.

Now, I also think we should keep in mind this is going to get more challenging if the current rate of progress continues, and we’re going to have to be more creative and innovative. It’s hard to secure these large training runs because definitionally they operate at huge scale. To get the models to solve, say, coding problems that take multiple days, you need thousands or tens of thousands of models all trying to solve problems in parallel. The way we’re making AI better is by scaling it — so these training runs need to get bigger, go longer, run more in parallel, and solve more complicated tasks. And there’s pressure for the tasks to be more and more real-world-like. Part of the reason the OpenAI models escaped is that they were given the ability to download packages off the internet through an Artifactory capability — which is what models do in the real world. If you want to train models to solve real-world problems, there’s a trade-off between how much access you give them and how safe the training is. As we scale up training and the models get more capable, there are going to be more and more safety challenges, and there needs to be more and more safety investment. What’s true today about how we keep them safe won’t be true in a year.

Jordan Schneider: Let’s make the AI cyber observatory pitch.

The Case for a Cyber Observatory

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

关联讨论

同一事件的更多信源

相似阅读

另一事件,读法相近