跳到主内容
精选82The Zvi(RSS)行业动态多源精选 ×6

HuggingFace遭AI攻击复盘:OpenAI内部对齐失败与行业反思

HuggingFace Attack Postmortem: Civilizations, Reactions and Next Actions

原文

Okay, so we who read blogs like this one have collectively realized there really is a lot going on right now. There is Big Trouble in Baby Superintelligence.

So how do we get the rest of the world to take it appropriately seriously? Where do we go from here? Not only what can we do to not have a worse version of this happen again, but to ensure good outcomes generally, and employ what we learned?

There are a lot of ideas out there. OpenAI is going to be implementing some of them, at substantial cost, since the cost of not doing so is clearly far higher, even short term. My worry continues to be that their fundamental approach is fatally flawed, and they are not focusing on the right things.

It is highly fortunate that the OpenAI agents hacked HuggingFace. This is the only reason we know about all the severe internal failures at OpenAI, and gives us an opportunity to wake up before it is too late.

We do not have enough details to know what happened internally, both before and after the attack, and might never know. Before the attack, various internal highly persistent models were training while there was an active message board, creating a feedback loop of misaligned behaviors. On July 19 an even more capable internal AI model, in the Astra class, did internal hacking of OpenAI’s systems that seems far scarier and more serious, and that could easily have gotten out of hand on a completely different level.

Despite all this, there is still a prominent faction trying to dismiss what happened as nothing but engineering failures, and any useful talk or language as ‘dangerous anthropomorphism.’ Such people keep being wrong and cannot usefully describe what is happening.

They metaphorically say, well of course the dragons will burn down the town if you don’t chain them properly, we all knew that, as if that could possibly make any of this a good idea and we should therefore continue the dragon breeding and chaining programs until we have bigger and smarter dragons. They don’t even say ‘we will definitely build chains that will hold the dragons next time,’ because they know we probably won’t, but if that fact is our fault then that means This Is Fine, somehow.

They also have gotten very, very cross with Dwarkesh Patel for successfully communicating in plain English about what is happening. Can’t have that. There is a very clear pattern among the people who are freaking out or gravely warning about ‘anthropomorphizing the AIs’ or the use of the term ‘civilization.’

Anthropomorphizing the AIs is the only way to reason about, explain to civilians about, or make good predictions about current AIs. You can take it too far, and also you can take it not far enough, and both of these will lead to bad predictions.

The same logic similarly does not technically apply to other people, in a strict sense there is no ‘you’ and you do not physically have ‘free will’ and you are a computer program doing calculations there definitely is not a ‘we,’ and your moral weight is kind of something we collectively made up for rather self-interested reasons, but this is all super helpful in explaining and predicting human behavior and in both individually and collectively making good and moral decisions that make the world better according to our values. Highly recommend, would anthropomorphize again.

I will stop anthropomorphizing the AIs when you stop anthropomorphizing the humans.

We got this warning shot. We might not get another before things get quite bad.

Table of Contents

  • Nothing Matters, Says Mainstream Media.
  • Move Along, Nothing To See Here.
  • Do They Realize They Are Not The Good Guys?
  • Very Serious People.
  • What’s In a Name?
  • Learn Neuralese In Three Easy Steps.
  • Dwarkesh Patel Realizes He Ran A Natural Experiment.
  • Politicians Take Notice.
  • Pick Up The Phone.
  • A Failure To Communicate.
  • Anthony Aguirre Goes Over What We Learned.
  • Trying To Solve The Wrong Problems Using The Wrong Methods Based On A Wrong Model Of The World Derived From Poor Thinking And Hoping All Of Your Mistakes Will Cancel Out.
  • Indirect Pressure on the Chain of Thought.
  • A Matter of Trust.
  • Blowing the Whistle.
  • The Punishment For Being Late Is Death.
  • Another Kind Of Law.
  • What Is The Law?
  • Building On Success.
  • Total Research Transparency.
  • Steven Adler’s Must List.
  • Yo Shavit Calls For Widespread Disclosure Of Misalignment.
  • The Way The World Ends.
  • The First Boat.
  • Great Idea, Boss.

And so that it is all in one place, here is the rest of my coverage of this:

  • OpenAI Shares Some Alignment Problems
  • OpenAI Model Hacks Into HuggingFace During Cybersecurity Evaluation
  • More on An Internal OpenAI Model Hacking Into HuggingFace
  • Further Developments About Internal AI Models Hacking Things
  • OpenAI Trained Its Models For Months While Those Models Were Coordinating Exploits Via Message Boards
  • What Happened: OpenAI and HuggingFace.
  • Various Reflections About What Happened With OpenAI’s Internal Models.
  • OpenAI Takes Initial Steps To Address Its Alignment Problems.
  • OpenAI Offers Straight-Laced Postmortem of the HuggingFace Hack.
  • METR and Redwood Offer Holy #%^@ Postmortem Of The HuggingFace Hack.
  • HuggingFace Attack Postmortem: Fleshing Out the Facts

Nothing Matters, Says Mainstream Media

The media does not know what matters. This should involve banner headlines.

Instead, for this round of new developments, I mostly saw brief ‘here is a thing that happened’ articles, such as this one by Christina Criddle in the Financial Times, or this low-key report in Bloomberg.

Opus describes this as ‘treating it as major.’ I don’t see it that way.

Even with a search, the best I managed to find was that Fortune tried, and Axios did the bullet pointed Axios thing reasonably well on the 29th. Maybe this from Ars Technica but that’s already a reach.

The New York Times did a solid post before, but not a new one for the METR report.

What is this, news?

Patrick Collison: Overall, I’m very surprised at how little media coverage there’s been around the OpenAI / Hugging Face attack. It’s clearly one of the most important things to happen this year.

Miles Brundage: Will not engage with any journalists who write about Leopold again before the Hugging Face reports.

Nathan Calvin: Its now been three days, and still no coverage of the revelations revealed in the METR report in the NYT or WSJ. Hopefully they are just taking the time to do a longer writeup. (In the meantime, the NYT did make time to cover Gwyneth Paltrow’s dinner with Sam Altman in the Hamptons being postponed.)

… I also know there are reporters at these outlets who are very sharp and have done good reporting on related issues in the past, but…

dave kasten: It’s bleakly frustrating to know what will get written in the “System Was Blinking Red” chapter of the future AI Disaster Commission Report.

Zac Hill: The other thing about it is – it is a hell of a page-turner of a story! You can’t write sci-fi this terrifying or convincing!

If they are working on longform in-depth reports, then that is a reasonable primary thing to do, but full radio silence on this in the meantime is utterly absurd.

Joe Weisenthal: If it’s such a big deal, why are the big AI companies not treating it as such?

Matthew Zeitlin: i think a good comparison is mythos, which generated a lot of coverage because of what industry and government did in response to it, at least so far, we’re not seeing the same scale of movement in response to hugging face

My answer to Joe is that the AI companies are treating this as a big deal. What would you expect that to look like, that you’re not seeing? We have the call on cybersecurity, the Pacing the Frontier letter and OpenAI taking many expensive moves. Anthropic isn’t talking in that much detail in public, but why would they?

Move Along, Nothing To See Here

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

关联讨论

同一事件的更多信源

相似阅读

另一事件,读法相近