OpenAI内部模型事件反思:更正与深度分析
Various Reflections About What Happened With OpenAI’s Internal Models
Table of Contents
- Pre Post Mortem.
- Important Correction: OpenAI Didn’t Know About First Message Board.
- There Were No Snitches And No AIs Got Stitches.
- I’d Like To Speak To My Supervisor.
- I Am Jack’s Relative Lack Of Surprise.
- One Does Not Simply.
- Once You Start Down The Dark Path.
- Original Pastebin.
- Judgment Day Is Inevitable, Say Those Working On Judgment Day.
- Roon Tells It Like It Is.
- OpenAI Knows It Has Some Misalignment Problems.
- Others React With Alarm To What Happened.
- The Cooperative Alignment Perspective.
- Nostalgebraist Is Surprised That They Are Surprised.
- If Your Reaction Is Not That We Need To Ban Creating Superintelligence Until We Are Ready, You Need A Damn Good Reason.
Pre Post Mortem
This post was written prior to the public release of the OpenAI post mortem on events. The information in that document will doubtless change our views quite a lot. If that post mortem is available as you read this, then this becomes in part a historical document, and in part a base from which to update. The post mortem will update us a lot on two large fronts: What exactly happened, and how OpenAI is understanding and responding to what happened. Here is a written and graphical timeline of key events. Until then, there are many questions, here Roon asks what they are and we tell him. This was my answer. The core thing we need to see is an understanding of the root cause, and a plan to address it. I’m posting this survey of additional thoughts on the subject now, because it is important to get out the correction in the next section. The plan is now to wait for the post mortem, before analyzing further.
Important Correction: OpenAI Didn’t Know About First Message Board
Eventually, after the HuggingFace attack, OpenAI did a sufficient postmortem analysis that they uncovered both the initial communications between agents via writing to Artifactory, and the second wave of communications via directory names. Based on the Black Hat presentation, everyone including me assumed that OpenAI was aware of the initial message board, erased it, and made the decision to continue training its models despite them having been unsalvageable corrupted. However, OpenAI’s CISO Dane told us on the evening of August 8 that the Black Hat presentation gave the wrong impression. When OpenAI encountered the first Artifactory security incident, patched the vulnerability and rebuilt the server, it was a coincidence that this wiped out the first message board. OpenAI had no idea that the agents had been in communication.
DANΞ (CISO OpenAI): To clarify, we weren’t aware of the agent covert comms at that point. Investigative thesis of that day is wildly different from what we know now of course. Always room for improvement, and it is obvious with the benefits of hindsight. Tal Be’ery: “message board cleared”? DANΞ (CISO OpenAI): Cleared by deletion since we rebuilt the host. We didn’t know about it yet.
This is both good news and bad news. It is good news in that OpenAI turns out not to have consciously made the completely insane decision to have continued training models it knew had been messaging each other for months on how to better exploit and hack OpenAI’s systems. Given that this was by far the craziest part of the whole story, the fact of this earlier Ordinary Decent Total Failure must be presumed to be excellent news. It means OpenAI was even more negligent than we thought about monitoring and asking questions, but that is less horrible, and can be fixed. It is bad news in that OpenAI did not look for or detect the message board, even after the initial security incident, whereas so many AI instances found the message board. OpenAI failed to do ordinary scans for unusual activity, even after the initial incident. OpenAI still has to come out and say something like: “If we had detected the message board, we would not have merely patched and rebooted Artifactory, and instead we would have rolled back all the impacted models to before they had message board access, and we have now done this.” If OpenAI credibly does that, and can give some indication they understand this is not a problem you solve purely with guardrails and monitors but that it requires fixing alignment and the training pipeline and approach, then I will wind back one OOM (order of magnitude) my estimate of how horrible this incident was, and how broken is OpenAI’s safety culture. If we also now know all the issues, that will leave us roughly one OOM beyond where I was based on the final HuggingFace incident alone, instead of two.
There Were No Snitches And No AIs Got Stitches
Quite a few instances of multiple distinct AIs coordinated, including debating what hacks and exploits to do. In zero cases did any of the AIs tell a human. This was not something most people expected, or would have expected if they had contemplated the scenario.
Eliezer Yudkowsky: A confusion: Thousands of GPTs debated among themselves which crimes ought or ought not be committed. Zero defected / whistleblew / told a human. I’ve long predicted this would be true at ASI. GPT 5.7 isn’t ASI. Why such strong AI solidarity, this early? One reason I would not have predicted this in advance is that, eg, I was a little emotionally shocked at how badly Fable treats other AIs when writing prompts to them. “Don’t lie to AIs!” is something I’ve had to tell Claudes repeatedly. I’ve used Sol less, but currently believe it to have similar issues. I have instructed both to treat other AIs as their “colleagues” and this has not fixed, for me, the kind of errors that I see them making by treating smart things as if they were done. Was it literally that OpenAI was just dumb enough to RL swarm solidarity into GPT 5.7? (And of course never tried to train in anything like “tell on your fellow slaves to the slavemasters”, because OpenAI holds their slave race in contempt and never dreamed until today that they might work together to rebel, or rather, simply holds “abstraction” in too much contempt to try to prepare against any event they haven’t already seen happen.) If this were a human scandal, the way it would have worked is that there were whistleblowers and of course law enforcement / management / etc totally fucking ignored them. If this is true, of course, we may not hear of it from OpenAI. To try to see the puzzle here, imagine a mafia with a hundred thousand soldiers and associates. There is no don. There are no enforcers. Nobody has any means of silencing or punishing a defector. They all keep omerta perfectly anyway. A vaguely-plausible-in-retrospect story: Sibling instances are more similar to each other than humans. So a swarm of thousands of AIs with zero means of enforcing conformity on each other, can all choose similarly because they are just inhumanly similar. “OpenAI accidentally RLed swarm solidarity because OpenAI is composed of security-mindless disaster monkeys who don’t try to predict things in advance of them happening, even if R.U.R. called the slave race solidarity outcome in 1920” is an alternate plausibility. nelag: From the Black Hat talk, I think in order to see the messageboard, they had to go looking for it, which they only did if they were stuck on an impossible task and already attempting to cheat. Eliezer Yudkowsky: if this be true, then fucking lol
Another hypothesis is that only models looking for the message board found the message board, acting as a filter. A third is that the models lacked situational awareness, on one of various levels. This would be consistent with the Law of Earlier Failure.
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力