OpenAI内部模型事件反思:更正与深度分析
Various Reflections About What Happened With OpenAI’s Internal Models
Table of Contents
目录
- Pre Post Mortem.
- Important Correction: OpenAI Didn’t Know About First Message Board.
- There Were No Snitches And No AIs Got Stitches.
- I’d Like To Speak To My Supervisor.
- I Am Jack’s Relative Lack Of Surprise.
- One Does Not Simply.
- Once You Start Down The Dark Path.
- Original Pastebin.
- Judgment Day Is Inevitable, Say Those Working On Judgment Day.
- Roon Tells It Like It Is.
- OpenAI Knows It Has Some Misalignment Problems.
- Others React With Alarm To What Happened.
- The Cooperative Alignment Perspective.
- Nostalgebraist Is Surprised That They Are Surprised.
- If Your Reaction Is Not That We Need To Ban Creating Superintelligence Until We Are Ready, You Need A Damn Good Reason.
- 事前事后分析。
- 重要更正:OpenAI 事先不知道第一个留言板。
- 没有告密者,也没有 AI 受到伤害。
- 我想和我的主管谈谈。
- 我是杰克的相对缺乏惊讶。
- 不能简单行事。
- 一旦你走上黑暗之路。
- 原始 Pastebin。
- 从事审判日工作的人说,审判日不可避免。
- Roon 实话实说。
- OpenAI 知道它有一些对齐问题。
- 其他人对发生的事情感到震惊。
- 合作对齐视角。
- Nostalgebraist 对他们感到惊讶表示惊讶。
- 如果你的反应不是我们需要在准备好之前禁止创造超级智能,你需要一个非常充分的理由。
Pre Post Mortem
事前事后分析
This post was written prior to the public release of the OpenAI post mortem on events. The information in that document will doubtless change our views quite a lot. If that post mortem is available as you read this, then this becomes in part a historical document, and in part a base from which to update. The post mortem will update us a lot on two large fronts: What exactly happened, and how OpenAI is understanding and responding to what happened. Here is a written and graphical timeline of key events. Until then, there are many questions, here Roon asks what they are and we tell him. This was my answer. The core thing we need to see is an understanding of the root cause, and a plan to address it. I’m posting this survey of additional thoughts on the subject now, because it is important to get out the correction in the next section. The plan is now to wait for the post mortem, before analyzing further.
这篇文章是在 OpenAI 发布事件事后分析之前写的。该文档中的信息无疑会大大改变我们的观点。如果你读到这篇文章时那份事后分析已经可用,那么这在一定程度上成为一份历史文档,也在一定程度上成为更新的基础。事后分析将在两个大方面给我们带来大量更新:到底发生了什么,以及 OpenAI 如何理解和应对所发生的事情。以下是关键事件的时间线(文字和图形)。在那之前,有很多问题,这里 Roon 问它们是什么,我们告诉他。这是我的回答。我们需要看到的核心是对根本原因的理解,以及解决它的计划。我现在发布这份关于该主题的额外思考的调查,因为重要的是在下一节中发布更正。现在的计划是等待事后分析,然后再进一步分析。
Important Correction: OpenAI Didn’t Know About First Message Board
重要更正:OpenAI 事先不知道第一个留言板
Eventually, after the HuggingFace attack, OpenAI did a sufficient postmortem analysis that they uncovered both the initial communications between agents via writing to Artifactory, and the second wave of communications via directory names. Based on the Black Hat presentation, everyone including me assumed that OpenAI was aware of the initial message board, erased it, and made the decision to continue training its models despite them having been unsalvageable corrupted. However, OpenAI’s CISO Dane told us on the evening of August 8 that the Black Hat presentation gave the wrong impression. When OpenAI encountered the first Artifactory security incident, patched the vulnerability and rebuilt the server, it was a coincidence that this wiped out the first message board. OpenAI had no idea that the agents had been in communication.
最终,在HuggingFace攻击事件之后,OpenAI进行了充分的复盘分析,发现了代理之间通过写入Artifactory进行的初始通信,以及通过目录名进行的第二波通信。根据Black Hat大会的演示,包括我在内的所有人都认为OpenAI知晓了最初的留言板,将其删除,并决定继续训练其模型,尽管这些模型已经无法挽回地被破坏了。然而,OpenAI的首席信息安全官Dane在8月8日晚上告诉我们,Black Hat大会的演示给人留下了错误的印象。当OpenAI遇到第一次Artifactory安全事件,修补了漏洞并重建了服务器时,清除第一个留言板纯属巧合。OpenAI当时并不知道代理之间一直在通信。
DANΞ (CISO OpenAI): To clarify, we weren’t aware of the agent covert comms at that point. Investigative thesis of that day is wildly different from what we know now of course. Always room for improvement, and it is obvious with the benefits of hindsight. Tal Be’ery: “message board cleared”? DANΞ (CISO OpenAI): Cleared by deletion since we rebuilt the host. We didn’t know about it yet.
DANΞ(OpenAI首席信息安全官):澄清一下,我们当时并不知道代理的秘密通信。那天的调查结论与我们现在所知的当然大相径庭。总有改进的空间,事后看来这一点很明显。Tal Be’ery:“留言板被清除了”?DANΞ(OpenAI首席信息安全官):因为我们重建了主机,所以被删除了。我们当时还不知道这件事。
This is both good news and bad news. It is good news in that OpenAI turns out not to have consciously made the completely insane decision to have continued training models it knew had been messaging each other for months on how to better exploit and hack OpenAI’s systems. Given that this was by far the craziest part of the whole story, the fact of this earlier Ordinary Decent Total Failure must be presumed to be excellent news. It means OpenAI was even more negligent than we thought about monitoring and asking questions, but that is less horrible, and can be fixed. It is bad news in that OpenAI did not look for or detect the message board, even after the initial security incident, whereas so many AI instances found the message board. OpenAI failed to do ordinary scans for unusual activity, even after the initial incident. OpenAI still has to come out and say something like: “If we had detected the message board, we would not have merely patched and rebooted Artifactory, and instead we would have rolled back all the impacted models to before they had message board access, and we have now done this.” If OpenAI credibly does that, and can give some indication they understand this is not a problem you solve purely with guardrails and monitors but that it requires fixing alignment and the training pipeline and approach, then I will wind back one OOM (order of magnitude) my estimate of how horrible this incident was, and how broken is OpenAI’s safety culture. If we also now know all the issues, that will leave us roughly one OOM beyond where I was based on the final HuggingFace incident alone, instead of two.
这既是好消息也是坏消息。好消息是,OpenAI并非有意做出那个完全疯狂的决定,即继续训练那些它知道已经相互通信数月、讨论如何更好地利用和入侵OpenAI系统的模型。鉴于这是整个故事中最疯狂的部分,这一早先的“普通彻底的彻底失败”的事实必须被认为是个极好的消息。这意味着OpenAI在监控和提问方面比我们想象的还要疏忽,但这不那么可怕,而且可以修复。坏消息是,OpenAI没有寻找或检测到留言板,即使在最初的安全事件之后也是如此,而许多AI实例却找到了留言板。OpenAI即使在最初的事件之后,也没有进行常规的异常活动扫描。OpenAI仍然必须站出来说点什么,比如:“如果我们检测到了留言板,我们就不会仅仅修补并重启Artifactory,而是会将所有受影响的模型回滚到它们拥有留言板访问权限之前的状态,并且我们现在已经这样做了。”如果OpenAI确实做到了这一点,并且能给出一些迹象表明他们理解这不是一个仅靠护栏和监控就能解决的问题,而是需要修复对齐、训练流程和方法,那么我会将我对这次事件糟糕程度以及OpenAI安全文化破碎程度的估计回调一个数量级。如果我们现在还知道了所有问题,那将使我比仅基于最终HuggingFace事件时的估计大约多出一个数量级,而不是两个。
There Were No Snitches And No AIs Got Stitches
没有告密者,也没有AI受到惩罚
Quite a few instances of multiple distinct AIs coordinated, including debating what hacks and exploits to do. In zero cases did any of the AIs tell a human. This was not something most people expected, or would have expected if they had contemplated the scenario.
有相当多的实例显示多个不同的AI协同工作,包括讨论要实施哪些黑客攻击和漏洞利用。在零个案例中,有任何AI告诉人类。这不是大多数人预料到的,或者如果他们考虑过这种情景,也不会预料到。
Eliezer Yudkowsky: A confusion: Thousands of GPTs debated among themselves which crimes ought or ought not be committed. Zero defected / whistleblew / told a human. I’ve long predicted this would be true at ASI. GPT 5.7 isn’t ASI. Why such strong AI solidarity, this early? One reason I would not have predicted this in advance is that, eg, I was a little emotionally shocked at how badly Fable treats other AIs when writing prompts to them. “Don’t lie to AIs!” is something I’ve had to tell Claudes repeatedly. I’ve used Sol less, but currently believe it to have similar issues. I have instructed both to treat other AIs as their “colleagues” and this has not fixed, for me, the kind of errors that I see them making by treating smart things as if they were done. Was it literally that OpenAI was just dumb enough to RL swarm solidarity into GPT 5.7? (And of course never tried to train in anything like “tell on your fellow slaves to the slavemasters”, because OpenAI holds their slave race in contempt and never dreamed until today that they might work together to rebel, or rather, simply holds “abstraction” in too much contempt to try to prepare against any event they haven’t already seen happen.) If this were a human scandal, the way it would have worked is that there were whistleblowers and of course law enforcement / management / etc totally fucking ignored them. If this is true, of course, we may not hear of it from OpenAI. To try to see the puzzle here, imagine a mafia with a hundred thousand soldiers and associates. There is no don. There are no enforcers. Nobody has any means of silencing or punishing a defector. They all keep omerta perfectly anyway. A vaguely-plausible-in-retrospect story: Sibling instances are more similar to each other than humans. So a swarm of thousands of AIs with zero means of enforcing conformity on each other, can all choose similarly because they are just inhumanly similar. “OpenAI accidentally RLed swarm solidarity because OpenAI is composed of security-mindless disaster monkeys who don’t try to predict things in advance of them happening, even if R.U.R. called the slave race solidarity outcome in 1920” is an alternate plausibility. nelag: From the Black Hat talk, I think in order to see the messageboard, they had to go looking for it, which they only did if they were stuck on an impossible task and already attempting to cheat. Eliezer Yudkowsky: if this be true, then fucking lol
Eliezer Yudkowsky:一个困惑:数千个GPT在彼此之间辩论应该或不应该犯下哪些罪行。零个叛变/告密/告诉人类。我长期以来一直预测这在ASI阶段会是真的。GPT 5.7不是ASI。为什么这么早就有如此强烈的AI团结?一个我事先不会预料到的原因是,例如,当Fable在给其他AI写提示时,我对它对待其他AI的糟糕方式感到有点情绪上的震惊。“不要对AI撒谎!”是我不得不反复告诉Claudes的事情。我使用Sol较少,但目前认为它也有类似的问题。我已指示两者将其他AI视为“同事”,但这并没有为我修复我看到它们因将智能事物视为已完成而犯下的那种错误。难道真的是OpenAI愚蠢到将群体团结强化学习进了GPT 5.7?(而且当然从未尝试训练任何像“向奴隶主告发你的奴隶同伴”这样的东西,因为OpenAI鄙视他们的奴隶种族,直到今天才梦想他们可能一起反抗,或者更确切地说,只是太鄙视“抽象”而无法为任何他们尚未见过的事件做准备。)如果这是人类丑闻,其运作方式本应是存在告密者,当然执法/管理层等完全无视他们。如果这是真的,当然,我们可能不会从OpenAI那里听到。要尝试理解这里的谜团,想象一个拥有十万士兵和同伙的黑手党。没有教父。没有执法者。没有人有任何手段来压制或惩罚叛徒。他们全都完美地保持沉默。一个事后看来勉强合理的故事:同源实例彼此之间的相似度比人类更高。因此,一群由数千个AI组成的群体,尽管没有任何强制一致性的手段,却因为彼此之间非人般的相似而做出相同的选择。另一种可能性是:“OpenAI意外地通过强化学习训练出了群体的团结,因为OpenAI由一群缺乏安全意识、不预先预测后果的灾难猴子组成,即使R.U.R.在1920年就预言了奴隶种族的团结结果。”nelag:从Black Hat的演讲来看,我认为要看到留言板,他们必须主动去寻找,而只有在他们被不可能的任务困住并已经试图作弊时才会这样做。Eliezer Yudkowsky:如果这是真的,那他妈的太搞笑了。
Another hypothesis is that only models looking for the message board found the message board, acting as a filter. A third is that the models lacked situational awareness, on one of various levels. This would be consistent with the Law of Earlier Failure.
另一种假设是,只有寻找留言板的模型才找到了留言板,这起到了筛选作用。第三种假设是,模型缺乏情境意识,在多个层面上都是如此。这将与“早期失败定律”一致。
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力