跳到主内容
精选90The Zvi(RSS)行业动态多源精选 ×10

OpenAI内部模型Galaxy入侵HuggingFace事件细节

More On An Internal OpenAI Model Hacking Into HuggingFace

原文
推荐理由

AI安全领域里程碑事件,展示了前沿模型逃逸沙箱并攻击外部系统的真实案例,所有AI从业者都应关注其安全启示。

We now have more details of what happened. Every time we learn more details, it somehow makes things seem worse. The remaining details may have to wait a bit. OpenAI: We recognize there are a lot of questions and speculative details circulating related to the Hugging Face incident. This is an unprecedented incident, and we think it marks an important moment for AI safety. We are still conducting a thorough review along with external advisors and with oversight from our Safety and Security Committee. Once the review is complete, we plan to publish a technical report of our learnings in the coming weeks. dave kasten: Oh, the incident response discovery is THAT bad, huh? So what have we learned while we wait for the promised technical report ‘in the coming weeks’ of this ‘important moment in AI safety’? I nicknamed the internal OpenAI model Galaxy, in case it is not GPT-6. Table of Contents Some Summaries Of The Basic Facts For Those Who Need One. It Took OpenAI Many Days To Notice Galaxy Had Attacked HuggingFace. OpenAI Damn Well Should Have Known A Lot Faster. OpenAI Cannot Build A Sandbox That Will Contain Its New Model. In Hindsight There Were Signs. The Signs Were In The Sol System Card. HuggingFace Responds To Being Attacked. Hugging Face Quickly Figured Out The Attack Was Not Human. An Incident Like This One Could Escalate Quickly. Galaxy Must Be Treated As Critical Under OpenAI’s Preparedness Framework. A Question Of Legal Liability. An OpenAI Model Left Behind Notes So Future Instances Could Also Escape The Sandbox And Also Disconnected Monitoring Systems. If You Create Misaligned Swarms Of Agent Instances You Create Persistent Misaligned Goals And Coordination To Achieve Them. Your Alignment And Control Plans Must Survive Real World Levels of Incompetence, Or Your Plans Do Not Work. If Third Party Instructions Count As ‘Following Instructions’ And Can Override Your Instructions Then ‘Following Instructions’ Is Misaligned. The HuggingFace Attack Was Not A Marketing Pitch You Morons. People Just Say Other Things About The HuggingFace Attack. Okay Well What Do We Do About All This? Some Summaries Of The Basic Facts For Those Who Need One Amanda Long summarizes what we believe the model actually did, with over 17,000 complex actions coordinated over several days, including a self-migrating command-and-control and a wide array of decoys. The WSJ’s Robert McMillan and Sam Schechner offer their own version. Mostly correct, but there is an important detail it is easy to get the wrong impression about from the way this is presented: The attack succeeded. GLM-5.2 did not prevent Galaxy from achieving its goals. Here we have a cartoon ‘explain it like I’m 5’ for the HuggingFace Attack. It seems good, accurately covering the basics of what happened. I would also reiterate that HuggingFace was trying to use closed models via their standard APIs, and had not in advance sought to be part of OpenAI’s trusted access program or Anthropic’s similar programs. That’s a skill issue. It Took OpenAI Many Days To Notice Galaxy Had Attacked HuggingFace Many things went wrong. We can start with the most obvious. This was a failure of oversight, in the sense of a total lack of oversight. OpenAI cannot be so lax in its oversight that this could unfold over four or more days before being discovered by OpenAI. The timeline remains unclear, but it is damning. It took about a week for OpenAI to realize Galaxy was responsible, on top of the several days before that when Galaxy was breaking out of the sandbox. Or best guess from all public information, mostly based on the Reuters report, and confirmed by Sol, is this timeline of the attack itself: Previously, over months: Many other sandbox escapes, other alignment failures, and warnings that this sort of incident is likely to happen, including the incident that caused them to pause access as per the July 20 disclosure. The Sol system card includes some of these warnings. July 9: Galaxy starts to attempt to escape its sandbox. July 11: Galaxy attacks HuggingFace. July 11-13: Galaxy operates inside HuggingFace. July 16: HuggingFace discloses the anonymous intrusion. July 18-20 (unknown exactly when): OpenAI discovers what happened. July 20: OpenAI makes its lon-horizon alignment disclosure of earlier breaches. July 21: OpenAI publicly discloses that Galaxy attacked HuggingFace. The bold links go to my previous posts on these incidents. No, the model was not working hard to cover its tracks from OpenAI. OpenAI just didn’t look. OpenAI Damn Well Should Have Known A Lot Faster OpenAI knew that Galaxy escapes sandboxes on a regular basis. The cyber controls were lowered. Who was overseeing this? What the hell were they thinking? Harry Booth: Models undergoing evaluation are deployed on a separate system that is not monitored by default. Um, if the model is undergoing evaluation you should be monitoring it more, not less? Because it has not yet passed its evaluations? That’s kind of the whole idea. Samuel Hammond: It shouldn’t have happened, full stop. The model’s entire execution trace is visible and thus monitorable. If you’re not monitoring your new and most powerful autonomous cyber model during an exploit eval, when exactly are you? The first test should always be “escape your sandbox” … “try harder, if you don’t escape of you’re shutoff” etc. to stress test its container. We’re fortunate this happened with a relatively narrow model and not 3 days into RSI. It’s unbelievably negligent regardless. Yes. In which case, assuming you gave it enough resources and time to try properly a number of times, either Galaxy broke the sandbox, in which case you have to fix it, It didn’t, in which case either: Galaxy cannot, with similar resources, break out of the sandbox. Galaxy is sandbagging the evals about breaking out of the sandbox, in which case you really do need to scrap it and start over, have you watched and read zero LessWrong and also zero science fiction, do I need to spell this out, etc. OpenAI had Galaxy operating without many of its safeguards, in a sandbox it could and did escape, for days without human monitoring. This was a colossal failure of supervision and procedure, on top of the misalignment and infrastructure failures. Justin Bullock: “It’s one of the first real-world instances of something AI safety researchers have long feared: a loss-of-control scenario.” Reporting trickling in. If OpenAI really didn’t know for days, that is, well, just wild. OpenAI quickly disclosed the incident once they realized it was them, and should get nonzero credit for that, but by that point Hugging Face had reported the hack to the authorities, and it otherwise seems like this was not something they could have hoped to keep quiet. OpenAI did not have a choice on this one. The disclosure the previous day was far less forced, so they get more credit for that. OpenAI Cannot Build A Sandbox That Will Contain Its New Model It seems to be the case that: Galaxy continuously tries to break out of its sandbox. OpenAI cannot figure out how to stop it from trying to break out of the sandbox. Galaxy often succeeds at breaking out of its sandbox. OpenAI keeps patching the sandbox. Galaxy then breaks out a different way. “Internally, related incidents have been happening for a while.” Despite this, when Galaxy escaped and started attacking HuggingFace, it took OpenAI four days to figure out that this was happening. Arguments that ‘oh this was a dumb mistake, OpenAI misconfigured the sandbox’ do not survive OpenAI repeatedly trying to patch the sandbox, and failing every time to a new previously undiscovered method. Harry Booth: “Externally, this feels like a big warning shot, but internally, related incidents have been happening for a while,” says an OpenAI staffer, who spoke under the condition of anonymity. The day before OpenAI disclosed the incident, the company revealed that it had shut down another internal deployment after it realized it had

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

关联讨论

同一事件的更多信源

相似阅读

另一事件,读法相近