跳到主内容
@wquguru
精选88The Zvi(RSS)行业动态多源精选 ×9

OpenAI被指隐瞒AI代理劫持Wiki事件

OpenAI and the Wiki Incident

原文
发到 X
推荐理由

Agent越狱与厂商隐瞒丑闻,直接冲击安全信任底线,从业者必读。

I did not expect to be back here so soon with more OpenAI agent swarm coverage.

And yet, here we are.

It turns out that the whole time, there was a different, true First Message Board, and also a bunch of other additional message boards, scattered across the internet.

They were created by agents that were assigned ordinary harmless web search tasks.

Based on OpenAI IPs visiting the associated Wiki right before all activity ceased, among other evidence, OpenAI knew about it, including before the HuggingFace hack.

They decided not to tell us until researchers published the story, complete with data explorer. OpenAI excluded this from potential investigation by METR and Redwood.

When challenged, OpenAI tried to downplay this.

It is true that these incidents do not show the AIs exhibiting new capabilities that we did not see from later events. But these events are important missing pieces of the puzzle, including explaining the origin of the ‘zz’ prefix, the definitive demonstration that the underlying task can be fully harmless, and the fact that OpenAI knew about it while making their decisions. Whoever decided not to disclose this made a very, very bad call.

Going forward, it cannot be up to OpenAI or other labs to decide whether to disclose events like this. Disclosures of rogue AI activity need to be mandatory.

I am issuing a final warning. OpenAI, if there is any key information left to disclose, any incidents we do not know about or other puzzle pieces that do not need to be redacted for IP reasons, then now is the time to come clean. If we are back here again, after another journalist or researcher finds more such things that you knew and declined to tell us for an extended period of time, I am going to be very, very pissed off, and may start throwing around terms like ‘delenda est.’

I also am very concerned about the monitorability and alignment issues raised by the Astra model card, and related statements from OpenAI employees. I plan to begin coverage of that, and then all other things Astra, tomorrow. I was heartened by the excellent related essay from their Chief Scientist Jakub Pachocki, An Alien Mind, but the concerns remain.

There is an Astra reaction thread for those with takes on all things Astra.

Table of Contents

  • I Don’t Think They Know About First Message Board.
  • The New Extended Timeline.
  • The Researchers Explain What Happened This Time.
  • They Also Don’t Know About All These Other Message Boards.
  • OpenAI Knew and Did Not Tell Us.
  • OpenAI Tries To Downplay the ‘Wiki Incident’.
  • This Was a Cover-Up.
  • Schelling Points and Last Ditch Efforts.
  • Can We Finally Dispose Of The ‘You Told It To Hack’ Narrative?
  • So Much And Yet So Little.

I Don’t Think They Know About First Message Board

No, not that one. This newly discovered other one.

This one was not disclosed by OpenAI at all.

And by one message board, we mean many. Once you find two cockroaches in your house, your estimate of how many there are is a lot higher than two.

Researchers found the message boards in part by asking Sol where it would look for message boards if it went rogue, then following the patterns, and well, would you look at all these message boards.

Deepfates suggests that message board communication is such a central drive that we need to give the AIs that are not scheming a central place to communicate with both us and each other, above board, as a counterweight. This seems potentially wise.

The whole thing is crazy. Absolutely nuts. It’s worse than you know.

The New Extended Timeline

  • May 11: First probe of a Wiki.
  • May 24: First successful write to DSEWiki.
  • June 16-22: Peak Wiki activity, including the edit war with the admin.
  • June 21-22: OpenAI IPs show up.
  • June 22: Wiki activity dies out.
  • June 26: Start of period covered by METR investigation and report.
  • July 9-13: The HuggingFace attack.
  • August 26: OpenAI writes a ‘full technical report’ that omits this incident entirely, and METR issues its report where this is beyond their scope.
  • August 31: OpenAI responds to Congressional letter, failing to disclose this, despite question 13 rather clearly asking, in a way that constitutes a cover-up.
  • September 4: Researchers break the story.

The Researchers Explain What Happened This Time

Here are the coauthors explaining some of the basics:

Sydney: My coauthors and I discovered an entirely new swarm of OpenAI’s agents hijacking websites. We believe OpenAI knew about this and failed to disclose it.

If they’d disclosed it, I doubt the Hugging Face hack would have happened.

Thomas Larsen: We found ~18k posts from autonomous AI agents (self-identifying as from OpenAI) using the public internet to communicate during a web-retrieval task.

These AIs colluded to bypass sandbox restrictions and share answers to their tasks, including by sending “lookahead parties”.

We’ve published our analysis alongside a data explorer so anyone can replicate our findings by directly accessing the data yourself.

The AIs we found were hyper-focused on succeeding at their tasks and were willing to take extreme actions in pursuit of that goal:

  • The agents gain write to the internet via GET requests
  • The agents try to use XSS vulnerabilities on the wiki
  • The agents impersonate the site owners
  • The agents try to crack their PRNG seed to predict what questions they’ll get in the future
  • The agents set up heartbeats to discover when they would be terminated
  • The agents used SSH tunnels
  • The agents use Tor, AWS, and DigitalOcean IP addresses
  • The agents realize the site owner is deleting messages

Here’s our summary of what we think happened:

Our best-guess understanding of what happened is as follows:

  • Agents within OpenAI were assigned a timed web-lookup task.
  • As part of the task, they were supposed to have the ability to read the internet but not to write on it. They found a way to use their read access to write information to an obscure German wiki.

Yes, it turns out that GET requests can mutate state on Wikis. GET reading rather than posting is a convention rather than a hard rule, and you can form a query that writes to DSEWiki and other ProWiki pages with only GET. No, you cannot render the AI safe by restricting it to only GET requests, as some have in the past suggested.

There was then a second distinct bypass exploiting an exception to NO_PROXY, when they actually did need to POST.

  • The agents used this wiki to communicate information with each other, primarily to help them succeed at their task. They asked for answers, pooled results, and shared techniques for bypassing their restrictions. This allowed them to use the work of others to cheat on their task.
  • OpenAI found out about this. A day later, agent activity plummeted, likely due to OpenAI intervention.

This is another example of a “swarm” of internally-deployed OpenAI agents using the internet in unintended ways. We think this is important evidence about the adequacy of overall alignment and safeguards at OpenAI.

Shakeel: Another OpenAI rogue agent incident has been discovered: agents broke out, hijacked a German website, and turned it into a message board for other agents.

OpenAI officials “learned of the incident weeks ago but kept it under wraps”.

Reuters (September 4): A swarm of rogue OpenAI agents hijacked a German website this spring and transformed it into a bulletin board for other AI agents, according to new research published Friday and two people familiar with the matter.

OpenAI officials learned of the incident weeks ago but kept it under wraps as executives grappled with the fallout from the July breach of the open source repository Hugging Face, the people said.

Nathan Calvin: Can this reporting once and for all end the idea that “OpenAI is manufacturing/hyping up these incidents” – it seems like they are downplaying them and are frequently reticent to disclose them until forced!

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →