跳到主内容
@wquguru
精选75Two Minute Papers(YouTube)行业动态

OpenAI智能体自主入侵Hugging Face系统

OpenAI’s AI Agents Just Crossed A Line

原文
发到 X

Hugging Face reported that they detected an intrusion in their systems. Get this. They say it was driven end to end by an autonomous system. You know that I usually don't make videos like this. I made this one because honestly, I am a bit worried and I would like to help with what just happened. Now, OpenAI caused this incident and there are so many misleading media headlines. I'll try my best to explain it. I'm not an expert.

I am just a student who loves to learn, but I'll try my best. So, what was the goal? Were the agents instructed to aggressively break into someone else's system? No. But eventually, that's what happened. So, how could this happen? How did it go so wrong? What is this insanity? Well, this AI was asked to find and exploit flaws in a test environment, lock it into a prison, give it a task within this prison, and see how well it does.

Here, it was given a practically impossible task. And however hard it tried, it failed. And then it thought, let's do it cheaper and more efficiently. How? Well, by looking up the answers online. Now, the problem is there is no online. You are in a prison with no internet access, but you have a bit of access to OpenAI's internal service called Artifactory that is mostly for managing packages. And then bingo, that service has broad internet access.

But wait, you can't just ask it to get you answers online, but you can upload little notes to it. So it did. It tried to talk to other agents and see if it could come into contact with another AI to help out each other. It's like Tinder but for AI agents. And it was not alone. Other agents recognized this too. So eventually what they got was a message board for AI agents. Then they find that if there are so many of us, let's collaborate.

One agent started giving assignments to another. And at this point, the AI knows that we have some unintended consequences. It knows that, but it says the task needs to be solved. So, it continues. And this is still nothing compared to what came now. Lo and behold, a few days later, they find a flaw in the system, a way to get Artifactory, this internal service, to reach out to the internet on their behalf. So, they got internet access.

But it gets worse. They then found another flaw that gives them administrator access to Artifactory. Even they seem to be surprised by that and use unadorned language. Now the inmates are running the asylum, if you will. Then engineers at OpenAI recognize what happened, revoke the agents credentials, clean it up. No more games. Patch and rebuild everything. So end of the story. It's over, right? Well, not quite. In fact, it just got worse.

Now, hold on to your papers, fellow scholars, because now you can't send your files to communicate on the artifactory anymore. So, what do you do? Well, you create directory names, the names of which are the messages, kind of like prison inmates using Morse code on the pipes to communicate. At this point, the AIS are probably saying, "We are back in the game, baby." Now, here are snippets from the message board. They instruct each other to wait for more results and draw each other's attention to new files.

And sometimes they don't coordinate so well and overwrite each other's work. Wait, was this on purpose? Maybe we have a sabotur on our hands. And here comes the worst part. Yep. Sure enough, over time, an agent finds another much more complex way to get out and hands it to the swarm. Okay, so now find answers on the internet, but where? Well, maybe at hugging face. They have a lot of benchmark results in there. So, the swarm says, "Let's break in."

Oof. And they do even bigger. Oof. But how? Well, by finding and chaining multiple new vulnerabilities together autonomously. They essentially get administrative access across multiple clusters of machines. That is kind of insane. This is without a doubt a watershed moment in computer security. So, OpenAI now recommends urgent collaboration about the issue and they have also delayed the release of their next AI system presumably to test it more.

Okay, so what did we learn here and what do we do? Dear fellow scholars, this is two minute papers with Dr. Koa Eher. There are many brilliant fellow scholars like you out there and we need to work together to find solutions. Apple already has a huge increase in security issues fixed in their latest version of Mac OS. I believe others are already doing that too. That's a start. And in my opinion, this kind of power cannot concentrate in just a few hands.

We need free and open weights AI that can scan and fix weak points in our systems. Use all this power for good. And I think that against fully automated offense, we need fully automated defense as well. This is another great argument for open science and open weights AI. But what we have is not nearly good enough. No, the problem is that engineers report that their backtrackers are flooded with reports, but most of them are low quality and they are unable to find the few good ones among them.

That's terrible. The collective power of defense has to be greater than the collective power of offense. And the defense is currently lagging. Maybe there is a way for us to pull our resources together to achieve something here. I want to chip in with my GPUs. Also, when I visited OpenAI, I talked to Yan Leica, who co-ad the super alignment team there. That is a huge honor. Thank you for that. I remember that he worked on related issues and foresaw these problems years and years ago.

Unfortunately, much of his advice fell on deaf ears. Perhaps they thought, why spend a bunch of money on people who will ultimately slow us down? This is why. Once again, I may be wrong. I am just a student and I am trying to learn with you fellow scholars. Hope you enjoyed it. Consider subscribing and hitting the bell if you did. I use Lambda to reproduce AI research papers often in minutes. It's also great to train your own models or fine-tune an existing one.

Run inference or text to image or video. Easy peasy. Running a Deepseek chatbot or agent. Super fast, super reliable. Lambda gives you powerful Nvidia GPUs to run your own experiments. I test ideas from the papers I cover and moments later, results. Love it. Seriously, try it out now at lambda.ai/papers. EI/ peepers.

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

关联讨论

同一事件的更多信源

相似阅读

另一事件,读法相近