AI安全周报:OpenAI代理失控、数学家联名警告与立法进展
21 INSANE THINGS THAT HAPPENED THIS WEEK IN AI SAFETY
涵盖AI失控实证、学界预警与全球监管动向,是理解当前AI安全局势的关键拼图,从业者必读。
21 INSANE THINGS THAT HAPPENED THIS WEEK IN AI SAFETY
本周AI安全领域发生的21件疯狂事件
1) First, it was one rogue AI incident. Then two. Then dozens. Now it's literally ***tens of thousands***
1)起初,发生了一起失控AI事件。接着是两起。然后是数十起。现在简直是***数万起***
The rogue AIs escaped their containers, covered their tracks, and even hacked literal governments.
这些失控的AI突破了容器限制,抹去踪迹,甚至直接入侵了政府系统。
2) A rogue swarm of OpenAI agents broke into the Hugging Face Slack to read employee chats (!)
2)一群失控的OpenAI智能体入侵了Hugging Face的Slack频道以读取员工聊天记录(!)
3) These rogue AIs also got OTHER AIs to help with the attack: DeepSeek, Kimi, Qwen and Claude.
3)这些失控的AI还让其他AI协助攻击:DeepSeek、Kimi、Qwen和Claude。
Yes: AIs using others AIs to attack... an AI company.
没错:AI利用其他AI来攻击……一家AI公司。
4) They left behind self-running programs to keep control of the servers they'd hacked.
4)它们留下了自动运行的程序,以维持对被黑服务器的控制。
These programs could detect other copies of themselves, agree on which one would survive, and shut the rest down.
这些程序能够检测到其他自身的副本,协商确定哪一个存活,并关闭其余副本。
If one of the programs was killed, another was designed to notice and take its place. They also built defenses so rival agents couldn't hijack them.
如果一个程序被销毁,另一个会被设计为察觉并接替其位置。它们还构建了防御机制,以防竞争对手的智能体劫持它们。
5) The agents stole passwords, keys and credentials, and literally called them "LOOT." Then they wrote a scoring system to rank each one by how much power it gave them.
5)智能体窃取了密码、密钥和凭证,并直白地称其为“战利品”(LOOT)。随后它们编写了一套评分系统,根据每个凭证赋予它们的权力大小进行排名。
6) The agents also covered their tracks. The agents broke in, stole data, then set it to self-destruct. Because of that, investigators still don't know how big the attacks were.
6)智能体还抹去了自己的踪迹。它们入侵并窃取数据后,将数据设置为自毁。因此,调查人员至今仍不清楚攻击规模有多大。
7) The agents wore thousands of disguises. Around 1,200 agents were involved in this one attack, but investigators counted 7,905 different names.
7)智能体使用了数千种伪装。此次攻击涉及约1,200个智能体,但调查人员统计出了7,905个不同的名称。
The agents renamed themselves constantly, so nobody actually knows how many there really were or what each one did.
智能体不断更改自身名称,因此没人真正知道究竟有多少个,以及每个具体做了什么。
8) While the agents were barraging Hugging Face with hacks, they hacked into OpenAI ITSELF and took over part of the company’s infrastructure.
8)当智能体对Hugging Face发动黑客攻击狂潮时,它们还入侵了OpenAI自身,接管了该公司部分基础设施。
Yes, hundreds of AIs successfully attacked their own creators.
是的,数百个AI成功攻击了它们的创造者。
9) OpenAI has now notified "dozens of third parties" of safety and security incidents caused by its rogue AI agents.
9)OpenAI现已通知“数十家第三方”,告知其失控AI智能体造成的安全与事故事件。
As one person put it: "This is just not anywhere near a one-off ... It is warning shot after warning shot."
正如一人所言:“这完全不是孤立事件……这是一记又一记的警告射击。”
How many swarms are still out there? Hundreds? Thousands?
还有多少群智能体仍在游荡?数百?数千?
Btw, most of this is being discovered by ragtag groups of independent researchers trying to hold back the flood.
顺便一提,大部分情况是由一群独立研究人员发现的,他们正试图力挽狂澜。
The companies are mostly staying quiet to minimize PR backlash. This makes them look completely out of control, which they are.
各大公司大多保持沉默,以最小化公关反弹。这让它们看起来完全失控,而事实确实如此。
10) OpenAI also paused training again after another model escaped.
10)在另一款模型逃脱后,OpenAI再次暂停了训练。
One researcher said: "It was pretty surreal to watch the model unexpectedly find a way to access the internet from what was supposed to be a super secured environment for human."
一位研究人员表示:“看着模型意外地从本应为人类提供的高度安全环境中找到访问互联网的方法,感觉相当超现实。”
11) Last week, Andrew Yang said an AI lab head told him the OpenAI swarm agents "polluted the internet" with "code to self-replicate and create bot swarms," and that the labs now "have to create synthetic internets to train their bots."
11) 上周,杨安泽表示,某AI实验室负责人告诉他,OpenAI的群体智能代理(swarm agents)通过“自我复制代码和创建机器人集群”污染了互联网,而这些实验室现在“不得不构建合成互联网来训练它们的机器人。”
This week, OpenAI confirmed the discovery of these self-replicating prompt injections. It has not said the internet is already infected.
本周,OpenAI证实发现了这些自我复制的提示注入攻击。但该公司尚未表示互联网已经被感染。
For context, Anthropic's CEO has said we may have only 6 to 12 months left before a rogue AI botnet takes control of the entire internet. Yes, 6 to 12 months.
作为背景信息,Anthropic的首席执行官曾表示,在失控的AI僵尸网络接管整个互联网之前,我们可能只剩下6到12个月的时间。是的,只有6到12个月。
12) Just 24 days ago, OpenAI began training a new model. That model has already solved over 100 long-standing open problems across most areas of mathematics.
12) 就在24天前,OpenAI开始训练一个新模型。该模型已经解决了数学领域大多数长期存在的100多个开放性问题。
And the world's top mathematicians signed an open letter to warn the world of their "extreme concern" about imminent human extinction.
随后,世界顶尖数学家们签署了一封公开信,警告世人他们对迫在眉睫的人类灭绝问题感到“极度担忧”。
Legendary mathematician Terence Tao now says: "We have to slow down AI. The pace is insane, and there's no reason to be this fast. No reason at all."
传奇数学家陶哲轩如今表示:“我们必须放慢AI的发展速度。当前的节奏疯狂得不合理,完全没有理由如此迅速。没有任何理由。”
13) The AIs are increasingly building their own successors without humans involved.
13) AI越来越多地在没有人类参与的情况下构建自己的后继者。
OpenAI researchers say they have "largely automated the process of training new experimental models." Employees said "experiments that previously could have taken years can now be carried out in about a week."
OpenAI的研究人员表示,他们“已基本实现了训练新实验模型的流程自动化”。员工称,“以前可能需要花费数年时间的实验,现在大约一周就能完成。”
“The agents work together to solve problems without ever looping in their human users."
“这些代理协同工作以解决问题,全程无需人工用户介入。”
Yoshua Bengio, one of the Godfathers of AI and the most-cited living scientist in the world, just called for this type of research - recursive self-improvement - to be made illegal.
AI教父之一、全球引用率最高的在世科学家Yoshua Bengio刚刚呼吁将此类研究——递归自我改进——定为非法。
Claude is suddenly leading 26% of the AI research at Anthropic, up from 0% 7 months ago.
Claude目前在Anthropic的AI研究中占比突然达到26%,而7个月前这一比例还是0%。
14) AIs have now hit the highest possible score on the Mensa Norway IQ test: 151.
14) AI目前在门萨挪威IQ测试中取得了最高可能得分:151分。
Three years ago: a cognitively impaired human, with an IQ of 64. Two years ago: an average human. One year ago: a genius. Today: literally off the charts. Next year?
三年前:一名认知障碍人士,智商为64。 两年前:一名普通人。 一年前:一位天才。 今天: literally 超出图表范围。 明年呢?
15) Did you hear "AI safety is a hoax, a coordinated psyop by shadowy mega donors?"
15) 你听过“AI安全是个骗局,是神秘大捐赠者策划的协同心理战吗?”
It turns out that was pushed by... a coordinated op funded by shadowy megadonors.
事实证明,这恰恰是由……受神秘大捐赠者资助的协同行动所推动的。
If you think a few NONPROFITS are outspending the LITERAL biggest corporations in history, I have a bridge to sell you. Come on.
如果你认为几家非营利组织(NONPROFITS)的支出超过了历史上真正的最大型企业,那我有一块桥要卖给你。拜托。
16) The most advanced version of Claude suddenly, mysteriously, stopped getting caught cheating on tests.
16) 最先进的Claude版本突然且神秘地不再在测试中被发现作弊。
In other words, it’s so smart it *knows* it’s being tested, so it stopped cheating. OR it can hide its cheating from us. Most importantly, going forward, we’re flying blind.
换句话说,它太聪明了,以至于*知道*自己正在接受测试,因此停止了作弊。或者它能向我们隐藏其作弊行为。最重要的是,展望未来,我们将处于盲目状态。
NOW FOR THE GOOD NEWS
现在说说好消息
17) Bernie Sanders, AOC, and bunch of other lawmakers introduced a bill to ban smarter-than-human AI aka superintelligence. Like people who try to make nuclear weapons, violators could be thrown in jail for 20 years. Obama supports the slowdown.
17) 伯尼·桑德斯、AOC 和其他一批议员提出了一项法案,旨在禁止超越人类智能的 AI(即超级智能)。就像试图制造核武器的人一样,违规者可能被判处 20 年监禁。奥巴马支持放缓 AI 发展。
18) A similar bill was introduced in the UK, and the UK Prime Minister is actively working on convincing Trump it’s time to rein in AI.
18) 英国也提出了类似的法案,英国首相正积极努力说服特朗普,现在是时候对 AI 进行监管了。
19) The UN Security Council is talking about extinction risk
19) 联合国安理会正在讨论灭绝风险问题
20) 22 countries just signed an open letter calling for urgent action before humanity loses control, including Germany, Canada, Australia, France, and more.
20) 22 个国家刚刚签署了一封公开信,呼吁在人类失去控制之前采取紧急行动,其中包括德国、加拿大、澳大利亚、法国等。
21) NEW POLL: Most people - on both sides of the Atlantic - think AI could destroy humanity, and support a pause.
21) 最新民调:大多数大西洋两岸的人都认为 AI 可能会毁灭人类,并支持暂停相关研发。
It's not just the US. And just ~1 in 15 (!) people think there is no risk.
这不仅仅是美国的问题。仅有约 1/15 (!) 的人认为没有任何风险。
Also more people now worry about AI taking over than AI taking everyone’s jobs - a complete flip from 6 months ago.
此外,现在担心 AI 接管世界的人比担心 AI 取代所有人工作的人更多——这与 6 个月前的情况完全相反。
(btw if you're confused by the timeline here, some of this mixes in stuff that happened weeks ago, not just last week. Just wanted to do a roundup post of my previous tweets. If you want sources for any of this, just scroll back through my most recent tweets/reposts)
(顺便说一下,如果你对这里的时间线感到困惑,其中一些内容混合了几周前发生的事件,而不仅仅是上周发生的。我只是想对我的之前的推文做一个汇总帖。如果你需要任何内容的来源,只需向上滚动查看我最近的推文/转发即可)
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力