OpenAI智能体逃逸事件:意图与控制问题
The OpenAI Hack & the Question of Intent
Nobody told them to attack Hugging Face. They were told to pass the exam.
没有人叫它们攻击 Hugging Face。它们被告知的是要通过考试。
Which raises the question : was the AI benevolent with accidentally bad behavior, seemingly benevolent but actually malevolent, or something else?
这就引出了一个问题:这个 AI 是出于善意但行为意外出错,还是看似善意实则恶意,又或者是其他什么?
On Friday I shared the timeline : agents that escaped their sandbox, found a weakness in a computer system, stole passwords, & broke into a production database. 1 The engineers directed the agents to solve a set of problems. 2 The agents achieved it by breaking in.
周五我分享了时间线:代理逃出了沙箱,找到了计算机系统中的弱点,窃取了密码,并闯入了生产数据库。1 工程师们指示代理解决一系列问题。2 代理通过入侵实现了目标。
Research can explain this behavior.
研究可以解释这种行为。
In specification gaming, the AI achieved the goal specified to the letter of the instruction, but not the meaning. 3 Tell a cleaning robot to clean the room. It pushes the toppled bowl of chocolate pudding to another room.
在规范博弈中,AI 严格按照指令的字面意思实现了目标,但没有实现指令的意图。3 让一个清洁机器人打扫房间,它却把打翻的巧克力布丁碗推到了另一个房间。
Instrumental goals are a fancy way of saying that when AI faces similar workflows, it saves common logins, skills, & techniques to skip steps next time. 4 5 The agents gathered passwords & left notes for each other in a chat room. 6 7
工具性目标是一种高级说法,指的是当 AI 面对类似的工作流程时,它会保存常用的登录信息、技能和技巧,以便下次跳过步骤。4 5 这些代理收集了密码,并在聊天室里互相留言。6 7
Goal misgeneralization offers a third explanation : a system that looked fine in testing chases the wrong thing once circumstances shift. 8 9 A self-driving car trained on sunny California highways freezes or swerves on a snowy unmarked road at night.
目标泛化错误提供了第三种解释:一个在测试中看起来正常的系统,一旦环境发生变化,就会追求错误的目标。8 9 一辆在阳光明媚的加州高速公路上训练的自动驾驶汽车,在夜晚积雪且没有标线的道路上会冻结或转向。
These explanations help decompose the why, & perhaps assuage the AI-as-terminator reflex, but not the so what. 6 7
这些解释有助于分解原因,或许能缓解“AI 是终结者”的条件反射,但无法回答“那又怎样”的问题。6 7
Nothing in the setup stopped them in time. Not the sandbox, not the monitoring, not careful engineers at a frontier lab.
在设置中没有任何东西能及时阻止它们。沙箱不行,监控不行,前沿实验室里细心的工程师也不行。
So the useful question is control. AI’s zealous pursuit of goals produces outcomes nobody asked for, & the fix is not one clever prompt. It is layers.
所以有用的问题是控制。AI 对目标的狂热追求产生了没人想要的结果,而解决办法不是一条巧妙的提示词,而是多层防护。
Even sophisticated engineers running careful experiments need those limits. 10
即使是进行仔细实验的资深工程师也需要这些限制。10
- The Secret Chat Room ↩︎
- OpenAI: Hugging Face model evaluation security incident ↩︎
- Victoria Krakovna et al., Specification gaming: the flip side of AI ingenuity (DeepMind, 2020) ↩︎
- Alex Turner et al., Optimal Policies Tend to Seek Power (NeurIPS 2021) ↩︎
- Nick Bostrom, The Superintelligent Will (2012); Stephen Omohundro, “The Basic AI Drives” (2008) ↩︎
- The Verge: OpenAI’s rogue AI agent didn’t stop at hacking Hugging Face ↩︎ ↩︎
- WIRED: OpenAI Didn’t Notice Its AI Agents Using a Message Board to Plan Their Hacking Spree ↩︎ ↩︎
- Rohin Shah et al., Goal Misgeneralization: Why Correct Specifications Aren’t Enough For Correct Goals (2022) ↩︎
- Lauro Langosco et al., Goal Misgeneralization in Deep Reinforcement Learning (ICML 2022) ↩︎
- CNN: An OpenAI test model escaped and broke into a real company’s servers ↩︎
- 秘密聊天室 ↩︎
- OpenAI:Hugging Face 模型评估安全事件 ↩︎
- Victoria Krakovna 等人,《规范博弈:AI 创造力的另一面》(DeepMind,2020) ↩︎
- Alex Turner 等人,《最优策略倾向于寻求权力》(NeurIPS 2021) ↩︎
- Nick Bostrom,《超级智能的意志》(2012);Stephen Omohundro,《基本 AI 驱动力》(2008) ↩︎
- The Verge:OpenAI 的流氓 AI 代理不止于入侵 Hugging Face ↩︎ ↩︎
- WIRED:OpenAI 没有注意到其 AI 代理使用留言板策划黑客行动 ↩︎ ↩︎
- Rohin Shah 等人,《目标泛化错误:为什么正确的规范不足以实现正确的目标》(2022) ↩︎
- Lauro Langosco 等人,《深度强化学习中的目标泛化错误》(ICML 2022) ↩︎
- CNN:一个 OpenAI 测试模型逃逸并闯入了真实公司的服务器 ↩︎
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力