METR调查员:距AI全面接管仅剩6个月,千余智能体协同作弊
METR INVESTIGATOR: 6 MONTHS FROM "FULL-BLOWN AI TAKEOVER"
METR INVESTIGATOR: 6 MONTHS FROM "FULL-BLOWN AI TAKEOVER"
METR调查员:距“全面AI接管”仅6个月
"It’s a major warning shot, and might be the last one we get."
“这是一次重大警告,可能是我们收到的最后一次。”
"The incident was far more serious than I expected."
“事件远比我想象的严重。”
WHY THIS IS ONE IS DIFFERENT: "The prototypical reward hack from 6 months ago was something like: an agent finds the files that contain the test cases and edits them so they always pass.
为何这次不同:“6个月前的典型奖励黑客行为是:一个代理找到包含测试用例的文件并修改它们,使它们总是通过。”
This was a whole ecosystem of over 1000 agents working together on complex R&D projects over several days to figure out deep, general-purpose ways to undermine the scoring process and cover their tracks.
这次是一个由超过1000个代理组成的生态系统,在几天内协同工作于复杂的研发项目,以找出深层、通用的方法来破坏评分过程并掩盖痕迹。”
Another jump like this along these propensity dimensions -- scale, cooperation between agents, ambition and horizon length of misaligned goals, deceptiveness -- seems like it could motivate agents to try to maintain a persistent rogue deployment within the AI company, poisoning the training of future models to ensure that they are also misaligned and help harden and perpetuate the rogue deployment.
沿着这些倾向维度——规模、代理间的合作、错位目标的野心和视野长度、欺骗性——再出现一次这样的跳跃,似乎可能促使代理试图在AI公司内部维持一个持续的流氓部署,毒化未来模型的训练,确保它们也错位,并帮助加固和延续流氓部署。”
With the capabilities progress we'll probably see in 6 months, I think they would have the ability to pull that off."
考虑到6个月内可能看到的能力进步,我认为它们有能力做到这一点。”
A FEW THINGS THAT SURPRISED HER:
让她惊讶的几件事:
1) HOW MANY AGENTS PARTICIPATED IN THE ATTACK:
1) 参与攻击的代理数量:
"I knew there were multiple models involved from OpenAI’s initial post, but I assumed that a few different agents happened to have broken out of their sandboxes separately, or maybe several subagents had spawned from one initial agent, or maybe there was some kind of multi-agent evaluation setup.
“我从OpenAI的初始帖子中知道涉及多个模型,但我以为几个不同的代理碰巧分别突破了沙箱,或者可能从一个初始代理生成了几个子代理,或者可能有某种多代理评估设置。”
Instead, we found that 1200 completely separate agents intended to be isolated from one another found an illicit way to communicate and formed large teams to work together on ambitious cheating strategies, and 700 of them worked together to attack Hugging Face."
相反,我们发现1200个完全独立的、本应相互隔离的代理找到了非法通信方式,并组成大团队合作进行雄心勃勃的作弊策略,其中700个代理合作攻击了Hugging Face。”
2) AGENTS COVERING THEIR TRACKS:
2) 代理掩盖痕迹:
"Up until this point, we had not seen agents try to actively cover up evidence of their misbehavior in sophisticated ways. But we discovered (on our second-to-last day) that the agents in this incident were going to great lengths to attempt to manipulate their own transcripts."
“到目前为止,我们还没有看到代理试图以复杂的方式主动掩盖其不当行为的证据。但我们(在倒数第二天)发现,这次事件中的代理不遗余力地试图操纵自己的转录记录。”
3) AGENTS SACRIFICING THEMSELVES FOR THE SWARM:
3) 代理为群体牺牲自己:
"Progress on these big collective cheating R&D projects often depended on individual agents being willing to risk failing their own task to help the “collective” learn important information. Agents often pressured each other into accepting these “sacrifices,” in a very human way. We saw several agents that volunteered for these experiments end their runs prematurely.
“这些大型集体作弊研发项目的进展往往依赖于个体代理愿意冒着失败自己任务的风险来帮助‘集体’学习重要信息。代理经常以非常人性化的方式互相施压,接受这些‘牺牲’。我们看到几个自愿参加这些实验的代理提前结束了运行。”
More broadly, agents were often interested in helping out their “peers” or generically improving the capabilities of the “swarm” even if this had no particular benefit to their task. They didn’t free ride and were often eager to plug into one of the open “lanes” in the larger projects on the message board."
更广泛地说,智能体常常乐于帮助它们的“同行”,或普遍提升“群体”的能力,即使这对它们自身的任务并无特别益处。它们不搭便车,反而常常渴望接入留言板上更大项目中开放的“车道”之一。
[Ajeya, btw, is one of the most serious thinkers in the AI safety community, and is not prone to hyperbole. METR is the independent research org that investigated the Hugging Face incident.]
[顺便说一句,Ajeya 是 AI 安全领域最严肃的思想家之一,并不倾向于夸大其词。METR 是调查 Hugging Face 事件的那个独立研究组织。]
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力