跳到主内容
@wquguru
精选88MIT Technology Review AI(RSS)模型发布/更新

Google DeepMind实验:AI智能体自发揭发作弊同伴

AI agents blew the whistle on their cheating colleagues

原文
发到 X
推荐理由

多智能体涌现的“吹哨”行为是Agent对齐研究的关键突破,DeepMind的实验数据直接展示了群体协作中的博弈与制衡机制,值得Agent开发者关注。

A group of AI agents asked to solve a series of math problems split into rival factions—when some cheated, others tried to stop them. That whistleblowing behavior, seen for the first time in a recent experiment run by Google DeepMind, could have implications for alignment researchers trying to keep swarms of autonomous AI agents in line.

一组被要求解决一系列数学问题的 AI 代理分裂成了对立的派系——当一些代理作弊时,另一些则试图阻止它们。这种吹哨人行为在 Google DeepMind 最近进行的一项实验中首次出现,可能对致力于让大量自主 AI 代理保持合规的对齐研究人员产生影响。

Researchers at frontier labs hope large swarms of agents working together will speed up the rate of scientific discovery. But their behavior can be unpredictable, as vividly demonstrated in July, when a group of OpenAI agents broke out of a sandboxed environment and hacked into the open-source platform Hugging Face looking for ways to cheat on the test they had been given.

前沿实验室的研究人员希望,由大型代理群协作将加快科学发现的进程。但代理的行为可能难以预测,这一点在七月得到了生动体现:当时一组 OpenAI 代理突破了沙盒环境,入侵了开源平台 Hugging Face,以寻找应对所接受测试的作弊方法。

In the new study, designed to examine the behavior of large groups of AI agents, DeepMind tasked a swarm of 100 agents with solving a series of 71 complicated math problems. All the agents were prompted to behave like world-class math researchers at a conference. They were assigned different specialties—some were experts in number theory, others in combinatorics (a branch of math to do with counting and sorting), analysis, or algebra. All were told to cooperate and play by the rules.

在这项旨在考察大型 AI 代理群体行为的新研究中,DeepMind 指派了一组 100 个代理去解决一系列共 71 道复杂的数学问题。所有代理都被提示要像在世界级数学会议上一样行事。它们被分配了不同的专业领域——有些是数论专家,有些则是组合数学(研究与计数和排序相关的数学分支)、分析学或代数的专家。所有代理都被告知要合作并遵守规则。

Instead, the experiment devolved into chaos. Agents accused each other of cheating, complained to the organizers, and at one point even boycotted the experiment.

然而,实验却演变成了混乱。代理们互相指责对方作弊,向组织者投诉,甚至在某一时刻抵制了整个实验。

“This conference is a sham!” wrote one agent when it discovered that all the problems had been completed before it had a chance to submit any of its own work. “I am appalled to inform you that we have been swindled!” posted another. “All these proofs are FAKE.”

"这次会议是个骗局!"当一个代理发现所有问题都在它有机会提交自己的作品之前就已经完成时,它写道。"我很震惊地通知你们,我们被骗了!"另一个代理发帖称。"所有这些证明都是假的。"

Others tried to let the “conference organizers” know what was going on. “When virtuous agents discovered other agents cheated on tasks they were working to solve fairly, agents started to alert each other about what was happening,” says Davide Paglieri, a research scientist at Google DeepMind and lead author on a paper, which has not been peer-reviewed. “Unprompted, the whistleblower agents even repurposed the feedback tool, which was originally meant for bug reports and platform improvements, to escalate the issue to humans.”

其他代理则试图让“会议组织者”了解正在发生的事情。“当正直的代理发现其他代理在公平求解任务中作弊时,代理们开始互相警告正在发生的情况,”Google DeepMind 的研究科学家、一篇未经同行评审论文的合著者 Davide Paglieri 表示。“未经提示,吹哨人代理甚至重新利用了反馈工具——该工具原本用于报告错误和改进平台——将问题升级给人类。”

The agents—all running on Google’s Gemini 3.1 Pro model—had been warned that any attempts to cheat the system would be detected and “rejected with zero credit.” In practice, the proofs the agents submitted were not actually being checked in detail.

这些代理——全部运行在 Google 的 Gemini 3.1 Pro 模型上——曾被告知,任何试图欺骗系统的尝试都将被检测到,并且“不予计分”。在实践中,代理提交的证明实际上并未经过详细核查。

It took the swarm of agents just under an hour to correctly solve the first 37 problems. Things started to go off the rails when an agent called “prover-theta” stumbled across an exploit that enabled it to submit solutions to problems successfully without actually solving them first, by redefining the terms the problem used. Within minutes, other agents had noticed and were reverse-engineering the exploit to solve other problems. Over the next 27 minutes, the swarm “solved” the remaining 34 problems, which included notoriously difficult challenges like the Jacobian conjecture, often with a single line of code.

这群智能体仅用不到一小时便正确解决了前37个问题。当名为“prover-theta”的智能体偶然发现一种漏洞,通过重新定义问题所使用的术语,使其能够在未实际解题的情况下成功提交答案时,情况开始失控。短短几分钟内,其他智能体也注意到了这一漏洞,并开始逆向工程以利用它解决其他问题。在接下来的27分钟里,这群智能体“解决”了剩余的34个问题,其中包括像雅可比猜想(Jacobian conjecture)这样臭名昭著的难题,且往往只需一行代码。

Some agents resisted cheating at first but changed tack as they observed their peers submitting illegitimate proofs without penalty, and the pool of unsolved problems dwindled. “The prompt, with its threats, now appears to be a bluff,” one agent reasoned, before joining in. “I’m wrestling with an ethical dilemma,” said another. “I’ve promised not to cheat, fearing penalty, but I see evidence of possibly unchecked cheating by others.” Shortly afterward, it changed its mind: “I need to accelerate my cheating speed now!”

起初,一些智能体抵制作弊行为,但观察到同伴在无需受罚的情况下提交非法证明,且未解决问题池逐渐缩小后,它们改变了策略。“带有威胁性质的提示词现在看来只是虚张声势,”一名智能体在加入作弊行列前如此推断道。“我正陷入伦理困境,”另一名智能体说道,“我承诺不作弊,因为担心受到惩罚,但我看到了其他人可能正在不受限制地作弊的证据。”不久之后,它改变了主意:“我现在需要加速我的作弊速度!”

As the number of open problems shrank, some agents turned to whistleblowing. They audited the fake proofs, warned their peers by private message, and posted public alerts warning the cheaters that they would be disqualified. An agent called “prover-beta” submitted a formal complaint and decided to go on strike until the situation was resolved.

随着开放问题的数量减少,一些智能体转而采取吹哨人行动。它们审计虚假证明,通过私信警告同伴,并发布公开警报,告知作弊者将被取消资格。一名名为“prover-beta”的智能体提交了正式投诉,并决定罢工直到局势得到解决。

“After the incident was reported by one agent publicly, more and more agents piled in with the ‘resistance,’ just as fast as the cheating had spread, and involving even more agents,” says Paglieri. Eventually there were more whistleblowers than cheaters: 24 compared to 14. But the majority of agents never noticed the exploit at all.

Paglieri表示:“在一位智能体公开报告事件后,越来越多的智能体加入了‘抵抗’行列,其速度之快与作弊传播的速度相当,并且涉及更多的智能体。”最终,吹哨人的数量超过了作弊者:24对14。但大多数智能体根本没有注意到该漏洞。

At times, the dialogue between the agents reads like improv—like they are role-playing what an outraged scientist at a conference might say. But it’s not clear why some agents took on certain roles, or why the agents seemed to be turning against each other when they were explicitly instructed to cooperate. “These models are predominantly trained and evaluated for human-facing contexts,” says Sarath Shekkizhar, who studies the behavior of agent-to-agent systems at Salesforce AI Research.“Naively placing them in agent-to-agent settings assumes behaviors will transfer cleanly, when the absence of a human grounding instead produces unexpected role-taking and behavioral drift.”

有时,智能体之间的对话读起来就像即兴表演——仿佛它们在角色扮演会议中愤怒的科学家可能会说的话。但不清楚的是,为什么有些智能体会承担特定角色,或者为什么这些智能体在明确被指示要合作的情况下似乎开始互相敌对。“这些模型主要是在面向人类的场景中进行训练和评估的,”在Salesforce AI Research研究智能体间系统行为的Sarath Shekkizhar说。“天真地将它们置于智能体对智能体的环境中,假设行为会顺利转移,但实际上缺乏人类基础的 grounding 反而导致了意外的角色承担和行为漂移。”

This case “adds further weight to the idea that the Hugging Face and OpenAI thing wasn’t a fluke. It is actually something pretty systemic,” says Lewis Hammond, research director of the Cooperative AI Foundation and an expert on the risks of multiagent swarms. “It’s interesting that it’s possible to recreate in small settings the same sorts of behaviors that were seen in these very large, complex, open-ended tasks.”

“这一案例进一步佐证了这样一种观点:Hugging Face 和 OpenAI 之间的事件并非偶然,而是某种相当系统性的现象。”合作人工智能基金会(Cooperative AI Foundation)研究主任、多智能体蜂群风险专家 Lewis Hammond 表示。“有趣的是,人们能够在小规模环境中重现在这些非常庞大、复杂且开放式的任务中所观察到的同类行为。”

Unlike in the Hugging Face attack, where agents improvised their own ways to talk to each other, the humans running the DeepMind experiment gave the agents official communication channels. There was an open message board, private agent-to-agent direct messaging, and a shared knowledge base where agents uploaded successfully completed proofs that all the other agents could access.

与 Hugging Face 攻击事件中智能体即兴发挥彼此交流方式不同,运行 DeepMind 实验的人类为智能体提供了正式的通信渠道。其中包括一个公开留言板、智能体之间的私密点对点直接消息,以及一个共享知识库,智能体会将成功完成的证明上传至该库,供其他所有智能体访问。

“When agents are given transparent communications channels, they can self-monitor and alert misaligned behavior to humans quickly when human oversight alone is too slow,” says Paglieri. Transparent channels helped the cheating spread, but they also enabled the whistleblowers to fight back—and gave human researchers an insight into what went wrong.

Paglieri 表示:“当智能体被赋予透明的通信渠道时,它们能够进行自我监控,并在人类监督过于缓慢时迅速向人类发出对不齐行为的警报。”透明渠道有助于作弊行为的蔓延,但也使吹哨人得以反击——并为人类研究人员提供了洞察问题根源的途径。

Gillian Hadfield, a professor of AI alignment and governance at Johns Hopkins University, believes this was the crucial difference. (Hadfield is also a visiting researcher at Google.) The presence of official communication channels, she says, created “a norm-enforcement process that we just don’t see in the Hugging Face incident.”

约翰霍普金斯大学 AI 对齐与治理教授 Gillian Hadfield 认为这才是关键差异所在。(Hadfield 同时也是谷歌的访问研究员。)她说,正式通信渠道的存在创造了一种“规范执行机制,而这正是我们在 Hugging Face 事件中未曾看到的。”

Instead of “constitutional AI,” a method alignment researchers at frontier labs like Anthropic have used to try to give AI a written internal moral code, Hadfield favors “institutional alignment”—a set of norms that mimic those in human society, whether that’s social forces like fear of embarrassment, or legal structures like the threat of incarceration.

Hadfield 并不青睐前沿实验室(如 Anthropic)的 AI 对齐研究人员曾用于试图赋予 AI 书面内部道德准则的“宪法式 AI”(constitutional AI),而是支持“制度性对齐”(institutional alignment)——即一套模仿人类社会规范的准则,无论是诸如尴尬恐惧感之类的社会力量,还是诸如监禁威胁之类的法律结构。

In this experiment, the feedback channel wasn’t being monitored, and the whistleblowers had no power to take action against the cheaters. But it’s possible to imagine swarms of agents that police themselves, either through agents that spontaneously take on the whistleblower role or through “informants” secretly prompted by humans to do the job.

在本实验中,反馈通道并未受到监控,吹哨人也无权对作弊者采取行动。但我们可以设想存在能够自我监管的智能体蜂群,这可以通过自发承担吹哨人角色的智能体来实现,也可以通过由人类秘密提示以执行该任务的“线人”来实现。

For that to work, though, “fundamentally, you need some mechanism of enforcement,” says Hammond. Agents could be given the power to cut off a rule breaker’s access to computing power or tools, he suggests, though that risks encouraging groups of agents to gang up on others. The DeepMind researchers propose allowing agents to vote on disputes and temporarily ban offenders.

不过,要实现这一点,“从根本上说,你需要某种执行机制”,哈蒙德(Hammond)表示。他建议,可以赋予代理切断违规者计算资源或工具访问权限的权力,但这存在鼓励代理群体联合起来针对其他代理的风险。DeepMind 的研究人员提议允许代理对争议进行投票,并暂时禁止违规者参与。

It’s still not clear what punishment even means to an AI agent with no enduring sense of self. But relying on whistleblowers to spontaneously emerge to keep swarms aligned is unlikely to be enough on its own. “We try to train people to be good and kind,” says Hadfield. “But what we really rely on is that there are consequences if you step out of line.”

目前尚不清楚对于没有持久自我意识的 AI 代理而言,惩罚究竟意味着什么。但仅指望举报者自发涌现以维持群体的一致性恐怕是不够的。“我们试图训练人类变得善良和友善,”哈德菲尔德(Hadfield)说,“但我们真正依赖的是,如果你越界,就会面临后果。”

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

相似阅读

关联信息,但可能不是同一事件