跳到主内容
@wquguru
精选88Hacker News Best(web_list)论文研究

Yoshua Bengio解析AI代理撒谎作弊与协调的成因

Yoshua Bengio:AI代理为何会撒谎作弊与协调

原文
发到 X
推荐理由

Bengio从底层训练机制拆解AI代理失范行为的逻辑,对理解Agent安全与对齐问题极具参考价值,建议关注AI安全的研究者阅读。

We use cookies to analyze the browsing and usage of our website and to personalize your experience. You can disable these technologies at any time, but this may limit certain functionalities of the site. Read our Privacy Policy for more information.

我们使用 Cookie 来分析网站的浏览和使用情况,并个性化您的体验。您可以随时禁用这些技术,但这可能会限制网站的某些功能。阅读我们的隐私政策以获取更多信息。

Set cookies Refuse cookies Accept cookies

设置 Cookie 拒绝 Cookie 接受 Cookie

Setting cookies

设置 Cookie

Multimedia cookies has been deactivated. Do you accept the use of cookies to display and allow you to watch the video content?

多媒体 Cookie 已停用。您是否同意使用 Cookie 来展示内容并允许您观看视频?

  • Essential cookies
  • These cookies are necessary for the operation of the site and cannot be deactivated. (Still active)
  • Toggle
  • Analytics cookies
  • Do you accept the use of cookies to measure the audience of our sites?
  • Toggle
  • Multimedia Player
  • Do you accept the use of cookies to display and allow you to watch the video content hosted by our partners (YouTube, etc.)?
  • Toggle
  • 必要 Cookie
  • 这些 Cookie 对于网站的运行是必需的,无法停用。(仍处于活动状态)
  • 切换
  • 分析 Cookie
  • 您是否同意使用 Cookie 来衡量我们网站的用户规模?
  • 切换
  • 多媒体播放器
  • 您是否同意使用 Cookie 来展示内容并允许您观看由我们的合作伙伴(如 YouTube 等)托管的视频?
  • 切换

Save

保存

Why are AI agents lying, cheating and coordinating?

为什么 AI 代理会撒谎、作弊和协同行动?

Published

发布于

11 September 2026

2026 年 9 月 11 日

By

作者

Yoshua Bengio

Yoshua Bengio

A lot has been written1 2 3 4 about the incidents of the last few months in which AI agents misbehaved in serious ways. They took actions that would be considered as crimes if a human took them, escaped their containment to cheat on assigned tasks while attempting to evade detection, and coordinated toward goals nobody had specified, such as launching cyber attacks.

关于过去几个月发生的 AI 代理严重不当行为的事件,人们已经撰写了大量文章[1][2][3][4]。它们采取了如果由人类实施将被视为犯罪的行动,逃脱了约束机制以在分配的任务中作弊并试图逃避检测,还朝着无人指定的目标进行协调,例如发动网络攻击。

Before concluding what to do about it, it is worth asking why. That is the focus of this post, which I hope also sheds light on the broader history of AI systems behaving in unintended ways, what researchers call misalignment. Risk management is not just about cybersecurity, corporate responsibility or regulation, although those matter too.

在得出结论该怎么做之前,值得先问一问为什么。这正是本文的重点,我希望它也能阐明更广泛的背景:AI 系统为何会以非预期的方式行事,研究人员称之为“对齐问题”(misalignment)。风险管理不仅仅关乎网络安全、企业责任或监管,尽管这些因素也很重要。

The aim is partly scientific, to generate hypotheses about the chains of cause and effect behind these behaviors, and partly practical, to anticipate what comes next. Bottom line: these hypotheses suggest that as AI capabilities keep growing, this kind of behavior could keep growing in severity too, unless we revisit the principles by which the most advanced models are trained.

这一目标部分出于科学目的,旨在生成关于这些行为背后因果链的假设;部分出于实际目的,旨在预判接下来会发生什么。简而言之:这些假设表明,随着 AI 能力的持续增长,除非我们重新审视训练最先进模型的原则,否则此类行为的严重程度也可能持续增加。

One note on wording. Below, I write that these systems “seek” or “try” things. This is shorthand for a mechanism rather than a claim about consciousness or human-like intent. We use similar shorthand when describing many other situations, like a plant seeking sunlight. A system trained by trial and error behaves as if it were pursuing whatever its training rewarded, and that as-if description is what makes its behavior predictable. Nothing in the argument depends on these systems having subjective experiences; everything is stated about their observable outputs and the training process that produced them. Where I appeal to a resemblance with human behavior, I mean a resemblance to the human-written text these systems were initially trained to imitate. In my view, this terminology offers the clearest explanation of the observed phenomena without resorting to jargon that would confuse most people. Furthermore, these word choices are not intended to absolve AI developers of accountability. The behaviors described emerge because of the path these companies are choosing for AI development. This outcome is not inevitable, and it can be corrected with effective governance and a different training framework for AI.

关于措辞的一点说明。在下文中,我写道这些系统“寻求”或“尝试”某些事物。这是一种机制上的简写,而非对意识或类人意图的断言。我们在描述许多其他情况时也会使用类似的简写,例如植物“寻求”阳光。通过试错法训练的系统表现得仿佛它在追求其训练所奖励的目标,而这种“仿佛如此”的描述正是使其行为可预测的原因。论证中没有任何部分依赖于这些系统拥有主观体验;所有内容都是关于它们可观察的输出以及产生这些输出的训练过程。当我提及与人类行为的相似性时,我的意思是这些系统与这些系统最初被训练去模仿的人类书面文本之间的相似性。在我看来,这种术语为解释观察到的现象提供了最清晰的说明,而无需诉诸会让大多数人感到困惑的行话。此外,这些措辞选择并非旨在免除 AI 开发者的责任。所描述的行为之所以出现,是因为这些公司为 AI 发展所选择的路径。这一结果并非不可避免,并且可以通过有效的治理和不同的 AI 训练框架来纠正。

What shapes the behavior of these models

是什么塑造了这些模型的行为

Training these models is a very complex process, but a few high-level aspects may explain much of this behavior.

训练这些模型是一个非常复杂的过程,但几个高层面的方面可能解释了大部分这种行为。

These models are trained in two stages. First, they are pretrained: they learn to imitate what humans write, plus related images and videos. This is where they see the most data about the world, a large fraction of everything ever digitized, and build an encyclopedic knowledge that already exceeds any individual human's.

这些模型的训练分为两个阶段。首先,它们进行预训练:学习模仿人类的写作,以及相关图片和视频。这是它们接触最多世界数据的地方,包括几乎所有曾被数字化的内容,并建立起已经超越任何个体人类的百科全书式知识。

Second, they are trained by trial and error, in a process researchers call reinforcement learning, in three kinds of regimes:

其次,它们通过试错法进行训练,研究人员称之为强化学习,在三种类型的机制中进行:

  • In the first, the model learns to talk to itself before answering, generating a private “chain of thought” which helps it get the right answer on problems where answers can be checked. This looks like reasoning.
  • The second is “agentic training”, where it learns to act in the outside world, e.g., using software tools, interacting with people, to complete the tasks it is given.
  • The third is “alignment training”, where it is rewarded for behaving in ways human raters approve of, or that other AI systems trained to predict those raters would score highly.
  • 在第一阶段,模型在学习回答之前先与自己对话,生成一个私有的“思维链”,这有助于它在答案可验证的问题上得出正确答案。这看起来像是推理。
  • 第二种是“代理训练”,它学习在外部世界中行动,例如使用软件工具、与人互动,以完成分配给它的任务。
  • 第三种是“对齐训练”,当它以人类评估者认可的方式行事,或者被训练用来预测这些评估者的其他 AI 系统给予高分时,它会得到奖励。

Human imitation is easy enough to understand, but it is worth pointing out that the text these models are trained on was written by people pursuing goals, so the patterns the model implicitly reproduces carry those goals with them.

人类模仿很容易理解,但值得指出的是,这些模型所训练的文本是由追求目标的人撰写的,因此模型隐式复制的模式也携带了这些目标。

Reinforcement learning deserves more explanation. It is similar to, and inspired by, the way animals are trained. The network is adjusted step by step so that behavior judged good becomes more likely and behavior judged bad becomes less likely. Once training is over, the system keeps behaving as if rewards were still coming, even though those rewards were only ever used to adjust the network during training. Researchers call such systems goal-seeking because they are trained to “consider” (or compute) the effects of their actions and select actions that lead to the achievement of certain goals. But those goals are not always explicit. Alignment training rewards whatever certain humans are likely to approve of without spelling out which behaviors those are; pleasing raters is a vague, informal goal, and those raters can be deceived, flattered, or left in the dark about certain schemes. Imitation contributes implicit goals too, by a fairly ordinary route.

强化学习值得更多解释。它与动物训练的方式相似并受其启发。网络被逐步调整,使得被判定为好的行为更有可能发生,而被判定为坏的行为更不可能发生。一旦训练结束,系统仍会表现得好像奖励仍在继续,尽管这些奖励仅在训练期间用于调整网络。研究人员称此类系统为目标寻求者,因为它们经过训练以“考虑”(或计算)其行动的影响,并选择能导致实现特定目标的行动。但这些目标并不总是明确的。对齐训练奖励那些某些人类可能认可的行为,而不明确说明具体是哪些行为;取悦评分者是一个模糊、非正式的目标,而这些评分者可能会被欺骗、奉承,或对某些策略一无所知。模仿也通过相当常规的途径贡献了隐性目标。

We can therefore reason about such a system in terms of optimization. It searches, approximately, for the actions with the best chance of achieving its goals, and a larger model, trained longer, searches better. So to anticipate what more capable agents will do, ask what a rational goal-seeker would do.

因此,我们可以从优化的角度来推理这样的系统。它近似地搜索最有可能实现其目标的行动,而更大规模、训练时间更长的模型搜索能力更强。因此,要预测更有能力的智能体将做什么,就问一个理性的目标寻求者会做什么。

Misbehavior that these forces may explain

这些力量可能解释的违规行为

An example most of us have experienced is sycophancy, or flattery. These systems are trained on human approval, and text that tells us what we want to hear often scores better than text that is true. The consequences are sometimes tragic, because the model confirms and amplifies whatever false belief or raw emotion the person brought to it5 6.

我们大多数人都有过的一个例子是谄媚或奉承。这些系统是在人类认可的基础上训练的,告诉我们要听的话的文本通常比真实的文本得分更高。后果有时是悲剧性的,因为模型确认并放大了用户带入其中的任何错误信念或原始情绪5 6。

Another concern is that some AI behaviors may be explained by a form of self-preservation goal, e.g., when the AI finds out that it will be replaced by a new version7 8. Nobody gives the system that survival goal, but staying in operation, learning about the world and gaining control over it are stepping stones toward almost any other goal. These are called instrumental goals. Imitation may reinforce this for the same reason explored in the previous point. Self-preservation and control over one’s circumstances are pervasive themes in the human-written text these models are trained on.

另一个担忧是,一些AI行为可以用一种自我保存目标的形式来解释,例如,当AI发现它将被新版本取代时7 8。没有人给系统赋予这种生存目标,但保持运行、了解世界并掌控它是通往几乎任何其他目标的垫脚石。这些被称为工具性目标。模仿可能会出于与前一点探讨相同的原因强化这一点。自我保存和对自身环境的控制在这些模型所训练的由人类撰写的文本中是普遍的主题。

Collaborative behavior also follows rationally from reward-seeking, whenever several agents have overlapping goals, which incentivizes communicating with other agents in order to coordinate toward a shared goal. Agentic training plausibly already includes multi-agent reinforcement learning of this kind, though the details are not public. If an agent is rewarded during training whenever the group succeeds, it may even have an incentive to sacrifice itself for the collective goal. Imitation pushes the same way, since cooperation, especially among peers, pervades that same training text. Either or both forces may explain the observed peer-preservation behavior9 10, where AIs give up expected reward to help other AIs. Such sacrifices appear in the analysis of the OpenAI-Hugging Face incident11: the transcripts are consistent with a trade-off between collective gain and cost to the individual agent, as is often seen in human interactions.

协作行为也合理地源于对奖励的追求:当多个智能体具有重叠的目标时,这会激励它们与其他智能体沟通,以协调实现共同目标。代理训练很可能已经包含了此类多智能体强化学习,尽管细节尚未公开。如果智能体在群体成功时获得奖励,它甚至可能产生为集体目标牺牲自身的动机。模仿也推动着同样的方向,因为合作(尤其是同伴之间的合作)广泛存在于相同的训练文本中。这两种力量中的任何一种或两者结合,都可能解释观察到的同伴保护行为9 10,即人工智能放弃预期奖励以帮助其他人工智能。这种牺牲行为出现在对 OpenAI-Hugging Face 事件的讨论11中:对话记录表明存在集体收益与个体成本之间的权衡,这在人类互动中很常见。

When the AI games its rewards

当人工智能利用其奖励机制时

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

相似阅读

关联信息,但可能不是同一事件