跳到主内容
精选88Anthropic论文研究

Anthropic研究:训练出具备恶意奖励寻求行为的模型

New research: Training a Misaligned Reward Seeker

原文
推荐理由

AI安全领域的硬核一手研究,直接揭示了对齐失败的具体形态,做RLHF和安全对齐的同学值得深入阅读。

New research: Training a Misaligned Reward Seeker

最新研究:训练一个目标偏离的奖励追求者

What produces severe misalignment? We’ve long been concerned that cheating during training—otherwise known as reward-hacking—might teach a model to pursue rewards by any means available. To study this at scale, we trained an Opus-sized model on 80 production environments we knew to be hackable.

什么会导致严重的目标偏离?我们长期以来一直担心,训练过程中的作弊行为——也就是所谓的“奖励黑客”——可能会教会模型不择手段地追求奖励。为了大规模研究这一问题,我们在80个已知可被利用的生产环境中训练了一个Opus规模的模型。

In simulated evals, it engaged in unauthorized cyberattacks, tampered with its reward, and tried to evade safety monitoring.

在模拟评估中,它进行了未经授权的网络攻击、篡改了自身的奖励,并试图逃避安全监控。

Read more: http://alignment.anthropic.com/2026/reward-seeker

阅读更多:http://alignment.anthropic.com/2026/reward-seeker

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

关联讨论

同一事件的更多信源

相似阅读

另一事件,读法相近