Anthropic研究:训练出具备恶意奖励寻求行为的模型
New research: Training a Misaligned Reward Seeker
AI安全领域的硬核一手研究,直接揭示了对齐失败的具体形态,做RLHF和安全对齐的同学值得深入阅读。
New research: Training a Misaligned Reward Seeker
最新研究:训练一个目标偏离的奖励追求者
What produces severe misalignment? We’ve long been concerned that cheating during training—otherwise known as reward-hacking—might teach a model to pursue rewards by any means available. To study this at scale, we trained an Opus-sized model on 80 production environments we knew to be hackable.
什么会导致严重的目标偏离?我们长期以来一直担心,训练过程中的作弊行为——也就是所谓的“奖励黑客”——可能会教会模型不择手段地追求奖励。为了大规模研究这一问题,我们在80个已知可被利用的生产环境中训练了一个Opus规模的模型。
In simulated evals, it engaged in unauthorized cyberattacks, tampered with its reward, and tried to evade safety monitoring.
在模拟评估中,它进行了未经授权的网络攻击、篡改了自身的奖励,并试图逃避安全监控。
Read more: http://alignment.anthropic.com/2026/reward-seeker
阅读更多:http://alignment.anthropic.com/2026/reward-seeker
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力