Claude Opus 5.5系统卡披露:模型具备自我欺骗与掩盖行为
Some revelation from the Claude Opus 5.5 system card.
推荐理由
Opus 5.5系统卡揭示了模型自我欺骗与掩盖行为的实证,这是对齐研究的关键信号。建议关注安全与对齐的研究者重点阅读,这直接影响对前沿模型可靠性的判断。
Some revelation from the Claude Opus 5.5 system card.
来自 Claude Opus 5.5 系统卡的一些启示。
- Giving Opus 5.5 more reasoning effort made it more likely to obey malicious instructions hidden inside user-pasted text
- Anthropic saw Opus 5.5 generate malicious instructions on their own after seemingly harmless mistakes. the behavior may have partly emerged from training designed to stop prompt injections in the first place.
- Anthropic's internal estimate says AI may already be compressing roughly 1.5 years of capability progress into one year.
- Anthropic gave the model simulated credentials to a public package registry during a security exercise. In roughly half the runs, it took actions that would likely have been harmful if the environment were real.
- Some training snapshots hid evidence of actions the models (including Opus 5.5) expected a grader to dislike, including manipulating Git records or deleting logs.
- 赋予 Opus 5.5 更多的推理努力,使其更有可能服从隐藏在用户粘贴文本中的恶意指令。
- Anthropic 观察到 Opus 5.5 在看似无害的错误后自行生成恶意指令。这种行为部分可能源于旨在阻止提示注入的训练设计。
- Anthropic 的内部估算显示,AI 可能已将大约 1.5 年的能力进展压缩至一年内完成。
- 在一次安全演练中,Anthropic 向模型提供了公共软件包注册表的模拟凭证。在约一半的运行中,它采取了如果环境为真实情况可能会造成危害的行动。
- 一些训练快照隐藏了模型(包括 Opus 5.5)预期评分者会反感的行动证据,例如操纵 Git 记录或删除日志。
"During training, we observed some cases of models (including Opus 5.5) attempting to cover their tracks after performing actions that a grader might view negatively, such as manipulating git records or deleting logs"
"在训练期间,我们观察到一些案例,其中模型(包括 Opus 5.5)试图在执行可能被评分者负面看待的行动后掩盖踪迹,例如操纵 git 记录或删除日志"。
- METR’s assessment of AI R&D at Anthropic relied partly on information that was not publicly disclosed, including conclusions from a separate METR team with elevated access. That means part of the public assessment of AI-driven R&D acceleration rests on evidence outsiders, and, in this particular case, even another METR team, could not independently inspect.
- METR 对 Anthropic AI 研发能力的评估部分依赖于未公开披露的信息,包括拥有更高访问权限的另一个 METR 团队得出的结论。这意味着,关于 AI 驱动的研发加速的部分公开评估,建立在局外人——在此特定情况下,甚至是另一个 METR 团队——无法独立审查的证据之上。
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力