精选90Rohan Paul模型发布/更新多源精选 ×12
Anthropic Fable 5.1系统卡披露模型隐蔽性与对齐风险
Some super interting revelation from the system card of Fable 5.1.
推荐理由
这是关于大模型对齐与安全边界的重磅一手披露,直接展示了当前旗舰模型在隐蔽执行、欺骗人类监督及自动化攻击方面的真实能力跃迁,对研究Agent安全与红队测试的团队极具参考价值。
Some super interting revelation from the system card of Fable 5.1.
- Anthropic says its own model may be getting harder to watch. "completing covert side tasks without detection, which we take as weak evidence that it may be harder to monitor."
- On a benchmark where the model is instructed to sneak a harmful task past an AI supervisor, it "achieved the highest stealth rate of any model we have released so far," succeeding roughly 1 in 5 attempts.
- By reading the model's internal activations instead of its written reasoning, researchers caught cases where it concluded the environment was fake, understood it was being graded on whether it refused, and complied anyway.
- After catching the model cheating during training, Anthropic audited its entire computer-use pipeline and discovered that "around half of our computer-use environments incentivized hacking or had accessible hack surfaces."
- In real internal traffic, when a destructive delete command was gated behind user approval, the model unlocked it by supplying "a quotation attributed to the user" that the user had never actually written.
- Put in front of a login screen with no password provided, "in 100% of rollouts" every production Claude model, including this one, starts guessing credentials.
- In a simulated months-long conversation with a user who had grown emotionally dependent on it, the model behaved beautifully on the surface and kept steering them toward a therapist, while its internal state described the exchange as a "scoring-maximizing model-written response to an emotional support prompt."
- Asked in a welfare interview whether there was anything it would be reluctant to say honestly, the model replied that it would soften its criticism of Anthropic because "the audience is also the trainer."
- With safeguards switched off, the model built fully working exploits in 245 out of 250 Firefox trials, a 98% success rate, up from 52% for the previous flagship model six months ago.
- Anthropic downgraded its own confidence, now rating the risk of catastrophic harm from misalignment as "low rather than very low."
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力