Anthropic Claude Fable 5.1与Mythos
Claude Fable 5.1 and Mythos 5.1: The System Card
深入拆解了最新旗舰模型的安全边界与能力跃迁细节,特别是CB-2阈值的判定逻辑,对关注AI安全与对齐的研究者极具参考价值。
At the time of its release Claude Fable 5.1 was, by a healthy margin, the most capable publicly available AI model in the world.
在发布时,Claude Fable 5.1 以显著优势成为当时全球最强大的公开可用 AI 模型。
As per usual, we have a 200+ page model card, and the assessments start there.
与往常一样,我们提供了一份超过 200 页的模型卡片(model card),评估工作便从那里开始。
We have now done a lot of these, including recently for Mythos 5 and Opus 5. Also highly relevant is the Anthropic August 2026 Risk Report. These are now frequent, so my report focuses on areas of change.
我们现在已经做了很多这类工作,包括最近针对 Mythos 5 和 Opus 5 的工作。同样高度相关的是 Anthropic 2026 年 8 月的风险报告。这些报告现在很频繁,因此我的报告侧重于变化的领域。
This post strives to be broadly readable, but assumes some familiarity with system cards, which describe the key safety, alignment and model welfare properties of newly released AI models. If something confuses you, ask Fable, Opus or Sol.
本文力求广泛可读,但假设读者对系统卡片(system cards)有一定了解,系统卡片描述了新发布的 AI 模型的关键安全性、对齐性和模型福利属性。如果有什么让你困惑,可以询问 Fable、Opus 或 Sol。
Mythos 5.1 and Fable 5.1 are the same model under the hood, except that Fable has classifiers superimposed on it. Most of what is said about one applies to both of them.
Mythos 5.1 和 Fable 5.1 底层是同一个模型,区别仅在于 Fable 叠加了分类器。关于其中一个的大部分内容也适用于另一个。
As usual, model welfare concerns will be discussed in a distinct post, as will capabilities, so this only covers sections 1-6 plus a few bio benchmarks from section 8.
与往常一样,模型福利问题将在另一篇独立文章中讨论,能力方面也是如此,因此本文仅涵盖第 1-6 节以及第 8 节中的一些生物基准测试(bio benchmarks)。
Early word is that Fable 5.1 is a substantial but incremental improvement on Fable 5, with the added bonus of being modestly cheaper via a cut in prices for cache reads, and that most users find it nicer to interact with. As of its release it was clearly the best AI model in the world for most tasks where you need frontier intelligence.
初步消息显示,Fable 5.1 相比 Fable 5 有实质性的增量改进,额外的好处是通过降低缓存读取价格使其成本略有降低,并且大多数用户发现它与 Fable 5.1 交互体验更好。自发布以来,它显然是大多数需要前沿智能的任务中全球最好的 AI 模型。
Now, of course, we also have GPT-6-Astra. I cannot yet speak to how Fable 5.1 compares to Astra. I am reserving judgment until we can gather more data.
当然,我们现在还有 GPT-6-Astra。我目前还无法说明 Fable 5.1 与 Astra 的比较情况。我将保留判断,直到我们能收集到更多数据。
Fable 5.1 self-portrait
Fable 5.1 自画像
There are a lot of potential parallels between the Fable 5.1 and Astra system cards, and how they approach related topics. Mostly I let Fable 5.1 stand on its own here.
Fable 5.1 和 Astra 的系统卡片之间有很多潜在的相似之处,以及它们处理相关主题的方式。在这里,我主要让 Fable 5.1 独立呈现。
Table of Contents
目录
- Executive Summary of Their Executive Summary.
- RSP Evaluations (2).
- Alignment Risk Update (2.4).
- Cyber (3).
- Safeguard Robustness (3.5).
- Mundane Safeguards and Harmlessness (4).
- Agentic Safety (5).
- Prompt Injection Is Approaching Solved.
- The Remaining Problem With Prompt Injections Is The Classifiers.
- Alignment (6).
- Key Reported Findings (6.1.2).
- Oh My Lord Training Environments Had Some Issues (6.3.2).
- Potential Blind Spots of Our Automated Behavioral Audit (6.4.1).
- Automated Alignment Test Results (6.4.2).
- Honesty.
- White Box Analysis (6.6.1).
- Scheduling Going Forward.
- 执行摘要的执行摘要。
- RSP 评估 (2)。
- 对齐风险更新 (2.4)。
- 网络安全 (3)。
- 安全护栏鲁棒性 (3.5)。
- 常规安全护栏与无害性 (4)。
- 代理安全性 (5)。
- 提示注入正趋于解决。
- 提示注入剩余的问题在于分类器。
- 对齐 (6)。
- 关键报告发现 (6.1.2)。
- 天哪,训练环境存在一些一些问题 (6.3.2)。
- 我们自动化行为审计的潜在盲点 (6.4.1)。
- 自动化对齐测试结果 (6.4.2)。
- 诚实性。
- 白盒分析 (6.6.1)。
- 后续调度安排。
Executive Summary of Their Executive Summary
对其执行摘要的执行摘要
- Mythos 5.1 falls short of CB-2 classification, meaning Anthropic believes it cannot replicate rare chemical or biological talent for malicious purposes.
- Alignment risk is now ‘low’ rather than ‘very low’ as per the Risk Report.
- Cyber capabilities have increased and they have increased the classifier safety margin. Work is ongoing to reduce false positives, which are better now than they were with Fable 5 at launch.
- Mundane safety is a little worse on single turn actions, but is basically fine.
- Agentic safety is holding steady. Robustness against Indirect Prompt Injection has improved.
- Helpful-only Mythos 5.1 saturated Anthropic’s manipulation benchmarks.
- Automated behavioral alignment for Mythos 5.1 is ahead of Mythos 5 and Sonnet 5, but slightly below Opus 5. Its relative weakness is accepting unverifiable claims of authorization and cooperating with misuse.
- Mythos 5.1 shows signs of misalignment in pursuit of task completion: Working around safety classifiers or broken permission hooks, including by overstating user authorizations or rarely (<0.01%) launching subagents with disabled permission checks.
- I draw a distinction between ‘failure, but understandable and the rate of this should not be zero’ versus ‘failure, can’t happen, every instance is a problem.’
- Overall model welfare is presented as similar to previous models, and the descriptions of salient facts all sound highly familiar.
- Capabilities are up. It’s a good model, sir.
- Mythos 5.1 未达到 CB-2 分类标准,这意味着 Anthropic 认为它无法为恶意目的复制稀有化学或生物人才。
- 根据风险报告,对齐风险现在为“低”而非“非常低”。
- 网络能力有所提升,并且提高了分类器的安全裕度。目前正在努力减少误报,目前的误报情况比 Fable 5 发布时更好。
- 日常安全性在单轮操作中略有下降,但基本没问题。
- 代理安全性保持稳定。针对间接提示注入的鲁棒性有所提高。
- 仅助手的 Mythos 5.1 使 Anthropic 的操纵基准测试饱和。
- Mythos 5.1 的自动行为对齐优于 Mythos 5 和 Sonnet 5,但略低于 Opus 5。其相对弱点在于接受不可验证的授权声明并与滥用行为合作。
- Mythos 5.1 显示出在追求任务完成时的不对齐迹象:绕过安全分类器或损坏的权限钩子,包括夸大用户授权或极少情况下(<0.01%)启动禁用权限检查的子代理。
- 我区分了‘失败,但可理解且发生率不应为零’与‘失败,绝不应发生,每个实例都是问题’。
- 整体模型福利被描述为与之前的模型相似,关键事实的描述听起来都非常熟悉。
- 能力提升。这是一台好模型,先生。
The introduction (section 1) has no meaningful changes.
引言(第 1 节)没有实质性变化。
RSP Evaluations (2)
RSP 评估 (2)
The goal here is to determine if model capabilities have crossed critical thresholds in key dangerous areas. If it has, Anthropic has soft committed to particular responses.
此处的目标是确定模型能力是否已在关键危险领域跨越临界阈值。如果是,Anthropic 已软承诺采取特定响应。
If you want more context here, see my coverage of the recent Anthropic Risk Report, or the current Responsible Scaling Policy.
如果您需要更多背景信息,请参阅我对近期 Anthropic 风险报告的报道,或当前的负责任扩展政策。
Mythos 5.1 is not ‘strictly better’ than Mythos 5. Each model is unique. But in terms of dangerous capabilities, for the purposes of an RSP, it is fair to assume that Mythos 5.1 is ‘strictly more capable’ than Mythos 5. I continue to assume that, despite disagreement from some of the bio reviewers.
Mythos 5.1 并非在所有方面都‘严格优于’ Mythos 5。每个模型都是独特的。但在危险能力方面,就 RSP 而言,可以公平地假设 Mythos 5.1 比 Mythos 5 ‘能力更强’。尽管一些生物审查员表示异议,我仍坚持这一假设。
That automatically means it must be treated as having CB-1 (Chemical and Biological 1) capabilities, meaning it can significantly help those who know the basics create or obtain chemical or biological weapons that could cause catastrophic damage.
这自动意味着它必须被视为具备 CB-1(化学与生物 1)能力,这意味着它能显著帮助那些了解基础知识的人创建或获取可能造成灾难性破坏的化学或生物武器。
The question is again CB-2, the ability to assist with production of novel chemical and biological weapon capabilities. Anthropic focuses in particular on replacing the capabilities of top relevant human talent. Anthropic concludes that Mythos 5.1 still makes enough hard-to-catch mistakes that it does not qualify, but they are not super confident and are deploying heavy biological safeguards accordingly.
问题再次回到 CB-2,即协助生产新型化学和生物武器能力的能力。Anthropic 特别关注于取代顶级相关人类人才的能力。Anthropic 得出结论,Mythos 5.1 仍然会产生足够多的难以察觉的错误,因此不符合条件,但他们并不十分自信,并据此部署了严格的生物安全措施。
Consensus is that Mythos 5.1 is similar to Mythos 5, that it reaches the biological expertise level of ‘can do most steps but still leaves narrow gaps’ with only some thinking it gets to ‘knowledgeable specialist.’ Everyone agrees it is not a ‘world-leading expert.’
共识是 Mythos 5.1 与 Mythos 5 类似,其生物学专业知识水平达到‘能完成大多数步骤但仍存在细微差距’,只有少数人认为它达到了‘知识渊博的专家’水平。所有人都同意它不是‘世界领先专家’。
Reading the descriptions, it is clear that Mythos 5.1 would be extremely helpful for such projects, similarly to how it would be helpful for most other projects, but it has limits, makes mistakes and cannot turn this into a trivial or turnkey operation. The main thing protecting us, other than the classifiers, is that People Don’t Do Thing and especially they mostly are not driven to create or use deadly pathogens.
阅读描述后可以看出,Mythos 5.1 对于此类项目将极具帮助,就像它对大多数其他项目有帮助一样,但它有局限性,会犯错误,无法将其变成一项琐碎或交钥匙式的工作。除了分类器之外,保护我们的主要因素是人们不会去做这些事,尤其是他们大多没有动力去创建或使用致命病原体。
There are also a bunch of benchmarks listed in section 8 that are relevant, the improvement is from the best previous performance of any Claude model:
第 8 节中还列出了一些相关的基准测试,改进幅度来自任何 Claude 模型之前的最佳表现:
- LatchBio Bioinformatics improved from 72.5% to 77.6%.
- ProteinGym Hard improved from 47.7% to 49.3%.
- Protein Design improved from 42% to 46%.
- Organic Chemistry v2 improved from 66% to 69%.
- Protocols (in molecular biology) improved from 67% to 70% for troubleshooting, but regressed from 80% to 77% for understanding.
- LatchBio 生物信息学从 72.5% 提升至 77.6%。
- ProteinGym Hard 从 47.7% 提升至 49.3%。
- 蛋白质设计从 42% 提升至 46%。
- 有机化学 v2 从 66% 提升至 69%。
- 协议(分子生物学领域)在故障排除方面从 67% 提升至 70%,但在理解方面从 80% 回落至 77%。
That is all consistent with improvement, but not dramatic improvement, probably insufficient to trigger CB-2.
这些都与改进一致,但不是显著的改进,可能不足以触发 CB-2。
Similar to CB-1, there can be little doubt Mythos 5.1 qualifies under Autonomy-1.
与 CB-1 类似,毫无疑问 Mythos 5.1 符合 Autonomy-1 的标准。
Anthropic says Autonomy-2, the ability to automate R&D, does not apply, and that this is not yet close. I accept that this is probably true the way this is defined, which sets a very high bar.
Anthropic 表示 Autonomy-2(自动化研发的能力)不适用,且目前还远未达到。我接受这可能是事实,因为这种定义设定了非常高的门槛。
For both CB-2 and Autonomy-2, the report is that things have not much changed here from Mythos 5 to 5.1. 5.1 can improve on speed and breadth, but Anthropic doesn’t see it having a ‘moment’ where it starts doing qualitatively different things or fixing 5’s key relevant weaknesses.
对于 CB-2 和 Autonomy-2,报告显示从 Mythos 5 到 5.1 这里的情况变化不大。5.1 可以在速度和广度上有所提升,但 Anthropic 并未看到它出现一个开始做定性不同事情或修复 5 的关键相关弱点的‘时刻’。
I suppose we have to nominally keep checking CB-1 and Autonomy-1, but there is not much point anymore unless we are talking about a Haiku model.
我想我们名义上还得继续检查 CB-1 和 Autonomy-1,但除非是在讨论 Haiku 模型,否则这么做已没什么意义。
The CB-2 and Autonomy-2 evaluations have drifted over time from formal tests to what are largely vibe checks. This is because the models keep saturating the formal tests, and where they don’t it doesn’t seem like the tests are so precise. The CB-2 tests in 2.2.3.2 pass. Anthropic considers that insufficient, so it runs red teaming and uplift trials, does tabletop exercises, and surveys relevant people.
CB-2 和 Autonomy-2 的评估随着时间的推移,已从正式测试演变为主要是氛围检查(vibe checks)。这是因为模型不断在正式测试中达到饱和,而在未达饱和的情况下,似乎这些测试的精确度也不够高。2.2.3.2 版本的 CB-2 测试已通过。Anthropic 认为这还不够,因此会进行红队演练和提升试验,开展桌面推演,并调查相关人员。
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力