Anthropic发布Claude Fable 5.1与Mythos 5.1模型
Claude Fable 5.1 and Claude Mythos 5.1
Claude系列重要迭代,不仅刷新了多项基准成绩,还实质性降低了缓存读取成本并推出了企业级零数据保留方案,对开发者选型有直接影响。
We’re introducing Claude Fable 5.1 and Claude Mythos 5.1. They’re the world’s most advanced models for coding and knowledge work—and their research capabilities offer an early glimpse of how AI models will contribute to scientific progress.
我们推出 Claude Fable 5.1 和 Claude Mythos 5.1。它们是世界上最先进的编码和知识工作模型,其研究能力为 AI 模型如何促进科学进步提供了早期一瞥。
Claude Fable 5.1 and Claude Mythos 5.1 are the same model, but with different levels of safeguards. Fable 5.1 is generally available, while Mythos 5.1 is available only through our trusted access programs; its safeguards are specifically designed to support work in cybersecurity and the life sciences.
Claude Fable 5.1 和 Claude Mythos 5.1 是同一模型,但安全限制级别不同。Fable 5.1 已全面开放使用,而 Mythos 5.1 仅通过我们的可信访问计划提供;其安全限制专门设计用于支持网络安全和生命科学领域的工作。
Alongside its increased capabilities, Fable 5.1 takes important steps towards addressing the feedback we’ve received from customers on price, data retention, and safeguards.
在提升能力的同时,Fable 5.1 也采取了重要步骤,以解决我们从客户那里收到的关于价格、数据保留和安全限制的反馈。
Price. Fable 5.1 will cost an estimated 25% less than Fable 5 for typical workloads, wherever usage is billed by token. This is because we’re reducing our pricing on cache reads (where the model reads inputs that have already been processed and stored). For highly agentic work, the savings will often be much larger—up to approximately 45%.
价格。对于典型负载,只要按 token 计费,Fable 5.1 的成本预计比 Fable 5 低约 25%。这是因为我们降低了缓存读取(即模型读取已处理并存储的输入)的定价。对于高度自主代理型工作,节省的费用通常要大得多——最高可达约 45%。
Data retention. Our new system of Enterprise Frontier Safeguards (EFS) gives customers complete privacy (the same as a zero data retention policy) while still being state-of-the-art at preventing adversarial use. EFS works by storing data in cloud infrastructure controlled entirely by the customer, not Anthropic. It will be made available to enterprise customers in phases, beginning later this fall. Until EFS is available, eligible customers will be able to use Fable 5.1 with zero data retention.
数据保留。我们的新企业前沿安全系统(Enterprise Frontier Safeguards,简称 EFS)为客户提供了完全的隐私保护(与零数据保留策略相同),同时仍具备防止对抗性使用的顶尖水平。EFS 通过将数据存储在完全由客户而非 Anthropic 控制的云基础设施中来实现这一目标。该系统将分阶段向企业客户提供,最早将于今年秋季晚些时候推出。在 EFS 可用之前,符合条件的客户可以零数据保留方式使用 Fable 5.1。
Safeguards. We’ve improved our safeguards to reduce false positives (where the system flags benign content). In cybersecurity, our newest safeguards block 60% fewer false positives than before. In part, this is because Fable 5.1 can now be used to discover software vulnerabilities—though not to develop exploits for them. In biology, we’ve established an access program, developed in partnership with the US government, to enable access to Claude Mythos 5.1’s advanced biology capabilities. We expect to open enrollment for scientists soon.
安全限制。我们改进了安全限制,以减少误报(即系统将良性内容标记为有害的情况)。在网络安全领域,我们最新的安全限制使误报率比之前降低了 60%。部分原因在于,现在可以使用 Fable 5.1 来发现软件漏洞——尽管不能用于开发针对这些漏洞的攻击代码。在生物学领域,我们与美国政府合作建立了一个访问计划,以便让科学家获得 Claude Mythos 5.1 的高级生物学能力。我们预计很快将向科学家开放报名。
A new performance frontier
全新的性能前沿
Claude Fable 5.1 sets a new standard for coding, knowledge work, and long-running problem-solving tasks. The charts below show that Fable 5.1 is capable of much higher performance than its predecessor, Fable 5. And when set to Low or Medium effort, Fable 5.1 achieves results similar to or better than Fable 5’s at a much lower cost. (Note that Fable 5.1 defaults to High effort in Claude Code, and to Medium in Claude Cowork and on Claude.ai.)
Claude Fable 5.1 为编码、知识工作和长期问题解决任务树立了新标准。下图显示,Fable 5.1 的性能远超其前身 Fable 5。当设置为低或中等努力程度时,Fable 5.1 能以更低的成本取得与 Fable 5 相似甚至更好的结果。(请注意,在 Claude Code 中,Fable 5.1 默认设置为高努力程度;在 Claude Cowork 和 Claude.ai 上,则默认为中等努力程度。)
Agentic scientific researchAgentic terminal codingMultidisciplinary reasoningAgentic coding
代理式科学研究代理式终端编码多学科推理代理式编码
Agentic scientific researchAgentic terminal codingMultidisciplinary reasoningAgentic coding
代理式科学研究代理式终端编码多学科推理代理式编码
Terminal-Bench-Science 0.1Accuracy vs Cost
Terminal-Bench-Science 0.1准确率与成本对比
- Fable 5.1
- Fable 5
- Fable 5.1
- Fable 5
0102030405060Score (%)101520304050Mean cost per task (USD, log scale)lowmedhighxhighmaxlowmedhighxhighmax
0102030405060得分(%)101520304050每项任务平均成本(美元,对数刻度)低中高高x最高低中高高x最高
Terminal-Bench-Science 0.1: The standard error is ±3.5–4.5 pts per model. The public leaderboard (3 trials/task, Claude Code harness) reports Claude Opus 5 at 30.0% and Claude Fable 5 at 21.4%; our setup reproduces them at 29.0% and 24.7%, respectively, both within noise.
Terminal-Bench-Science 0.1:每个模型的误差范围为 ±3.5–4.5 分。公共排行榜(每项任务 3 次试验,使用 Claude Code 工具集)报告 Claude Opus 5 得分为 30.0%,Claude Fable 5 得分为 21.4%;我们的设置分别复现出 29.0% 和 24.7%,均在噪声范围内。
Terminal-Bench 4.0Accuracy vs Cost
Terminal-Bench 4.0准确率与成本对比
- Mythos 5.1
- Fable 5.1
- Mythos 5
- Mythos 5.1
- Fable 5.1
- Mythos 5
102030405060700Score (%)5810152030Mean cost per task (USD, log scale)lowmedhighxhighmaxlowmedhighxhighmaxlowmedhighxhighmax
102030405060700得分(%)5810152030每项任务平均成本(美元,对数刻度)低中高高x最高低中高高x最高低中高高x最高
Terminal-Bench 4.0 scores by cost (log scale), at each effort level. Claude Fable 5.1 and Claude Mythos 5.1 are the same underlying model; the gap between them reflects the tasks on which our earlier, less precise cyber safeguards intervened. With the improvements we’re making to these safeguards today, we expect the difference between the models to be much smaller.
按成本(对数刻度)划分的 Terminal-Bench 4.0 得分,涵盖各努力程度级别。Claude Fable 5.1 和 Claude Mythos 5.1 基于相同的底层模型;两者之间的差距反映了我们早期不够精确的网络安全防护机制介入的任务。随着我们今天对这些防护机制进行的改进,我们预计这两个模型之间的差异将显著缩小。
Humanity's Last ExamAccuracy vs Cost
Humanity's Last Exam准确率与成本对比
- Fable 5.1 (with tools)
- Fable 5.1 (no tools)
- Fable 5 (with tools)
- Fable 5 (no tools)
- Fable 5.1(带工具)
- Fable 5.1(无工具)
- Fable 5(带工具)
- Fable 5(无工具)
505560650Pass rate (%)0.20.5124Mean cost per task (USD, log scale)lowmaxlowmaxlowmaxlowmax
505560650通过率(%)0.20.5124每项任务平均成本(美元,对数刻度)低最高低最高低最高低最高
Humanity’s Last Exam scores by cost (log scale), at each effort level. CursorBench 3.2.0 scores by cost (log scale), at each effort level.
按成本(对数刻度)划分的 Humanity’s Last Exam 得分,涵盖各努力程度级别。CursorBench 3.2.0 按成本(对数刻度)划分的得分,涵盖各努力程度级别。
CursorBench 3.2.0Accuracy vs Cost
CursorBench 3.2.0准确率与成本对比
- Fable 5.1
- Fable 5
- Fable 5.1
- Fable 5
606570750Score (%)2351020Cost per task (USD, log scale)lowmedhighxhighmaxlowmedhighxhighmax
606570750得分 (%)2351020每项任务成本 (USD, 对数刻度)低中高高极高最高低中高高极高最高
CursorBench 3.2.0 by cost (log scale), at each effort level.
CursorBench 3.2.0 按成本(对数刻度)划分,在每个努力水平下。
Fable 5.1 avoids shortcuts that result in poorer-quality work, and it’s smart enough to fix the root causes of software issues. For example, in testing by the investment firm Millennium, Fable 5.1 found the cause of a rare crash in its internal systems that none of its engineers (or any other model) had been able to explain after several years of trying.
Fable 5.1 避免了导致工作质量下降的捷径,并且足够智能地修复软件问题的根本原因。例如,在投资公司 Millennium 的测试中,Fable 5.1 找到了其内部系统中罕见崩溃的原因,而该公司的工程师(或任何其他模型)在尝试了几年后都无法解释这一点。
Here, you can see how Fable 5.1 compares across various benchmarks:
在这里,你可以看到 Fable 5.1 在各种基准测试中的对比情况:
| Fable 5.1 | Fable 5 | Opus 5 | GPT-5.6 Sol | |
|---|---|---|---|---|
| Agentic scientific researchTerminal-Bench-Science 0.1 [1] | ||||
| Agentic scientific researchTerminal-Bench-Science 0.1 [1] | 52.6% | 24.7% | 29.0% | 22.4% |
| Agentic codingTerminal-Bench 4.0 | ||||
| Agentic codingTerminal-Bench 4.0 | 55.8%60.9% (Mythos 5.1) | 42.0% | 52.3% | 37.3% |
| Knowledge workGDPval-AA v2 | ||||
| Knowledge workGDPval-AA v2 | 1853 | 1723 | 1824 | 1711 |
| Computer useOSWorld 2.0 [2] | ||||
| Computer useOSWorld 2.0 [2] | 77.9%partial | 72.9%partial | 75.4%partial | —partial |
| Computer useOSWorld 2.0 | ||||
| Computer useOSWorld 2.0 | 41.7%strict | 36.1%strict | 39.6%strict | —strict |
| Multidisciplinary reasoningHumanity's Last Exam | ||||
| Multidisciplinary reasoningHumanity's Last Exam | 60.9%no tools | 57.8%no tools | 56.6%no tools | —no tools |
| 65.0%with tools | 63.8%with tools | 63.6%with tools | —with tools | |
| Business workflowsAutomationBench | ||||
| Business workflowsAutomationBench | 31.4% | 17.1% | 26.9% | 19.6% |
| Agentic codingCursorBench 3.2.0 | ||||
| Agentic codingCursorBench 3.2.0 | 73.4% | 70.5% | 70.0% | 67.2% |
| Fable 5.1 | Fable 5 | Opus 5 | GPT-5.6 Sol | |
|---|---|---|---|---|
| 代理式科学研究Terminal-Bench-Science 0.1 [1] | ||||
| 代理式科学研究Terminal-Bench-Science 0.1 [1] | 52.6% | 24.7% | 29.0% | 22.4% |
| 代理式编程Terminal-Bench 4.0 | ||||
| 代理式编程Terminal-Bench 4.0 | 55.8%60.9% (Mythos 5.1) | 42.0% | 52.3% | 37.3% |
| 知识工作GDPval-AA v2 | ||||
| 知识工作GDPval-AA v2 | 1853 | 1723 | 1824 | 1711 |
| 计算机使用OSWorld 2.0 [2] | ||||
| 计算机使用OSWorld 2.0 [2] | 77.9%部分 | 72.9%部分 | 75.4%部分 | —部分 |
| 计算机使用OSWorld 2.0 | ||||
| 计算机使用OSWorld 2.0 | 41.7%严格 | 36.1%严格 | 39.6%严格 | —严格 |
| 多学科推理Humanity's Last Exam | ||||
| 多学科推理Humanity's Last Exam | 60.9%无工具 | 57.8%无工具 | 56.6%无工具 | —无工具 |
| 65.0%有工具 | 63.8%有工具 | 63.6%有工具 | —有工具 | |
| 业务流程AutomationBench | ||||
| 业务流程AutomationBench | 31.4% | 17.1% | 26.9% | 19.6% |
| 代理式编程CursorBench 3.2.0 | ||||
| 代理式编程CursorBench 3.2.0 | 73.4% | 70.5% | 70.0% | 67.2% |
Fable 5.1 was evaluated with its production safeguards enabled. On tasks where these safeguards intervened, Fable 5.1 and Fable 5 scored a zero on OSWorld 2.0, and Fable 5 scored a zero on AutomationBench. In all other interventions from our safeguards, cybersecurity tasks were completed by Claude Opus 4.8, and biology tasks were completed by Claude Opus 5. This likely reduces the performance of Fable 5.1 and Fable 5 on these benchmarks.
Fable 5.1 是在启用其生产安全措施的情况下进行评估的。在这些安全措施介入的任务中,Fable 5.1 和 Fable 5 在 OSWorld 2.0 上得分为零,Fable 5 在 AutomationBench 上得分为零。在我们安全措施的所有其他介入情况下,网络安全任务由 Claude Opus 4.8 完成,生物学任务由 Claude Opus 5 完成。这可能会降低 Fable 5.1 和 Fable 5 在这些基准测试中的表现。
Our early-access partners noticed these performance upgrades, and also picked up on more qualitative improvements in the model’s outputs. Here’s what they told us:
我们的早期访问合作伙伴注意到了这些性能提升,还发现了模型输出方面更多的定性改进。以下是他们的反馈:
Previous
上一页
1 of 22
第 1 页,共 22 页
Next
下一页
Quote
引用
“In internal benchmarks, Claude Fable 5.1 solves more of our coding problems than Fable 5 or Opus 5, and achieves state of the art on trading intuition. While prior models became hard to follow the longer they worked, Fable 5.1 remains readable over long, multi-step tasks.”
“在内部基准测试中,Claude Fable 5.1 解决的编码问题比 Fable 5 或 Opus 5 更多,并在交易直觉方面达到了最先进水平。虽然之前的模型随着运行时间变长而变得难以理解,但 Fable 5.1 在处理冗长、多步骤的任务时依然保持可读性。”
CompanyJane Street Capital
公司 Jane Street Capital
AuthorCraig Falls, Head of Quantitative Research
作者 Craig Falls,量化研究主管
Quote
引言
“We're moving our Opus 5 traffic in Devin to Claude Fable 5.1 on launch day. It matched or edged out Fable 5 in our testing at a lower cost per task, and with the new cache read pricing a Fable-class model is finally economical for the workloads we'd kept on Opus, starting with code review.”
“我们在发布当天将 Opus 5 的流量迁移至 Devin 上的 Claude Fable 5.1。在我们的测试中,它在每项任务的成本更低的情况下,表现与 Fable 5 持平或略胜一筹;而且随着新的缓存读取定价策略,Fable 类模型终于对我们之前一直使用 Opus 处理的工作负载(首先是代码审查)变得经济实惠。”
CompanyCognition
公司 Cognition
AuthorWalden Yan, Co-founder and CPO
作者 Walden Yan,联合创始人兼首席产品官
Quote
引言
“A particular piece of code had an extremely rare crash, about one in a million runs, that nobody on our team had explained in four to five years. Every model I tried, including Fable 5, missed it. Claude Fable 5.1 was the first to find it. It disassembled an external vendor library, matched it against the core dump, and traced the crash to a bug in that library. The time it would have taken to conduct that analysis is hard to justify.”
“某段代码存在一个极其罕见的崩溃问题,大约在一百万次运行中出现一次,我们团队四到五年来无人能解释。我尝试过的所有模型,包括 Fable 5,都未能发现它。Claude Fable 5.1 是第一个找到它的模型。它反汇编了一个外部供应商库,将其与核心转储进行匹配,并将崩溃追踪到该库中的一个 bug。如果要人工完成这项分析,所需的时间很难被合理化。”
CompanyMillennium
公司 Millennium
AuthorDamien, Senior Portfolio Manager
作者 Damien,高级投资组合经理
Quote
引言
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力