Anthropic Claude Fable 5.1:降价、零数据留存与能力评测
Claude Mythos 5.1 and Fable 5.1: Capabilities
Fable 5.1在缓存定价和零数据留存上的调整直接改变企业使用成本与合规门槛,加上Agent能力的实质性增强,值得开发者关注其实际落地效果。
This is the weirdest situation in which to write a capabilities review.
撰写这篇能力评测,处境可谓最怪异不过。
Introducing the world’s most powerful model, by a substantial margin. No wait, this just in, we also have someone else introducing the world’s most powerful model.
隆重推出全球最强大的模型,优势显著。等等,最新消息,我们还有另一位在介绍全球最强大的模型。
Claude Fable 5.1 and GPT-6 Astra are both excellent models. This much, we know.
Claude Fable 5.1 和 GPT-6 Astra 都是优秀的模型。这一点毋庸置疑。
Fable 5.1 comes with reduced cache prices, the option of zero data retention and substantially more lenient classifiers than Fable 5.
Fable 5.1 带来了降低的缓存价格、零数据保留选项,以及比 Fable 5 宽松得多的分类器。
Early signs are, with large error bars, that the jump from Sol to Astra is bigger and more exciting than the jump from Fable 5 to Fable 5.1. This may be similar to how the scaling move from Opus to Fable was a big deal.
初步迹象(尽管误差范围较大)表明,从 Sol 到 Astra 的飞跃比从 Fable 5 到 Fable 5.1 的飞跃更大、更令人兴奋。这可能与从 Opus 到 Fable 的扩展性提升曾引起轰动类似。
With the exception of token use, Fable 5.1 got almost universally positive feedback in absolute terms. Reports are that Fable 5.1 is highly well-rounded. Writing is greatly improved. The Claudisms seem to have improved, although some are very much still there. It admits mistakes. People enjoy their conversations. Several people noted it simplifies code. The safety classifiers are less obnoxious.
除 token 使用量外,Fable 5.1 在绝对意义上几乎获得了普遍好评。据报告,Fable 5.1 表现非常全面。写作能力大幅提升。Claudisms(Claude 特有的表达习惯)似乎有所改善,尽管其中一些仍然存在。它承认错误。人们享受与它的对话。多人指出它能简化代码。安全分类器不再那么令人反感。
Fable 5.1 loves being proactive and doing all the things. If you give it a high effort level, it will find things to do with those tokens. Often they will be useful things.
Fable 5.1 喜欢主动出击并做所有事情。如果你给它设定高努力级别,它会利用这些 token 寻找做事的机会。通常它们都是有用的事情。
My own experience has been that Fable 5.1 and Astra are both excellent. In the one case I’ve had the chance to compare responses on a hard question, both answers were different and very good, complementing each other. My plan is to ‘dual wield’ and ask both all non-trivial queries, at least for a while.
我自己的体验是,Fable 5.1 和 Astra 都表现出色。在我有幸对比两者对难题的回答时,两个答案各不相同且都非常优秀,彼此互补。我的计划是‘双持’它们,至少在一段时间内,对所有非琐碎查询同时提问。
Mostly this review of necessity looks at Fable 5.1 in isolation, rather than being able to compare it to Astra, and most of those who react are doing likewise. There will be full coverage of Astra in turn, which will give more opportunity for comparisons.
由于缺乏与 Astra 进行直接比较的机会,这篇评测不得不主要孤立地审视 Fable 5.1,大多数做出反应的人也是如此。随后将对 Astra 进行全面报道,这将提供更多比较的机会。
Concept by Fable 5.1, executed by Gemini
概念由 Fable 5.1 构思,执行由 Gemini 完成
Table of Contents
目录
- The Official Pitch.
- Our Price Cheap.
- Zero Data Retention and Reduced Safeguards.
- Official Benchmarks.
- Other People’s Benchmarks.
- The System Prompt.
- The Blurb Pitches.
- The Every Review Is In and It’s Very Good.
- Positive Reactions.
- Our Price Cheap But Only Per Token.
- Negative Reactions.
- Early Whispers.
- Weapon of Choice.
- 官方宣传语。
- 我们的价格低廉。
- 零数据保留与降低的安全防护。
- 官方基准测试。
- 他人的基准测试。
- 系统提示词。
- 简介宣传语。
- 所有评测均已出炉,且评价极高。
- 正面反馈。
- 我们的价格低廉,但仅按 token 计费。
- 负面反馈。
- 早期传闻。
- 首选武器。
The Official Pitch
官方发布
As usual, Anthropic basically said ‘here is the new model with the higher number.’
和往常一样,Anthropic 基本上是说“这是带有更高编号的新模型。”
The coding department thinks it is very good at coding, also Our Price Cheap.
编程部门认为它在编程方面非常出色,而且我们的价格很便宜。
Boris Cherny (Claude Code Creator, Anthropic): Fable 5.1 is our best model yet for coding, data analysis, computer use, design, presentations, Tag, and the hardest long-running agentic work.
Boris Cherny(Claude Code 创始人,Anthropic):Fable 5.1 是我们迄今为止最好的编程、数据分析、计算机使用、设计、演示文稿、Tag 以及最困难的长期代理工作模型。
This model is a pleasure to work with, and I’ve been using it for everything.
这个模型用起来很愉快,我一直在用它处理所有事情。
We have also reduced prices for Enterprise, API, and SDK customers. Cache reads on Fable 5.1 are now $0.25 per million tokens (previously: $1). Up to 38% cheaper for a typical Claude Code session.
我们还降低了企业客户、API 和 SDK 客户的定价。Fable 5.1 的缓存读取现在为每百万 token 0.25 美元(之前为 1 美元)。典型的 Claude Code 会话成本降低高达 38%。
We have been working on how often safeguards intervene. Our latest biology safeguards intervene on benign requests 85% less often than the ones we shipped with Fable 5, and Claude Code users should see around 60% fewer cyber interventions per session. Expect more improvements soon.
我们一直在研究安全机制介入的频率。我们最新推出的生物安全机制对良性请求的介入频率比随 Fable 5 发布的版本低 85%,而 Claude Code 用户每个会话的网络安全干预次数预计减少约 60%。预计很快会有更多改进。
Last but not least, Fable 5.1 writes better and has better tone. We heard your feedback, and are actively working on reducing Claude-speak. Solid progress with 5.1, more to come.
最后但同样重要的是,Fable 5.1 写作更好,语气更佳。我们听到了你们的反馈,并正在积极努力减少“Claude 腔调”。在 5.1 中取得了扎实进展,未来还会有更多改进。
Alex Albert highlights its ability to generate videos through code. His example is he took a picture of a property lot, and told Fable to design a house, render it, and produce a cinematic walkthrough, short video at link.
Alex Albert 强调了其通过代码生成视频的能力。他的例子是拍了一张房产地块的照片,然后告诉 Fable 设计一栋房子,渲染它,并制作一段电影般的漫游短片,视频链接见此处。
They spend a section on scientific research. That’s also an OpenAI point of emphasis.
他们花了一节篇幅讨论科学研究。这也是 OpenAI 强调的重点。
Our Price Cheap
我们的价格很便宜
Anthropic is doing a strange thing with Fable 5.1 pricing. The headline price is unchanged from Fable 5, but the effective price is lower because they are reducing the pricing on cache reads from $1 to $0.25 per million tokens.
Anthropic 在 Fable 5.1 的定价上做了一件奇怪的事情。标题价格与 Fable 5 保持不变,但由于将缓存读取的定价从每百万 token 1 美元降至 0.25 美元,实际价格更低。
This will lower typical per-token costs by 25% and highly agentic work costs by ‘up to approximately 45%,’ consistent with Boris’s 38% cheaper for typical Claude Code use:
这将使典型每 token 成本降低 25%,高度代理工作的成本降低‘高达约 45%,’这与 Boris 提到的典型 Claude Code 使用成本降低 38% 一致:
I presume this aligns incentives, by lining up with actual costs, but I would have been inclined to put some of the discount in headline costs instead. People are simple creatures, you have to talk to them on their level sometimes.
我推测这通过将激励措施与实际成本对齐来发挥作用,但我本来倾向于将部分折扣体现在标题成本中。人们是简单的生物,有时你必须用他们的语言与他们沟通。
LLM Index: Claude Fable 5.1 is live on OpenRouter.
LLM Index:Claude Fable 5.1 已在 OpenRouter 上线。
Anthropic’s new model arrives at $10.00/M input, $50.00/M output.
Anthropic 的新模型以输入每百万 token 10.00 美元、输出每百万 token 50.00 美元的价格到来。
No predecessor price or history hook attached. The clean read is the listing itself: a premium Claude shelf entry, now priced on OpenRouter.
没有前代价格或历史挂钩。清晰的解读就是列表本身:一个高端 Claude 货架条目,现在在 OpenRouter 上定价。
Zero Data Retention and Reduced Safeguards
零数据保留和安全机制减少
Fable 5 never got above about 11% of Anthropic dollar spend on Ramp, despite being the clearly best model out there.
尽管 Fable 5 显然是目前最好的模型,但它从未在 Ramp 上 Anthropic 的美元支出中超过约 11%。
Two of the big objections were the safeguards having a large blast radius that stopped ordinary work, and companies, often due to various regulatory concerns, failing to abide Anthropic’s data retention policy, where they required records be kept for 30 days.
两个主要的反对意见是:安全措施具有较大的爆炸半径,会中断普通工作;以及公司往往因各种监管问题,未能遵守 Anthropic 的数据保留政策,该政策要求记录保存 30 天。
The individual Dude, who could abide, had huge coding edge.
那些能够遵守规定的个人用户,在编码方面拥有巨大的优势。
Anthropic has heard you, and had some time to get more comfortable, and for the government to calm down versus when it was so nervous it forced Fable offline over a supposed ‘jailbreak.’ Anthropic is working on a new system that will allow ‘eligible customers’ to use Fable 5.1 with zero outside data retention, via letting the customer store the data. Until then, full zero data retention is available for those customers.
Anthropic 听到了你们的声音,并且有了一些时间变得更加从容,政府也冷静了下来,不像之前那样紧张到因为所谓的‘越狱’而迫使 Fable 下线。Anthropic 正在开发一个新系统,允许‘符合条件的客户’通过让客户自行存储数据的方式,使用零外部数据保留的 Fable 5.1。在此之前,对于那些客户来说,完全零数据保留是可用的。
They’ve greatly reduced the false positive rate on the classifiers (at least 60%).
它们已将分类器的误报率大幅降低(至少 60%)。
That is too many changes at once, including model improvements and lower prices, but will be a fascinating natural experiment. We should see massive adoption of Fable 5.1 within the Anthropic ecosystem, far more than the 11% for Fable 5. If we do not, then people really are purely balking at the headline price without thinking that through.
这包括模型改进和降价在内的太多变化同时发生,但这将是一个迷人的自然实验。我们应该看到 Fable 5.1 在 Anthropic 生态系统内被大规模采用,远超 Fable 5 的 11%。如果没有做到这一点,那么人们确实只是在看到标题价格时纯粹地退缩,而没有深入思考。
Official Benchmarks
官方基准测试
Anthropic shares a ton of benchmarks. Mostly we see modest improvement from Fable 5 or Opus 5 to Fable 5.1, with some small regressions. I’m listing them so you can get a quick gestalt. There does not seem to be a clear pattern.
Anthropic 分享了大量的基准测试。我们主要看到从 Fable 5 或 Opus 5 到 Fable 5.1 的适度提升,但也有一些小幅倒退。我列出这些数据是为了让你快速了解整体情况。似乎没有明显的模式。
Fable 5.1 came out prior to GPT-6, so here is where things stood at that time.
Fable 5.1 在 GPT-6 发布之前推出,因此以下是当时的情形。
All ‘slash line’ scores I give, e.g. 50%/70%, refer to scores without and then with tools.
我给出的所有‘斜杠线’分数,例如 50%/70%,分别指不使用工具和使用工具时的分数。
All improvement or regression numbers by default are versus the best score of either Opus 5 or Mythos 5.
所有默认的提升或倒退数值都是与 Opus 5 或 Mythos 5 中的最佳分数相比。
All numbers are rounded to the given significant figures, as per my judgment.
所有数字均根据我的判断四舍五入到给定的有效数字位数。
Life sciences evals overall show modest improvement and are covered in post one.
生命科学评估总体显示适度提升,并在第一篇帖子中涵盖。
Parenthesis on Terminal Bench 4.0 is for Mythos, other scores are for Fable.
Terminal Bench 4.0 括号内的分数是针对 Mythos 的,其他分数则是针对 Fable 的。
I have added Astra to the chart, based on what we could find.
基于我们能找到的信息,我已将 Astra 添加到图表中。
The Anthropic ECI score is 162.0, versus 159.5 for Mythos 5 and 160.7 for Opus 5, exactly on the Mythos-era trend line.
Anthropic ECI 分数为 162.0,而 Mythos 5 为 159.5,Opus 5 为 160.7,正好处于 Mythos 时代趋势线上。
DeepSWE v1.1 score was 67.4% over 5 trials, but with no frame of reference.
DeepSWE v1.1 分数在 5 次试验中为 67.4%,但缺乏参考框架。
FrontierCode 1.1 Extended scores are worse at higher effort levels. Anthropic attributes that to Fable 5.1 being unable to stop itself from making additional helpful edits at higher effort levels, which get it marked as incorrect.
FrontierCode 1.1 Extended 在高努力等级下的分数更差。Anthropic 将其归因于 Fable 5.1 无法阻止自己在高努力等级下进行额外的有益编辑,从而导致被标记为错误。
FrontierSWE v2 score was 0.57 versus 0.52 for Opus 5, 0.48 for Fable 5 and 0.32 for Sol.
FrontierSWE v2 分数为 0.57,而 Opus 5 为 0.52,Fable 5 为 0.48,Sol 为 0.32。
Terminal-Bench-Science 0.1 was a big jump to 52.6% from 24.7%.
Terminal-Bench-Science 0.1 从 24.7% 大幅跃升至 52.6%。
CursorBench 3.2 maxed out at 73.4% versus a max out of 70.5% for Fable 5, at a substantially lower total cost. Sol maxes out at 67.2%.
CursorBench 3.2 最高达到 73.4%,而 Fable 5 的最高分为 70.5%,且总成本显著更低。Sol 的最高分为 67.2%。
CritPT-Corrected slightly improved from 85.5% to 88.4%.
CritPT-Corrected 略有提升,从 85.5% 增至 88.4%。
ArXivMath slightly improved from 91%/91% to 91%/94%.
ArXivMath 略有提升,从 91%/91% 变为 91%/94%。
ProgramBench improved slightly from 86.3% to 87.6%.
ProgramBench 略有提升,从 86.3% 增至 87.6%。
As per the chart, Humanity’s Last Exam improved from 57.8%/63.8% to 60.9%/65%.
如图所示,Humanity’s Last Exam 从 57.8%/63.8% 提升至 60.9%/65%。
They run various multi-agent tests, but don’t provide good points of comparison, so I can’t tell how good the results are.
他们运行了各种多智能体测试,但缺乏良好的对比基准,因此我无法判断结果的好坏。
Chartography improved from 37%/84% to 43%/86%.
Chartography 从 37%/84% 提升至 43%/86%。
BenchCAD Vision2Code improved from 38%/67% to 44%/84%.
BenchCAD Vision2Code 从 38%/67% 提升至 44%/84%。
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力