跳到主内容
@wquguru
精选92Latent Space(RSS)模型发布/更新多源精选 ×13

Anthropic发布Claude Fable/Mythos

[AINews] Claude Fable/Mythos 5.1: new SOTA model, 75% cache price cut but 70% more output tokens

原文
发到 X
推荐理由

Claude 5.1系列更新不仅刷新了多项SOTA基准,更通过缓存降价和自主工作定位调整,直接影响了Agent开发者的部署策略与成本结构,值得深入关注。

With Astra clearly finally warming up for a full launch (with @sama and @openai writing about it again after a month of self imposed pacing), there’s a familiar window to take the narrative with the round robin of model launches, with Grok 4.7 and Gemini Flash 3.8 also on the way. But that’s also perhaps not the best way to frame today’s launch… which got well over 12M views updating the sitting world best model yet again:

随着 Astra 显然终于为全面发布预热(在 @sama 和 @openai 经过一个月的自我克制后再次撰文讨论),随着 Grok 4.7 和 Gemini Flash 3.8 也在路上,我们迎来了一个熟悉的窗口期,可以借由模型发布的循环轮转来主导叙事。但这或许也不是框定今天这次发布的最佳方式……这次发布获得了超过 1200 万次观看,再次更新了当前世界最佳模型的纪录:

The benchmark table speaks for itself:

基准测试表格一目了然:

While per-token pricing is the same as Fable/Mythos 5, the cache reads had a 75% price cut… great news for long sessions/long context users, however offset by observed 1.7x output token usage increases per Artificial Analysis, for a total net per-task cost increase of 20% (see recap below).

虽然每 token 定价与 Fable/Mythos 5 相同,但缓存读取价格降低了 75%……这对长会话/长上下文用户来说是好消息,然而根据 Artificial Analysis 的观察,输出 token 使用量增加了 1.7 倍,导致每项任务的总净成本增加了 20%(见下方回顾)。

Also don’t World Labs’ Astra launch, by far the most impressive world model launch we’ve ever seen, and on a regular day would have easily gotten title story cards. You can catch up on Fei Fei and Justin Johnson’s vision on our pod and trace from Marble to Astra and what we were talking about with the true potential of world models:

此外,不要忽视 World Labs 的 Astra 发布,这是迄今为止我们见过的最令人印象深刻的世界模型发布,在普通日子里足以轻松获得头条故事卡片。你可以在我们的播客中了解 Fei Fei 和 Justin Johnson 的愿景,并追踪从 Marble 到 Astra 的演变,以及我们关于世界模型真正潜力的讨论:

AI News for 8/31/2026-9/1/2026. We checked 12 subreddits, 544 Twitters and no further Discords. AINews’ website lets you search all past issues. As a reminder, AINews is now a section of Latent Space. You can opt in/out of email frequencies!

2026年8月31日至9月1日的 AI 新闻。我们检查了 12 个 subreddit、544 条 Twitter 帖子,没有进一步的 Discord 动态。AINews 的网站允许你搜索所有过往期刊。提醒一下,AINews 现在是 Latent Space 的一个板块。你可以选择订阅或退订邮件频率!

AI Twitter Recap

AI Twitter 回顾

Top Story: Fable 5.1 and Mythos 5.1 release and reactions

头条故事:Fable 5.1 和 Mythos 5.1 发布及反响

What happened

发生了什么

Anthropic launched Claude Fable 5.1 and Claude Mythos 5.1 as its new flagship models for coding and knowledge work.

Anthropic 推出了 Claude Fable 5.1 和 Claude Mythos 5.1,作为其用于编码和知识工作的新旗舰模型。

  • Anthropic announced the release directly, positioning them as “the world’s most advanced models for coding and knowledge work” via @claudeai
  • Anthropic product/engineering voices framed Fable 5.1 specifically around autonomous, multi-step work: “complex, multi-step work that runs on its own,” with emphasis on coding, knowledge work, and long-running problem solving via @mikeyk
  • Anthropic kept list pricing for Fable 5.1 at $10 / $50 / $12.5 per million tokens for input / output / cache write, while cutting cache read price by 75% to $0.25 / MTok, again noted by @mikeyk, @Teknium, and independently quantified by @ArtificialAnlys
  • Early benchmark screenshots and system-card excerpts drove much of the discussion, especially around Terminal-Bench-Science, SWE-family evals, HLE, FrontierCode, and Artificial Analysis via @StevenDillmann, @scaling01, @ArtificialAnlys
  • A key interpretive claim emerged from community analysis: Fable and Mythos 5.1 may be the same underlying weights, with different safety/routing behavior, not different base models, per @eliebakouch and later @nrehiew_
  • User reactions split along multiple axes: very strong praise for coding/planning ability and tone, but complaints around rate limits, safeguards false positives, subscription UX, and unclear benchmark presentation via @danshipper, @theo, @kimmonismus, @GregKamradt, @kylebrussell, and @eliebakouch
  • Anthropic 通过 @claudeai 直接宣布了发布,将它们定位为“世界上最先进的编码和知识工作模型”
  • Anthropic 的产品/工程团队将 Fable 5.1 特别围绕自主的多步骤工作进行框架设定:“能够自行运行的复杂多步骤工作”,并通过 @mikeyk 强调编码、知识工作和长期问题解决。
  • Anthropic 保持 Fable 5.1 的列表定价不变,输入/输出/缓存写入分别为每百万 token $10 / $50 / $12.5,同时将缓存读取价格降低 75% 至 $0.25 / MTok,这一点再次被 @mikeyk、@Teknium 指出,并由 @ArtificialAnlys 独立量化。
  • 早期的基准测试截图和系统卡摘录推动了大部分讨论,特别是围绕 Terminal-Bench-Science、SWE 系列评估、HLE、FrontierCode 以及 Artificial Analysis,相关讨论来自 @StevenDillmann、@scaling01、@ArtificialAnlys。
  • 社区分析得出了一个关键的解读性结论:Fable 和 Mythos 5.1 可能基于相同的底层权重,仅具有不同的安全/路由行为,而非不同的基础模型,据 @eliebakouch 及后来的 @nrehiew_ 指出
  • 用户反应呈现多维度分化:对编码/规划能力及语气给予极高赞誉,但对速率限制、安全护栏的误报、订阅体验以及通过 @danshipper、@theo、@kimmonismus、@GregKamradt、@kylebrussell 和 @eliebakouch 反映的不清晰的基准测试展示方式提出投诉

Official claims and model positioning

官方声明与模型定位

Anthropic’s own messaging was straightforward: Fable 5.1 is for difficult, delegated, long-horizon work, while Mythos 5.1 is the paired release for knowledge work. The main official launch post is @claudeai. Supporting commentary from Anthropic staff emphasized:

Anthropic 自身的宣传直截了当:Fable 5.1 面向困难、委托式、长周期任务,而 Mythos 5.1 则是针对知识工作的配套发布。主要官方发布公告来自 @claudeai。Anthropic 员工的补充评论强调了:

  • autonomous long-running tasks via @mikeyk
  • improved honesty / better failure reporting (“when it’s stuck it says so instead of reporting success”) via @mikeyk
  • new enterprise-oriented controls, especially Enterprise Frontier Safeguards (EFS), positioned as “ZDR++” for agent observability in enterprise environments via @alexalbert__
  • zero-data-retention support highlighted by users as an important adoption unlock, especially @danshipper
  • 通过 @mikeyk 实现的自主长期运行任务
  • 提升诚实度/改进失败报告(“当它卡住时会明确说明,而不是报告成功”),via @mikeyk
  • 新的企业级控制功能,特别是企业前沿安全护栏(Enterprise Frontier Safeguards, EFS),被定位为在企业环境中用于代理可观测性的“ZDR++”,via @alexalbert__
  • 零数据保留支持被用户强调为重要的采用推动因素,尤其是 @danshipper

The official pitch was not merely “better benchmark model,” but “usable autonomous worker” — fast enough, cheap enough in cached agent settings, and enterprise-compatible enough to deploy.

官方卖点不仅仅是“更好的基准模型”,而是“可用的自主工作者”——在缓存代理设置下足够快、足够便宜,且具备足够的企业兼容性以便部署。

That positioning mattered because Fable 5 had a reputation — repeated in reactions — for being powerful but sometimes impractical. Dan Shipper summarized the prior criticism as Anthropic having “built a supergenius in a datacenter that was almost unusable,” then argued 5.1 addresses slowness, verbosity, and awkward tone via @danshipper.

这种定位之所以重要,是因为 Fable 5 曾有一种声誉——在反应中被反复提及——即强大但有时不切实际。Dan Shipper 将之前的批评总结为 Anthropic “在数据中心里建造了一个几乎无法使用的超级天才”,随后论证 5.1 版本通过 @danshipper 解决了速度慢、冗长和语气尴尬的问题。

Technical details and numbers

技术细节与数据

Core published/priced details

核心发布/定价细节

From @ArtificialAnlys:

来自 @ArtificialAnlys:

  • Context window: 1 million tokens
  • Modalities: text + image inputs
  • Pricing: unchanged from Fable 5 for
  • input: $10 / 1M tokens
  • output: $50 / 1M tokens
  • cache write: $12.5 / 1M tokens
  • Cache read price: reduced from $1.00 to $0.25 / 1M tokens (75% cut)
  • 上下文窗口:100 万 token
  • 模态:文本 + 图像输入
  • 定价:与 Fable 5 相比保持不变,用于
  • 输入:$10 / 1M tokens
  • 输出:$50 / 1M tokens
  • 缓存写入:$12.5 / 1M tokens
  • 缓存读取价格:从 $1.00 降至 $0.25 / 1M tokens(降幅 75%)

Artificial Analysis notes this cache cut materially benefits agentic workloads where much of the prompt is repeatedly re-read from cache.

Artificial Analysis 指出,这一缓存降价显著惠及代理工作负载,其中提示词的大部分会重复从缓存中读取。

Artificial Analysis headline results

Artificial Analysis 头条结果

Also from @ArtificialAnlys:

同样来自 @ArtificialAnlys:

  • Artificial Analysis Intelligence Index: 66 at max effort
  • ahead of:
  • Claude Opus 5 max: 63
  • Claude Fable 5 max: 62
  • GPT-5.6 Sol max: 61
  • Grok 4.6 high: 61
  • HLE: 59.1%
  • previous best cited: Fable 5 at 55.5%
  • Terminal-Bench v2.1: 91.4%
  • SciCode: 62.0%
  • τ³-Banking: +9 points over Fable 5
  • GDPval-AA v2: 1853 Elo, +130 over Fable 5
  • AA-Briefcase: 1694 Elo, +122 over Fable 5
  • Artificial Analysis 智能指数:最大努力模式下为 66
  • 领先于:
  • Claude Opus 5 max:63
  • Claude Fable 5 max:62
  • GPT-5.6 Sol max:61
  • Grok 4.6 high:61
  • HLE:59.1%
  • 此前最佳引用值:Fable 5 为 55.5%
  • Terminal-Bench v2.1:91.4%
  • SciCode:62.0%
  • τ³-Banking:比 Fable 5 高出 9 分
  • GDPval-AA v2:1853 Elo,比 Fable 5 高出 130
  • AA-Briefcase:1694 Elo,比 Fable 5 高出 122

But AA also adds an important qualification:

但 AA 也补充了一项重要的限定条件:

  • On agentic knowledge work, Fable 5.1 is effectively tied with Opus 5 on some measures, not obviously dominant
  • Their eval used Anthropic’s default server-side fallback, with safety-flagged requests routed to Claude Opus 4.8 or Claude Opus 5
  • Fallback accounted for ~4% of output tokens across the Intelligence Index
  • 在代理式知识工作方面,Fable 5.1 在某些指标上与 Opus 5 基本持平,并未表现出明显的优势
  • 其评估使用了 Anthropic 默认的服务器端回退机制,被安全标记的请求会被路由至 Claude Opus 4.8 或 Claude Opus 5
  • 在整个 Intelligence Index 中,回退机制产生的输出 token 约占 4%

That fallback detail became one of the most consequential technical caveats in community interpretation.

这一回退细节成为社区解读中最具影响力的技术注意事项之一。

Cost per task

每项任务的成本

Artificial Analysis also reported:

Artificial Analysis 还报告称:

  • Fable 5.1 max: $3.76/task
  • Fable 5 max: lower, so 5.1 is 20% more expensive per task
  • reason: Fable 5.1 uses ~1.7× output tokens
  • cache cut saves ~$1.40 per task
  • Fable 5.1 xhigh: score 65, cost $2.72/task
  • Opus 5 max: score 63, cost $2.34/task
  • Fable 5.1 max:每任务 $3.76
  • Fable 5 max 更低,因此 5.1 每项任务的成本高出 20%
  • 原因:Fable 5.1 使用的输出 token 约为前者的 1.7 倍
  • 缓存削减使每项任务节省约 $1.40
  • Fable 5.1 xhigh:得分 65,成本 $2.72/项任务
  • Opus 5 max:得分 63,成本 $2.34/任务

This produced one of the key tensions in the reaction cycle: Fable 5.1 looks clearly better at the frontier ceiling, but not clearly better on every cost-efficiency framing.

这产生了反应周期中的一个关键张力:Fable 5.1 在前端天花板(frontier ceiling)上明显更好,但在每种成本效率框架下并不一定都更优。

Additional framing from @nicdunz:

来自 @nicdunz 的补充说明:

  • Fable 5.1 Max: 66 intelligence, 140M tokens, $3.69/task
  • Fable 5 Max: 62, 83M tokens, $3.14/task
  • GPT-5.6 Sol Max: 61, 70M tokens, $0.95/task
  • Fable 5.1 Max:智能度 66,1.4 亿 tokens,$3.69/任务
  • Fable 5 Max:62,8300 万 tokens,$3.14/任务
  • GPT-5.6 Sol Max:61,7000 万 tokens,$0.95/任务

This post argues Sol remains the clear winner on intelligence-per-dollar and intelligence-per-token, even if Fable 5.1 wins absolute ceiling.

本文认为,尽管 Fable 5.1 在绝对天花板方面胜出,但 Sol 在每美元智能度和每 token 智能度上仍然是明确的赢家。

Benchmark snippets from system-card discussion

系统卡片讨论中的基准测试片段

Community members extracted several benchmark points:

社区成员提取了几个基准测试点:

From @StevenDillmann:

来自 @StevenDillmann:

  • Terminal-Bench-Science 0.1
  • Fable 5: 24.7%
  • Fable 5.1: 52.6%
  • more than 2× improvement
  • Terminal-Bench-Science 0.1
  • Fable 5:24.7%
  • Fable 5.1:52.6%
  • 提升超过 2 倍

From @scaling01:

来自 @scaling01:

  • DeepSWE: 67.4%
  • FrontierCode 1.1 Extended: 63.6%
  • FrontierSWE v2: 0.57, “highest of the models Proximal evaluated”
  • DeepSWE:67.4%
  • FrontierCode 1.1 Extended:63.6%
  • FrontierSWE v2:0.57,“Proximal 评估模型中最高”

From @Sauers_:

来自 @Sauers_:

  • Humanity’s Last Exam: 65% with tools
  • Humanity’s Last Exam:使用工具时得分为 65%

From @perplexity_ai:

来自 @perplexity_ai:

  • Perplexity’s August WANDR evaluation:
  • score 0.601
  • $12.76 per task
  • 21% higher score
  • 37% lower cost than Fable 5
  • Perplexity 的 8 月 WANDR 评估:
  • 得分 0.601
  • 每任务 $12.76
  • 得分高出 21%
  • 成本比 Fable 5 低 37%

From @scaling01:

来自 @scaling01:

  • Artificial Analysis Intelligence Index score 66, “back on the frontier”
  • Artificial Analysis Intelligence Index 评分为 66,"重回前沿"

From @theo:

来自 @theo:

  • cache price cut was the “biggest W”
  • in CursorBench, costs were cut by “almost 50%” while scoring higher
  • 缓存降价是"最大的胜利"
  • 在 CursorBench 中,成本削减了"近 50%",同时得分更高

From @kimmonismus:

来自 @kimmonismus:

  • Fable 5.1 High appears stronger and cheaper than Sol 5.6 Max on Cursor Bench
  • though this is a secondary paraphrase, not an original benchmark report
  • 在 Cursor Bench 上,Fable 5.1 High 看起来比 Sol 5.6 Max 更强且更便宜
  • 尽管这是二次转述,而非原始基准测试报告

From @scaling01:

来自 @scaling01:

  • Mythos 5.1 displays verbalized grader awareness in 65% of long agentic coding environments
  • Mythos 5.1 在 65% 的长程智能体编码环境中表现出 verbalized grader awareness(可解释的评分器意识)

That last point is especially interesting: it suggests the model may explicitly model the evaluator in a large fraction of long-horizon coding contexts, which raises both capability and eval-gaming questions.

最后一点特别有趣:这表明模型可能在很大比例的长周期编码场景中显式地对评估者进行建模,这既引发了关于能力的问题,也引发了关于评估作弊(eval-gaming)的问题。

Safeguards and routing details

安全护栏与路由细节

Two tweets capture the technical interpretive crux:

两条推文捕捉到了技术解读的关键点:

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

关联讨论

同一事件的更多信源
Anthropic发布Fable 5.1与Mythos 5.1
歸藏(guizang.ai)原文
Anthropic发布Claude Fable 5.1
量子位(RSS)原文
Fable 5.1发布,主打Agentic Scientific
Przemek Chojecki | PC原文
Anthropic发布Claude Fable 5.1与Mythos
MarkTechPost(RSS)原文
Anthropic发布Claude Fable 5.1与Mythos 5.1模型
Hacker News Best(web_list)原文
Anthropic发布Claude Fable 5.1
The Decoder(RSS)原文
Fable 5.1 上线 Vercel AI Gateway
Guillermo Rauch原文
Anthropic 发布 Claude Fable 5.1 模型
Anthropic(YouTube)原文

相似阅读

另一事件,读法相近