Claude Fable 5.1发布:缓存降价75%且代理基准性能翻倍
Claude Fable 5.1 is cheaper, substantially stronger on several agentic benchmark…
Fable 5.1在代理场景的成本优化和性能跃迁显著,缓存降价对长链路Agent开发者极具吸引力,建议评估迁移。
Claude Fable 5.1 is cheaper, substantially stronger on several agentic benchmarks, more concise, and apparently far less trigger-happy (says Anthropic).
Claude Fable 5.1 更便宜,在多项代理基准测试中实力显著更强,输出更简洁,而且显然没那么容易触发限制(Anthropic 如是说)。
the tl;dr
精简摘要
How much cheaper? -Input/output pricing remains $10/$50 per million -tokens. -Cache reads fall 75% to $0.25. -Anthropic estimates ~25% lower costs for typical workloads and up to ~45% for highly agentic, context-heavy work.
便宜了多少? -输入/输出定价仍为每百万 token $10/$50。 -缓存读取价格下降 75% 至 $0.25。 -Anthropic 估计典型工作负载成本降低约 25%,而高度依赖代理、上下文繁重的工作负载可降低高达约 45%。
BUT: not 45% cheaper for Claude subscriptions. ("...wherever usage is billed by token")
但是:Claude 订阅费用并非便宜 45%。(“……仅在按 token 计费的使用场景中”)
How much better than Fable 5? -Scientific agent benchmark: 52.6% vs. 24.7% - more than 2× -AutomationBench: 31.4% vs. 17.1% — an 84% relative gain -GDPval-AA: 1,853 vs. 1,723 -CursorBench: 73.4% vs. 70.5%
比 Fable 5 强多少? -科学代理基准测试:52.6% vs. 24.7% —— 提升超过 2 倍 -AutomationBench:31.4% vs. 17.1% —— 相对提升 84% -GDPval-AA:1,853 vs. 1,723 -CursorBench:73.4% vs. 70.5%
So: dramatic gains on some long-running tasks, modest improvements elsewhere, not a uniform intelligence jump.
因此:在某些长期任务上取得显著进步,在其他方面则有小幅提升,并非整体智能水平的均匀跃升。
Verbosity also seems improved, although there is no standardized score. Rogo reports equal accuracy with 20% fewer tokens. Red Hat found its updates more concise and easier to follow. Every says it used half as many tokens as Opus 5 while running about twice as fast.
冗长程度似乎也有所改善,尽管没有标准化评分。Rogo 报告称准确率相当但 token 使用量减少了 20%。Red Hat 发现其更新内容更简洁且更易理解。Every 表示其在运行速度约为 Opus 5 两倍的情况下,使用的 token 数量仅为后者的一半。
And fewer unnecessary red flags: -~60% fewer cyber-safeguard interventions per Claude Code session -Biology safeguards reportedly trigger 85% less often on benign elementary biology and medical questions -Vulnerability discovery is now allowed, while exploit generation, penetration testing and binary scanning remain restricted or redirected
不必要的警告标志也减少了: -每次 Claude Code 会话中的网络安全干预措施减少约 60% -据报道,在 benign(无害)的基础生物学和医学问题上,生物学安全限制触发的频率降低了 85% -漏洞发现现在被允许,而利用代码生成、渗透测试和二进制扫描仍受到限制或引导至其他用途
So far, sounds like a promising release. Although it clearly shows they care much more about business and enterprise users than us subscription pesants. Anyway: Testing time!
到目前为止,这看起来是一个充满希望的版本。尽管它清楚地表明,他们更关心商业和企业用户,而不是我们这些订阅制的平民百姓。不管怎样:开始测试吧!
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力