跳到主内容
@wquguru
精选88Chubby♨️模型发布/更新多源精选 ×13

Claude Fable 5.1发布:缓存降价75%且代理基准性能翻倍

Claude Fable 5.1 is cheaper, substantially stronger on several agentic benchmark…

原文
发到 X
推荐理由

Fable 5.1在代理场景的成本优化和性能跃迁显著,缓存降价对长链路Agent开发者极具吸引力,建议评估迁移。

Claude Fable 5.1 is cheaper, substantially stronger on several agentic benchmarks, more concise, and apparently far less trigger-happy (says Anthropic).

Claude Fable 5.1 更便宜,在多项代理基准测试中实力显著更强,输出更简洁,而且显然没那么容易触发限制(Anthropic 如是说)。

the tl;dr

精简摘要

How much cheaper? -Input/output pricing remains $10/$50 per million -tokens. -Cache reads fall 75% to $0.25. -Anthropic estimates ~25% lower costs for typical workloads and up to ~45% for highly agentic, context-heavy work.

便宜了多少? -输入/输出定价仍为每百万 token $10/$50。 -缓存读取价格下降 75% 至 $0.25。 -Anthropic 估计典型工作负载成本降低约 25%,而高度依赖代理、上下文繁重的工作负载可降低高达约 45%。

BUT: not 45% cheaper for Claude subscriptions. ("...wherever usage is billed by token")

但是:Claude 订阅费用并非便宜 45%。(“……仅在按 token 计费的使用场景中”)

How much better than Fable 5? -Scientific agent benchmark: 52.6% vs. 24.7% - more than 2× -AutomationBench: 31.4% vs. 17.1% — an 84% relative gain -GDPval-AA: 1,853 vs. 1,723 -CursorBench: 73.4% vs. 70.5%

比 Fable 5 强多少? -科学代理基准测试:52.6% vs. 24.7% —— 提升超过 2 倍 -AutomationBench:31.4% vs. 17.1% —— 相对提升 84% -GDPval-AA:1,853 vs. 1,723 -CursorBench:73.4% vs. 70.5%

So: dramatic gains on some long-running tasks, modest improvements elsewhere, not a uniform intelligence jump.

因此:在某些长期任务上取得显著进步,在其他方面则有小幅提升,并非整体智能水平的均匀跃升。

Verbosity also seems improved, although there is no standardized score. Rogo reports equal accuracy with 20% fewer tokens. Red Hat found its updates more concise and easier to follow. Every says it used half as many tokens as Opus 5 while running about twice as fast.

冗长程度似乎也有所改善,尽管没有标准化评分。Rogo 报告称准确率相当但 token 使用量减少了 20%。Red Hat 发现其更新内容更简洁且更易理解。Every 表示其在运行速度约为 Opus 5 两倍的情况下,使用的 token 数量仅为后者的一半。

And fewer unnecessary red flags: -~60% fewer cyber-safeguard interventions per Claude Code session -Biology safeguards reportedly trigger 85% less often on benign elementary biology and medical questions -Vulnerability discovery is now allowed, while exploit generation, penetration testing and binary scanning remain restricted or redirected

不必要的警告标志也减少了: -每次 Claude Code 会话中的网络安全干预措施减少约 60% -据报道,在 benign(无害)的基础生物学和医学问题上,生物学安全限制触发的频率降低了 85% -漏洞发现现在被允许,而利用代码生成、渗透测试和二进制扫描仍受到限制或引导至其他用途

So far, sounds like a promising release. Although it clearly shows they care much more about business and enterprise users than us subscription pesants. Anyway: Testing time!

到目前为止,这看起来是一个充满希望的版本。尽管它清楚地表明,他们更关心商业和企业用户,而不是我们这些订阅制的平民百姓。不管怎样:开始测试吧!

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

关联讨论

同一事件的更多信源

相似阅读

关联信息,但可能不是同一事件