跳到主内容
精选85Rohan Paul模型发布/更新多源精选 ×12

Anthropic 发布 Claude Opus 5,半价性能超旗舰

Anthropic launched Claude Opus 5 at half the price of its most powerful model Fa…

原文
推荐理由

Opus 5 以半价实现超越旗舰的性能,且自主修复能力突出,做 Agent 的同学务必关注定价和实际任务完成率。

Anthropic launched Claude Opus 5 at half the price of its most powerful model Fable 5, while at the same time beating Feble 5 in almost all benchmarks.

Costs the same as Opus 4.8, $5/$25 per mn input/ output tokens, but exactly half as much as Fable 5.

Frankly I am trying to figure out at this price of Opus 5 and with this great benchmark result why would I ever need to use Fable 5.

Opus 5 also has less restrictive cyber classifiers than Fable, automatic fallback options, mid-conversation tool changes and no special data-retention requirement for general access.

The cheaper model is winning specifically on economically valuable activities—coding, operating computers, research and completing business processes—not merely trivia or academic question answering.

Anthropic repeatedly describes Opus 5 as more willing to verify, recover and continue until a task is actually complete. Examples include:

  • Creating its own computer-vision pipeline to reconstruct a 3D machine component from raw pixels - Finding the root cause of an open-source bug and fixing an edge case missed by the existing community patch - Building a market-data integration and then creating its own test harness when no live validation feed was available - Completing Zapier’s full churn-prevention workflow when previous models failed

These are signs of adaptive problem-solving: the model notices that a required capability or validation mechanism is missing and constructs one instead of stopping.

Anthropic also describes Opus 5 as more thorough about validating its work and recovering when necessary. The launch gives examples in which it constructed its own vision pipeline, identified the underlying cause of a software bug and built a test harness when no live data feed existed.

That is potentially more commercially meaningful than a modest benchmark gain. An agent that finishes 70% of jobs without intervention can be dramatically more valuable than one that produces slightly better individual responses but frequently stalls.

GPT-5.6 Sol leads DeepSWE: 72.7% versus Opus 5’s 68.8%

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

关联讨论

同一事件的更多信源
Opus 5 发布
gabriel原文
Opus 5 在 ARC-AGI-3 得分翻四倍
Teortaxes▶️ (DeepSeek 推特🐋铁粉 2023 – ∞)原文
Anthropic Opus 5 抗提示注入能力显著提升
Simon Willison 博客(RSS)原文
Opus 5自估41%概率为道德主体,较前代提升
AI Notkilleveryoneism Memes ⏸️原文
Anthropic 发布 Claude Opus 5,…
Anthropic Newsroom(web_list)原文
Anthropic发布Opus 5,更便宜且限制更少
TechCrunch AI(RSS)原文

相似阅读

另一事件,读法相近