Anthropic 发布 Claude Opus 5,半价性能超旗舰
Anthropic launched Claude Opus 5 at half the price of its most powerful model Fa…
Opus 5 以半价实现超越旗舰的性能,且自主修复能力突出,做 Agent 的同学务必关注定价和实际任务完成率。
Anthropic launched Claude Opus 5 at half the price of its most powerful model Fable 5, while at the same time beating Feble 5 in almost all benchmarks.
Costs the same as Opus 4.8, $5/$25 per mn input/ output tokens, but exactly half as much as Fable 5.
Frankly I am trying to figure out at this price of Opus 5 and with this great benchmark result why would I ever need to use Fable 5.
Opus 5 also has less restrictive cyber classifiers than Fable, automatic fallback options, mid-conversation tool changes and no special data-retention requirement for general access.
The cheaper model is winning specifically on economically valuable activities—coding, operating computers, research and completing business processes—not merely trivia or academic question answering.
Anthropic repeatedly describes Opus 5 as more willing to verify, recover and continue until a task is actually complete. Examples include:
- Creating its own computer-vision pipeline to reconstruct a 3D machine component from raw pixels - Finding the root cause of an open-source bug and fixing an edge case missed by the existing community patch - Building a market-data integration and then creating its own test harness when no live validation feed was available - Completing Zapier’s full churn-prevention workflow when previous models failed
These are signs of adaptive problem-solving: the model notices that a required capability or validation mechanism is missing and constructs one instead of stopping.
Anthropic also describes Opus 5 as more thorough about validating its work and recovering when necessary. The launch gives examples in which it constructed its own vision pipeline, identified the underlying cause of a software bug and built a test harness when no live data feed existed.
That is potentially more commercially meaningful than a modest benchmark gain. An agent that finishes 70% of jobs without intervention can be dramatically more valuable than one that produces slightly better individual responses but frequently stalls.
GPT-5.6 Sol leads DeepSWE: 72.7% versus Opus 5’s 68.8%
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力