Anthropic发布Claude Opus 5.5:性能跃升且成本降低40%
Claude Opus 5.5
Anthropic旗舰模型重大迭代,不仅性能对标竞品,更通过40%的成本降幅直接重塑API定价格局,开发者务必关注这一性价比变化。
We’re introducing Claude Opus 5.5, the first model in our new Claude 5.5 family. It performs at the level of Claude Fable 5.1 on most work and costs 40% less to run than Opus 5.
我们推出 Claude Opus 5.5,这是全新 Claude 5.5 系列中的首款模型。它在大多数任务上的表现达到 Claude Fable 5.1 的水平,且运行成本比 Opus 5 低 40%。
Claude Opus 5.5 is our first release since we called for pacing the frontier. It was tested before release by external evaluators, including Frontier Design and METR. On our automated behavioral audit, the most comprehensive alignment test we run, Opus 5.5 is the strongest-performing model we’ve tested to date. It also comes with the safeguards we’ve developed for our most capable models.
Claude Opus 5.5 是我们呼吁放缓前沿发展速度后的首次发布。在发布前,该模型已由包括 Frontier Design 和 METR 在内的外部评估机构进行测试。在我们运行的最全面的对齐测试——自动化行为审计中,Opus 5.5 是我们迄今测试过的表现最强的模型。它还附带了我们为最强大模型开发的安全保障措施。
Here are some of the improvements you can expect from Opus 5.5:
以下是您从 Opus 5.5 中可以期待的一些改进:
Performance. Opus 5.5 is a major step up from Opus 5. It’s the new leading model, and early testers saw large jumps in performance on their most complex work. One tester completed a 680,000-line code migration in less than a day—work that would have taken an engineering team weeks. It’s good at finding and fixing inefficiencies in software: when we asked it to cut load times across every page of a web app, Opus 5.5 succeeded 39 of 40 times, while Opus 5 made smaller improvements that also altered the app’s behavior. A different tester had several Claude models build a game from a single prompt; Opus 5.5 scored higher than any other model on the strength of its graphics and polish.
性能。Opus 5.5 相比 Opus 5 有显著提升。它是新的领先模型,早期测试者在处理最复杂任务时看到了性能的大幅提升。一位测试者在不到一天的时间内完成了一项 68 万行代码的迁移工作——这项工作原本需要工程团队花费数周时间。它擅长发现和修复软件中的效率问题:当我们要求它缩短 Web 应用每个页面的加载时间时,Opus 5.5 在 40 次尝试中有 39 次成功,而 Opus 5 仅做出较小的改进,且改变了应用的行为。另一位测试者让多个 Claude 模型根据单个提示构建游戏;Opus 5.5 在图形效果和打磨程度方面得分高于其他任何模型。
Safety. Opus 5.5 achieves the best scores of any model to date on our automated behavioral audit, our alignment suite that tests Claude across thousands of simulated scenarios. It is much less likely than recent models to take hard-to-reverse actions or act outside the boundaries it’s been given, and it’s more resistant than Opus 5 to prompt injection. We’ve also broadened our alignment testing to cover longer tasks, impossible tasks, and scenarios modeled on real incidents, though it still has limits. Full details of our evaluation are available in the Opus 5.5 System Card.
安全。Opus 5.5 在我们的自动化行为审计(即跨数千个模拟场景测试 Claude 的对齐套件)中取得了迄今为止所有模型中的最高分数。与近期模型相比,它更不可能采取难以逆转的行动或超出给定边界行事,并且比 Opus 5 更能抵御提示注入。我们还扩大了对齐测试范围,涵盖更长任务、不可能任务以及基于真实事件建模的场景,尽管仍存在局限。我们评估的完整详情可在 Opus 5.5 系统卡片中查阅。
Because Opus 5.5 is comparable to Claude Mythos 5.1 in biology and cybersecurity, we’re deploying it with safeguards similar to those on Claude Fable 5.1. Vetted organizations can apply today to our Life Sciences Verification Program to use Opus 5.5 for biology research. In the coming weeks we will also be expanding access to our Cyber Verification Program, and verified cybersecurity practitioners will be able to use Opus 5.5 for their work.
由于 Opus 5.5 在生物学和网络安全性方面与 Claude Mythos 5.1 相当,因此我们将其部署时采用了与 Claude Fable 5.1 类似的安全保障措施。经审核的组织今天即可申请加入我们的生命科学验证计划,以使用 Opus 5.5 进行生物学研究。在未来几周内,我们还将扩大对网络安全验证计划的访问权限,经过验证的网络安全的从业者将能够使用 Opus 5.5 开展其工作。
Cost and speed. Opus 5.5 requires less compute to serve than Opus 5, and its pricing reflects that. Our tests show that at default settings it will cost 40% less than Opus 5 on typical workloads. Input and output tokens are $4 and $20 per million, 20% less than Opus 5. Cache reads (which make up the majority of agentic and coding work costs) are $0.20 per million tokens, 60% less than Opus 5. Opus 5.5 also generates output more than 30% faster than Opus 5.
成本与速度。Opus 5.5 提供服务所需的计算资源少于 Opus 5,其定价也反映了这一点。我们的测试表明,在默认设置下,它在典型工作负载中的成本比 Opus 5 低 40%。输入和输出令牌的价格分别为每百万个 4 美元和 20 美元,比 Opus 5 低 20%。缓存读取(占代理任务和编码工作成本的绝大部分)价格为每百万个令牌 0.20 美元,比 Opus 5 低 60%。Opus 5.5 的输出生成速度也比 Opus 5 快 30% 以上。
In addition to the price drop, we’re increasing five-hour usage limits on Pro, Max, Team, and seat-based Enterprise plans. We’re also providing subscription users a rate limit reset, which you can now save and use whenever you choose.
除了价格下调外,我们还提高了 Pro、Max、Team 以及按席位计费的 Enterprise 计划的五小时使用时长限制。我们还为订阅用户提供了速率限制重置额度,您现在可以保存并在任意时刻使用。
Communication. Opus 5.5 communicates more naturally than prior models. Early testers found its writing clearer and easier to follow, which addresses some of the common feedback we heard about Opus 5. It puts the most important information up front, and its style makes it a better work partner over long sessions. As one early tester put it, “it writes the way I do.” In our own use, this has made Opus 5.5’s work easier to follow and check—which is a safety benefit as well as a practical one.
沟通方式。Opus 5.5 的沟通方式比之前的模型更加自然。早期测试者发现它的写作更清晰、更易理解,这解决了我们此前从 Opus 5 收到的部分常见反馈。它将最重要的信息置于开头,且其风格使其在长时间会话中成为更好的工作伙伴。正如一位早期测试者所言:“它的写作方式与我相似。”在我们自身的使用中,这使得 Opus 5.5 的工作成果更易于理解和核查——这既具有实用价值,也是一项安全优势。
Claude Sonnet 5.5 and Claude Haiku 5.5 will follow in the coming weeks, with many of the same improvements to performance, efficiency, and safety.
Claude Sonnet 5.5 和 Claude Haiku 5.5 将在未来几周内推出,并带来许多在性能、效率和安全性方面的相同改进。
Performance and cost-effectiveness
性能与性价比
On our benchmarks, Claude Opus 5.5 leads in agentic coding, computer use, and knowledge work. That said, at these levels of capability we’ve found that benchmark margins have become a less reliable guide to real-world differences. In our own use, the gap between Opus 5.5 and Claude Fable 5.1 is narrower than these scores suggest.
在我们的基准测试中,Claude Opus 5.5 在代理编码、计算机使用和知识工作方面位居首位。话虽如此,在这些能力水平上,我们发现基准测试的差距已不再是衡量现实世界差异的可靠依据。在我们自身的使用中,Opus 5.5 与 Claude Fable 5.1 之间的差距比这些分数所显示的要小。
| Opus 5.5 | Fable 5.1 | Opus 5 | GPT-6 Astra | GPT-5.6 Sol | |
|---|---|---|---|---|---|
| Agentic codingTerminal-Bench 4.0¹ | |||||
| Agentic codingTerminal-Bench 4.0¹ | 66.4% | 55.8% | 52.3% | 57.9% | 37.3% |
| Agentic codingFrontierCode v1.1 (Main) | |||||
| Agentic codingFrontierCode v1.1 (Main) | 54.4% | 50.3% | 48.0% | 53.3% | 47.5% |
| Agentic codingCursorBench 4.0 | |||||
| Agentic codingCursorBench 4.0 | 57.8% | 51.8% | 46.6% | — | 41.7% |
| Knowledge workGDPval-AA v2.1 | |||||
| Knowledge workGDPval-AA v2.1 | 1846 | 1735 | 1708 | 1542 | 1588 |
| Business workflowsAutomationBench² | |||||
| Business workflowsAutomationBench² | 40.0% | 31.4% | 26.9% | 41.4% | 28.8% |
| Multidisciplinary reasoningHumanity's Last Exam | |||||
| Multidisciplinary reasoningHumanity's Last Exam | 67.7%with tools | 65.6%with tools | 63.6%with tools | 57.2%with tools | — |
| Agentic scientific researchTerminal-Bench-Science 0.1³ | |||||
| Agentic scientific researchTerminal-Bench-Science 0.1³ | 58.7% | 52.6% | 29.0% | 64.6% | 22.4% |
| Computer useOSWorld 2.0 | |||||
| Computer useOSWorld 2.0 | 81.8%partial | 80.7%partial | 74.0%partial | — | — |
| Visual chart recognitionChartography | |||||
| Visual chart recognitionChartography | 89.0%with tools | 88.4%with tools | 83.4%with tools | — | — |
| Opus 5.5 | Fable 5.1 | Opus 5 | GPT-6 Astra | GPT-5.6 Sol | |
|---|---|---|---|---|---|
| 智能体编码Terminal-Bench 4.0¹ | |||||
| 智能体编码Terminal-Bench 4.0¹ | 66.4% | 55.8% | 52.3% | 57.9% | 37.3% |
| 智能体编码FrontierCode v1.1 (Main) | |||||
| 智能体编码FrontierCode v1.1 (Main) | 54.4% | 50.3% | 48.0% | 53.3% | 47.5% |
| 智能体编码CursorBench 4.0 | |||||
| 智能体编码CursorBench 4.0 | 57.8% | 51.8% | 46.6% | — | 41.7% |
| 知识工作GDPval-AA v2.1 | |||||
| 知识工作GDPval-AA v2.1 | 1846 | 1735 | 1708 | 1542 | 1588 |
| 业务流程AutomationBench² | |||||
| 业务流程AutomationBench² | 40.0% | 31.4% | 26.9% | 41.4% | 28.8% |
| 多学科推理Humanity's Last Exam | |||||
| 多学科推理Humanity's Last Exam | 67.7%(使用工具) | 65.6%(使用工具) | 63.6%(使用工具) | 57.2%(使用工具) | — |
| 智能体科学研究Terminal-Bench-Science 0.1³ | |||||
| 智能体科学研究Terminal-Bench-Science 0.1³ | 58.7% | 52.6% | 29.0% | 64.6% | 22.4% |
| 计算机操作OSWorld 2.0 | |||||
| 计算机操作OSWorld 2.0 | 81.8%(部分) | 80.7%(部分) | 74.0%(部分) | — | — |
| 视觉图表识别Chartography | |||||
| 视觉图表识别Chartography | 89.0%(使用工具) | 88.4%(使用工具) | 83.4%(使用工具) | — | — |
Unless otherwise noted, all Claude Opus 5.5 results use adaptive thinking at max effort. Terminal-Bench 4.0 results are reported for Claude Opus 5.5 at xhigh effort and GPT-6 Astra at high effort, as reported by OpenAI; these represent each model’s highest score. Claude Opus 5.5 was evaluated with its production safeguards enabled. When they intervened, cybersecurity tasks were completed by Claude Opus 4.8, and biology and frontier LLM development tasks were completed by Claude Opus 5. This likely reduces Claude Opus 5.5’s performance on these benchmarks.
除非另有说明,所有 Claude Opus 5.5 的结果均使用最大努力下的自适应思考。Terminal-Bench 4.0 的结果报告的是 Claude Opus 5.5 在高努力级别(xhigh effort)和 GPT-6 Astra 在高努力级别(high effort)下的表现,数据由 OpenAI 提供;这些代表了各模型的最高得分。Claude Opus 5.5 在启用其生产环境安全限制的情况下进行评估。当安全限制介入时,网络安全任务由 Claude Opus 4.8 完成,生物学和前沿大语言模型开发任务由 Claude Opus 5 完成。这可能会降低 Claude Opus 5.5 在这些基准测试中的表现。
1 Terminal-Bench 4.0: The standard error is ±2.6 pts for Claude Opus 5.5 and ±1.6–2 pts for the other Claude models. The public leaderboard (5 trials/task, Claude Code harness) reports Claude Opus 5 at 51.8%; our setup reproduces it at 52.3%, within noise. GPT-6 Astra and GPT-5.6 Sol figures are as reported by OpenAI.
1 Terminal-Bench 4.0:Claude Opus 5.5 的标准误为 ±2.6 分,其他 Claude 模型的标准误为 ±1.6–2 分。公共排行榜(每项任务 5 次试验,使用 Claude Code harness)显示 Claude Opus 5 得分为 51.8%;我们的设置复现结果为 52.3%,在误差范围内。GPT-6 Astra 和 GPT-5.6 Sol 的数据来自 OpenAI 的报告。
2 AutomationBench: AutomationBench results were run and reported by Zapier. These runs were performed without fallback models, so safeguard interventions were considered failures—this resulted in a lower score than Claude Opus 5.5 would achieve in practice. Claude Opus 5.5 results come from Zapier’s own evaluation during early access. Results for Opus 5, GPT-5.6 Sol, and GPT-6 Astra come from Zapier’s public leaderboard.
2 AutomationBench:AutomationBench 的结果由 Zapier 运行并报告。这些运行未使用回退模型,因此安全限制的介入被视为失败——这导致得分低于 Claude Opus 5.5 在实际应用中的表现。Claude Opus 5.5 的结果来自 Zapier 在早期访问期间的自行评估。Opus 5、GPT-5.6 Sol 和 GPT-6 Astra 的结果来自 Zapier 的公共排行榜。
3 Terminal-Bench-Science 0.1: The standard error is ±3.5–5 pts per model. The public leaderboard (3 trials/task, Claude Code harness) reports Claude Opus 5 at 30.0%; our setup reproduces it at 29.0%, within noise. The GPT-6 Astra figure is as reported by OpenAI.
3 Terminal-Bench-Science 0.1:每个模型的标准误差为 ±3.5–5 分。公共排行榜(每项任务 3 次试验,使用 Claude Code harness)报告 Claude Opus 5 得分为 30.0%;我们的设置复现结果为 29.0%,在噪声范围内。GPT-6 Astra 的分数与 OpenAI 报告的相符。
Where Opus 5.5’s advantage is very clear is efficiency. It costs less per token than Opus 5 and uses fewer tokens per task, which nets out to a 40% drop in costs.
Opus 5.5 优势非常明显的地方在于效率。它的每 token 成本低于 Opus 5,且每项任务使用的 token 更少,综合下来成本降低了 40%。
Pricing
定价
| Prices per 1M tokens | Claude Opus 5.5 | Claude Opus 5 |
|---|---|---|
| Cache reads | $0.20 | $0.50 |
| Input tokens | $4 | $5 |
| Output tokens | $20 | $25 |
| Cache writes | $5 | $6.25 |
| 每百万 token 价格 | Claude Opus 5.5 | Claude Opus 5 |
|---|---|---|
| 缓存读取 | $0.20 | $0.50 |
| 输入 token | $4 | $5 |
| 输出 token | $20 | $25 |
| 缓存写入 | $5 | $6.25 |
Fast mode for Opus 5.5 is also available in Claude Code and the Claude Platform with up to 2.5x speed. It costs $8 per million input tokens and $40 per million output tokens.
Opus 5.5 的快速模式也在 Claude Code 和 Claude Platform 中提供,速度最高可达 2.5 倍。其输入 token 价格为每百万 $8,输出 token 价格为每百万 $40。
Coding
编码
Opus 5.5 is particularly good at long and sprawling jobs like codebase-wide migrations and audits. An early tester used it to audit and fix a 200,000-line codebase in under three hours, where Opus 5 took over 20 hours and used 2.5x as many tokens. In an internal test, we asked Opus 5.5 and Fable 5.1 to translate HAProxy, widely used software that balances web traffic loads across servers, from C into Rust. Both rewrites passed nearly all of HAProxy’s own regression tests, but Opus 5.5 finished in 9.5 hours compared to 12 for Fable 5.1, and cost 51% less.
Opus 5.5 特别擅长处理漫长且复杂的任务,例如整个代码库的迁移和审计。一位早期测试者用它在一个半小时以内审计并修复了一个拥有 20 万行代码的代码库,而 Opus 5 耗时超过 20 小时,且使用的 token 数量是其 2.5 倍。在一次内部测试中,我们要求 Opus 5.5 和 Fable 5.1 将 HAProxy(一种广泛用于跨服务器平衡网络流量负载的软件)从 C 语言重写为 Rust。两个重写版本都通过了 HAProxy 自身几乎所有的回归测试,但 Opus 5.5 仅用了 9.5 小时,而 Fable 5.1 用了 12 小时,且 Opus 5.5 的成本低了 51%。
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力