Anthropic发布Claude Sonnet 5.5:速度提升30%以上
Sonnet 5.5
Introducing Claude Sonnet 5.5, the second model in the Claude 5.5 family. It’s a clear upgrade over Claude Sonnet 5, runs 30%+ faster, and costs up to 30% less for most work.
介绍 Claude Sonnet 5.5,这是 Claude 5.5 系列的第二款模型。它相比 Claude Sonnet 5 有了显著提升,运行速度快了 30% 以上,且大多数工作任务的成本降低了多达 30%。
Sonnet 5.5 is a faster, lower-cost complement to Claude Opus 5.5. Where Opus 5.5 is built for complex work requiring careful judgment, Sonnet 5.5 is strongest at well-scoped everyday tasks, fixing bugs, and creating polished documents, slides, and spreadsheets. It’s also got a sharp eye for design. Claude Haiku 5.5, built for high-volume and cost-sensitive applications, will join the Claude 5.5 family in the coming weeks.
Sonnet 5.5 是 Claude Opus 5.5 的更快、更低成本的补充。Opus 5.5 专为需要谨慎判断的复杂任务而设计,而 Sonnet 5.5 在范围明确的日常任务、修复错误以及制作精美的文档、幻灯片和电子表格方面表现最强。它还具备敏锐的设计眼光。Claude Haiku 5.5 专为高吞吐量和成本敏感型应用而构建,将在未来几周内加入 Claude 5.5 系列。
Sonnet 5.5 improves over Sonnet 5 on:
Sonnet 5.5 在以下方面优于 Sonnet 5:
Performance. Sonnet 5.5 scores 70.6% on Terminal-Bench 4.0, an agentic coding evaluation, compared to Sonnet 5’s 10.3%. It scores two points below Opus 5.5 on GDPval-AA, a test of real-world work across a variety of occupations. And it’s strong on long-horizon work and image understanding—it’s the first Sonnet model to beat Pokémon Red working only from screenshots.
性能。Sonnet 5.5 在 Terminal-Bench 4.0(一项代理式编码评估)中得分 70.6%,而 Sonnet 5 的得分为 10.3%。在 GDPval-AA(一项涵盖多种职业的真实工作测试)中,其得分比 Opus 5.5 低两个百分点。它在长周期任务和图像理解方面也表现出色——它是首个仅凭截图就能通关《宝可梦 红》的 Sonnet 模型。
Collaboration. Like Opus 5.5, Sonnet 5.5 writes more clearly than our previous generation of models; early testers described it as a better partner for collaboration than Sonnet 5. Its speed also makes it well suited to fast iteration on less complex tasks.
协作能力。与 Opus 5.5 一样,Sonnet 5.5 的写作清晰度优于我们上一代模型;早期测试者表示,它作为协作伙伴的表现优于 Sonnet 5。其速度也使其非常适合对不太复杂的任务进行快速迭代。
Cost. Sonnet 5.5 is priced the same as Sonnet 5 at $2 per million input tokens, $10 per million output tokens, and $0.20 per million tokens for cache reads, but it typically needs far fewer tokens to do the same work. In our testing, it costs up to 30% less per task than its predecessor.
成本。Sonnet 5.5 的定价与 Sonnet 5 相同,即每百万输入 token 2 美元,每百万输出 token 10 美元,缓存读取每百万 token 0.20 美元,但完成相同工作通常所需的 token 数量要少得多。在我们的测试中,其每项任务的成本比前代产品最多降低 30%。
Speed. Sonnet 5.5 generates outputs 30%+ faster than Sonnet 5, making it our fastest Sonnet model to date.
速度。Sonnet 5.5 的输出生成速度比 Sonnet 5 快 30% 以上,成为迄今为止最快的 Sonnet 模型。
Alignment and safety. On our automated behavioral audit, Sonnet 5.5 improves on or matches Sonnet 5 on most measures of alignment. Because its cybersecurity capabilities are comparable to Opus 5’s, it’s the first Sonnet model to launch with cyber safeguards and fallbacks like those we’ve developed for our most capable models. Its biology safeguards are the same as Sonnet 5’s. Both safeguards target a narrow set of high-risk requests; routine software development and most life sciences work are unaffected.
对齐与安全。在我们的自动化行为审计中,Sonnet 5.5 在大多数对齐指标上改善或持平于 Sonnet 5。由于其网络安全能力与 Opus 5 相当,它是首款搭载与我们最强大模型相同的网络防护和回退机制的 Sonnet 模型。其生物安全防护与 Sonnet 5 相同。这两类防护均针对少量高风险请求;常规软件开发和大多数生命科学相关工作不受影响。
Performance
性能
Sonnet 5.5 improves on Sonnet 5 across domains—in some cases dramatically. On several evaluations, Sonnet 5.5 at Max effort even performs comparably to Opus 5.5. However, benchmark scores capture only one facet of a model’s capabilities; in our own testing, and in that of external testers, Opus 5.5 remains clearly stronger at complex, open-ended work requiring sustained judgment.
Sonnet 5.5 在多个领域相比 Sonnet 5 有所改进,在某些情况下提升显著。在多项评估中,Sonnet 5.5 在 Max 模式下甚至能与 Opus 5.5 相媲美。然而,基准分数仅捕捉模型能力的一个方面;在我们自己的测试以及外部测试者的测试中,Opus 5.5 在处理需要持续判断的复杂、开放式任务时仍然明显更强。
| Sonnet 5.5 | Sonnet 5 | Opus 5.5 | GPT-6 Sol | |
|---|---|---|---|---|
| Agentic codingTerminal-Bench 4.0 | ||||
| Agentic codingTerminal-Bench 4.0 | 70.6% | 10.3% | 66.4%¹ | — |
| Agentic codingFrontierCode 1.1 (Main) | ||||
| Agentic codingFrontierCode 1.1 (Main) | 46.2%Max² | 42.4% | 54.4% | 49.3% |
| 52.1%Xhigh | ||||
| Agentic codingCursorBench 4.0 | ||||
| Agentic codingCursorBench 4.0 | 55.5% | 34.1% | 57.8% | — |
| Knowledge workGDPval-AA v2.1³ | ||||
| Knowledge workGDPval-AA v2.1³ | 1844 | 1449 | 1846 | 1487⁴ |
| Knowledge workAA-Briefcase v1.1³ | ||||
| Knowledge workAA-Briefcase v1.1³ | 1811 | 1359 | 1822 | 1483⁴ |
| Multidisciplinary reasoningHumanity’s Last Exam | ||||
| Multidisciplinary reasoningHumanity’s Last Exam | 64.5%with tools | 54.9%with tools | 67.7%with tools | — |
| Computer useOSWorld 2.1 | ||||
| Computer useOSWorld 2.1 | 80.1%partial | 57.0%partial | 81.8%partial | — |
| Visual chart recognitionChartography | ||||
| Visual chart recognitionChartography | 61.6%no tools | 15.6%no tools | 64.4%no tools | 53.6%⁴no tools |
| Sonnet 5.5 | Sonnet 5 | Opus 5.5 | GPT-6 Sol | |
|---|---|---|---|---|
| 代理式编码Terminal-Bench 4.0 | ||||
| 代理式编码Terminal-Bench 4.0 | 70.6% | 10.3% | 66.4%¹ | — |
| 代理式编码FrontierCode 1.1 (Main) | ||||
| 代理式编码FrontierCode 1.1 (Main) | 46.2%Max² | 42.4% | 54.4% | 49.3% |
| 52.1%Xhigh | ||||
| 代理式编码CursorBench 4.0 | ||||
| 代理式编码CursorBench 4.0 | 55.5% | 34.1% | 57.8% | — |
| 知识工作GDPval-AA v2.1³ | ||||
| 知识工作GDPval-AA v2.1³ | 1844 | 1449 | 1846 | 1487⁴ |
| 知识工作AA-Briefcase v1.1³ | ||||
| 知识工作AA-Briefcase v1.1³ | 1811 | 1359 | 1822 | 1483⁴ |
| 多学科推理Humanity’s Last Exam | ||||
| 多学科推理Humanity’s Last Exam | 64.5%with tools | 54.9%with tools | 67.7%with tools | — |
| 计算机使用OSWorld 2.1 | ||||
| 计算机使用OSWorld 2.1 | 80.1%partial | 57.0%partial | 81.8%partial | — |
| 视觉图表识别Chartography | ||||
| 视觉图表识别Chartography | 61.6%no tools | 15.6%no tools | 64.4%no tools | 53.6%⁴no tools |
For details on how we run our evaluations, see the Sonnet 5.5 System Card.
有关我们如何运行评估的详细信息,请参阅 Sonnet 5.5 系统卡片。
The charts below plot each model’s score against its cost per task at every effort level. As effort goes up, models typically work for longer, leading to a higher cost per task but generally also a higher score. The closer a point is to the top left of the chart, the more capability it delivers per dollar.
下图展示了每个模型在不同努力水平下的得分与其每项任务成本的关系。随着努力程度的增加,模型通常需要更长的运行时间,导致每项任务的成本更高,但通常也能获得更高的分数。点越靠近图表的左上角,意味着每美元所能提供的能力越强。
On several benchmarks, Sonnet 5.5 at Low or Medium effort beats Sonnet 5’s best score for about a tenth of the cost per task. It complements Opus 5.5 best when running at lower effort settings, where it costs less per task. At higher settings, it can perform comparably at a similar cost.
在多项基准测试中,Sonnet 5.5 在 Low 或 Medium 模式下的表现优于 Sonnet 5 的最佳成绩,而每项任务的成本仅为后者的约十分之一。它在较低的努力设置下与 Opus 5.5 形成最佳互补,因为此时每项任务的成本更低。在较高的设置下,它可以在相似的成本下实现相当的性能。
Agentic terminal codingAgentic coding: FrontierCodeAgentic coding: CursorBenchKnowledge work: AA-Briefcase
代理式终端编码代理式编码:FrontierCode代理式编码:CursorBench知识工作:AA-Briefcase
Agentic terminal codingAgentic coding: FrontierCodeAgentic coding: CursorBenchKnowledge work: AA-Briefcase
代理式终端编码代理式编码:FrontierCode代理式编码:CursorBench知识工作:AA-Briefcase
Terminal-Bench 4.0Accuracy vs. cost
Terminal-Bench 4.0准确率与成本
- Sonnet 5.5
- Opus 5.5
- Sonnet 5
- GPT-5.6 Sol
010203040506070Score (%)12510Cost per attempt (USD, log scale)LowMedHighXhighMax
010203040506070得分 (%)12510每次尝试成本 (USD, 对数刻度)LowMedHighXhighMax
Terminal-Bench 4.0 measures how well a model can complete complex, multi-step professional tasks within a command-line interface. At Medium effort, the default in the Claude apps, Sonnet 5.5 far exceeds Sonnet 5’s best score for less than a tenth of the cost per task.
Terminal-Bench 4.0 衡量模型在命令行界面中完成复杂的多步专业任务的能力。在中等努力程度(Claude 应用中的默认设置)下,Sonnet 5.5 以不到十分之一的单次任务成本,远超 Sonnet 5 的最佳得分。
Terminal-Bench and OpenAI did not report GPT-6 Sol performance publicly, so we report GPT-5.6 Sol here.
Terminal-Bench 和 OpenAI 未公开报告 GPT-6 Sol 的性能,因此我们在此报告 GPT-5.6 Sol 的数据。
FrontierCode v1.1, main setAccuracy vs. cost
FrontierCode v1.1,主集准确率与成本对比
- Sonnet 5.5
- Opus 5.5
- Sonnet 5
- GPT-6 Sol
3035404550550Score (%)0.250.501251020Cost per task (USD, log scale)LowMedHighXhighMax
30 35 40 45 50 55 得分 (%) 0.25 0.50 1 2 5 10 单次任务成本 (美元,对数刻度) 低 中 高 极高 最大
FrontierCode measures whether an agent’s code changes would be merged. At High effort, the default on the Claude Platform, Sonnet 5.5 matches GPT-6 Sol’s best score for about a fifth of the cost per task.²
FrontierCode 衡量智能体的代码更改是否会被合并。在高努力程度(Claude 平台上的默认设置)下,Sonnet 5.5 达到了与 GPT-6 Sol 最佳得分相当的水平,而单次任务成本仅为约五分之一。²
CursorBench 4.0Accuracy vs. cost
CursorBench 4.0 准确率与成本对比
- Sonnet 5.5
- Opus 5.5
- Sonnet 5
- GPT-5.6 Sol
20304050600Score (%)0.5012510Cost per task (USD, log scale)LowMedHighXhighMax
20 30 40 50 60 0 得分 (%) 0.50 1 2 5 10 单次任务成本 (美元,对数刻度) 低 中 高 极高 最大
CursorBench evaluates coding agents on ambiguous, multi-file tasks taken from real Cursor sessions. Sonnet 5.5 at Low effort exceeds Sonnet 5’s best score for less than a tenth of the cost per task.
CursorBench 评估编码智能体在处理来自真实 Cursor 会话的模糊多文件任务时的表现。在低努力程度下,Sonnet 5.5 以不到十分之一的单次任务成本,超越了 Sonnet 5 的最佳得分。
CursorBench 4.0 does not report GPT-6 Sol performance publicly, so we report GPT-5.6 Sol here.
CursorBench 4.0 未公开报告 GPT-6 Sol 的性能,因此我们在此报告 GPT-5.6 Sol 的数据。
AA-Briefcase v1.1Accuracy vs. cost
AA-Briefcase v1.1 准确率与成本对比
- Sonnet 5.5
- Opus 5.5
- Sonnet 5
- GPT-6 Sol
900110013001500170019000Elo0.100.200.501251020Cost per task (USD, log scale)LowMedHighXhighMax
900 1100 1300 1500 1700 1900 0 Elo 0.10 0.20 0.50 1 2 5 10 20 单次任务成本 (美元,对数刻度) 低 中 高 极高 最大
On AA-Briefcase, a new benchmark of long-horizon knowledge work, Sonnet 5.5 at Medium effort bests Sonnet 5’s best score for about one ninth of the cost per task.³
在 AA-Briefcase 这一新的长周期知识工作基准测试中,Sonnet 5.5 在中努力程度下以约九分之一的单次任务成本,超越了 Sonnet 5 的最佳得分。³
Coding
编码
Sonnet 5.5’s jump in performance is particularly noticeable in coding. At High effort on FrontierCode, it scores 10 points higher than Sonnet 5 at the same setting, at about one fifteenth of the cost per task. On CursorBench, which tests models on tasks from real Cursor coding sessions, its best score is within about two points of Opus 5.5.
Sonnet 5.5 在编码方面的性能提升尤为显著。在 FrontierCode 的高努力设置下,其得分比同设置下的 Sonnet 5 高出 10 分,而每项任务的成本仅为后者的约十五分之一。在 CursorBench(该基准测试基于真实的 Cursor 编码会话任务来评估模型)上,其最佳得分与 Opus 5.5 相差约两分。
Early testers appreciated how quickly Sonnet 5.5 can understand a codebase. They were also struck by its efficiency: in head-to-head runs, it batched tool calls together more than Sonnet 5, leading to fewer steps and lower costs.
早期测试者赞赏 Sonnet 5.5 理解代码库的速度之快。他们也对它的效率印象深刻:在直接对比运行中,它将工具调用批量处理得比 Sonnet 5 更多,从而减少了步骤并降低了成本。
Epic GamesEveryCodeRabbitSpaceXAIBase44UnityCreator
Epic GamesEveryCodeRabbitSpaceXAIBase44UnityCreator
Quote
引言
“In Epic’s early testing, Claude Sonnet 5.5 cleared the same quality bar you’d expect from a higher-tier model, holding up on a system design audit and a data flow review. The new model managed tens of thousands of lines of code for gameplay system architecture, kept responses snappy, handled multi-hour tasks, and delivered with less prescriptive prompting.”
“在 Epic 的早期测试中,Claude Sonnet 5.5 达到了你预期中更高级别模型所具备的质量标准,在系统设计审计和数据流审查中表现稳健。新模型能够处理数万行用于游戏系统架构的代码,保持响应敏捷,处理耗时数小时的任务,并以较少的指令性提示交付成果。”
CompanyEpic Games
公司Epic Games
AuthorDaniel Vogel, Chief Operating Officer
作者Daniel Vogel,首席运营官
Quote
引言
“Claude Sonnet 5.5 cooks. Fast at coding and can be steered quickly in iterative workflows. But it can still work long if it needs to. It’s got some of Opus 5.5’s natural writing upgrades, which makes it more fun to work with.”
“Claude Sonnet 5.5 表现出色。编码速度快,且在迭代工作流中能迅速调整方向。但如果需要,它也能长时间工作。它具备 Opus 5.5 的一些自然写作升级特性,这使得协作更加愉快。”
CompanyEvery
公司Every
AuthorTyler Nishida, Designer
作者Tyler Nishida,设计师
Quote
引言
“Claude Sonnet 5.5 shows better judgment than Sonnet 5 across different levels of complexity, while spending significantly fewer output tokens. Sonnet 5’s tendency to reach for web search too often and its high token use are both gone in this new model. We plan to move simple and moderate reviews over now, and more in the coming weeks.”
“Claude Sonnet 5.5 在不同复杂度级别下展现出比 Sonnet 5 更好的判断力,同时显著减少了输出 token 的使用量。Sonnet 5 过于频繁地依赖网络搜索以及高 token 使用的问题在新模型中均已消失。我们计划现在将简单和中等复杂度的审查任务迁移过来,并在未来几周内迁移更多任务。”
CompanyCodeRabbit
公司CodeRabbit
AuthorDavid Loker, VP of AI
作者David Loker,AI 副总裁
Quote
引言
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力