Anthropic发布Claude Haiku 5.5:成本降75%且支持可调
Claude Haiku 5.5
Haiku系列迎来重大迭代,成本腰斩且新增可调努力度机制,对追求极致性价比的Agent规模化部署有直接参考价值,建议关注其定价策略变化。
Introducing Claude Haiku 5.5: the cheapest, fastest, and most capable small model we’ve ever released.
介绍 Claude Haiku 5.5:这是我们发布过的最便宜、最快且能力最强的小型模型。
Claude Haiku 5.5 is designed for high-volume, cost-sensitive tasks. It reliably handles quick and repetitive workloads (like summaries, compactions, database queries, and classification requests). It pairs well with Opus 5.5 and Sonnet 5.5 as a subagent on coding work. And, since it’s also our fastest model to date, it works especially well for speed-sensitive tasks like live customer support and browser use.¹
Claude Haiku 5.5 专为高吞吐量、对成本敏感的任务而设计。它能可靠地处理快速且重复的工作负载(如摘要生成、数据压缩、数据库查询和分类请求)。它与 Opus 5.5 和 Sonnet 5.5 配合良好,可作为编码工作中的子代理。此外,由于它也是迄今为止我们最快的模型,因此特别适用于对速度敏感的任务,如实时客户支持和浏览器使用。¹
Haiku 5.5 is available at a much lower price than Haiku 4.5. On average, it now costs around 75% less to run.²
Haiku 5.5 的价格远低于 Haiku 4.5。平均而言,其运行成本降低了约 75%。²
Along with this launch, we’re making improvements to the value of our model range. We’re halving the price of Claude Sonnet 5.5’s cache reads, which means Sonnet 5.5 now runs around 20% cheaper on most agentic work. And we’re introducing a new monthly API credit for our Claude Max and Team subscribers, designed to support our users in building new agents and applications that run on the Claude Platform.
随着此次发布,我们对模型系列的价值进行了优化。我们将 Claude Sonnet 5.5 的缓存读取价格减半,这意味着在大多数智能体(agentic)工作中,Sonnet 5.5 的运行成本降低了约 20%。我们还为 Claude Max 和 Team 订阅用户推出了一项新的月度 API 额度,旨在支持我们的用户在 Claude 平台上构建新的智能体和应用程序。
Performance
性能
Here’s how Claude Haiku 5.5 performs across a range of benchmarks:
以下是 Claude Haiku 5.5 在各种基准测试中的表现:
| Haiku 5.5 | Haiku 4.5 | GPT-6 Luna | Sonnet 5.5For reference | ||
|---|---|---|---|---|---|
| Knowledge workGDPval-AA v2.1 | |||||
| Knowledge workGDPval-AA v2.1 | 1620 | 735 | 1437 | 1840 | |
| Knowledge workAA-Briefcase v1.1 | |||||
| Knowledge workAA-Briefcase v1.1 | 1578 | 614 | 1336 | 1824 | |
| Computer useOSWorld 2.1 | |||||
| Computer useOSWorld 2.1 | 72.4%Offline subset | 15.7%Offline subset | 48.9%Offline subset | 83.9%Offline subset | |
| Multidisciplinary reasoningHumanity’s Last Exam | |||||
| Multidisciplinary reasoningHumanity’s Last Exam | 45.9%no tools | 10.2%no tools | — | 56.9%no tools | |
| 57.4%with tools | 18.7%with tools | — | 64.5%with tools | ||
| Agentic codingTerminal-Bench 4.0 | |||||
| Agentic codingTerminal-Bench 4.0 | 39.2% | 0.0% | 16.4% | 70.6% | |
| Agentic codingFrontierCode 1.1 (Main) | |||||
| Agentic codingFrontierCode 1.1 (Main) | 46.4% | — | 42.4% | 52.1%Xhigh | |
| Visual reasoningChartography | |||||
| Visual reasoningChartography | 46.4%no tools | 6.4%no tools | 29.1%no tools | 61.6%no tools |
| Haiku 5.5 | Haiku 4.5 | GPT-6 Luna | Sonnet 5.5参考值 | ||
|---|---|---|---|---|---|
| 知识工作GDPval-AA v2.1 | |||||
| 知识工作GDPval-AA v2.1 | 1620 | 735 | 1437 | 1840 | |
| 知识工作AA-Briefcase v1.1 | |||||
| 知识工作AA-Briefcase v1.1 | 1578 | 614 | 1336 | 1824 | |
| 计算机使用OSWorld 2.1 | |||||
| 计算机使用OSWorld 2.1 | 72.4%离线子集 | 15.7%离线子集 | 48.9%离线子集 | 83.9%离线子集 | |
| 多学科推理Humanity’s Last Exam | |||||
| 多学科推理Humanity’s Last Exam | 45.9%无工具 | 10.2%无工具 | — | 56.9%无工具 | |
| 57.4%有工具 | 18.7%有工具 | — | 64.5%有工具 | ||
| 智能体编码Terminal-Bench 4.0 | |||||
| 智能体编码Terminal-Bench 4.0 | 39.2% | 0.0% | 16.4% | 70.6% | |
| 智能体编码FrontierCode 1.1 (Main) | |||||
| 智能体编码FrontierCode 1.1 (Main) | 46.4% | — | 42.4% | 52.1%Xhigh | |
| 视觉推理Chartography | |||||
| 视觉推理Chartography | 46.4%无工具 | 6.4%无工具 | 29.1%无工具 | 61.6%无工具 |
For details on how we run our evaluations, see the Haiku 5.5 System Card.
有关我们如何执行评估的详细信息,请参阅 Haiku 5.5 System Card。
Haiku 5.5 is our first Haiku-class model to come with an adjustable effort setting. This means that, as with our other models, users can decide whether to optimize for cost or intelligence. The charts below show how Haiku 5.5 performs on three benchmarks at each effort setting:
Haiku 5.5 是我们首款配备可调节努力程度设置的 Haiku 级模型。这意味着,与其他模型一样,用户可以选择优化成本或智能性。下图显示了 Haiku 5.5 在每个努力程度设置下在三个基准测试中的表现:
Computer use: OSWorldKnowledge work: GDPval-AAMultidisciplinary reasoning: Humanity’s Last Exam
计算机使用:OSWorld 知识工作:GDPval-AA 多学科推理:Humanity’s Last Exam
Computer use: OSWorldKnowledge work: GDPval-AAMultidisciplinary reasoning: Humanity’s Last Exam
计算机使用:OSWorld;知识工作:GDPval-AA;多学科推理:Humanity’s Last Exam
OSWorld 2.1 (offline subset)Accuracy vs. cost
OSWorld 2.1(离线子集)准确率与成本对比
- Haiku 5.5
- Haiku 4.5
- Sonnet 5.5
- GPT-6 Luna
0102030405060708090Partial-credit score (%)0.050.100.200.50125Cost per attempt (USD, log scale)LowMedHighXhighMax
0102030405060708090部分得分率 (%) 0.050.100.200.50125每次尝试成本 (美元,对数刻度) 低中高高极高最高
OSWorld 2.1 measures how well agents can operate a real computer to finish long, multi-step tasks.
OSWorld 2.1 衡量智能体操作真实计算机以完成长流程、多步骤任务的能力。
GDPval-AA v2.1Accuracy vs. cost
GDPval-AA v2.1 准确率与成本对比
- Haiku 5.5
- Haiku 4.5
- Sonnet 5.5
- GPT-6 Luna
8001000120014001600180020000Elo, as reported0.0050.010.020.050.100.200.50125Cost per task (USD, log scale)LowMedHighXhighMax
800100012001400160018002000 Elo评分(报告值) 0.0050.010.020.050.100.200.50125每项任务成本 (美元,对数刻度) 低中高高极高最高
Artificial Analysis’s GDPval-AA v2.1 evaluates agents on real-world professional work across 44 occupations.
Artificial Analysis 的 GDPval-AA v2.1 评估智能体在 44 个职业领域的实际专业工作表现。
Humanity’s Last Exam (no tools)Accuracy vs. cost
Humanity’s Last Exam(无工具)准确率与成本对比
- Haiku 5.5
- Haiku 4.5
- Sonnet 5.5
010203040506070Score (%)0.0050.010.020.050.100.200.501Cost per attempt (USD, log scale)LowMedHighXhighMax
010203040506070得分率 (%) 0.0050.010.020.050.100.200.501每次尝试成本 (美元,对数刻度) 低中高高极高最高
Humanity’s Last Exam (HLE) is a test of expert-level academic knowledge and reasoning.
Humanity’s Last Exam (HLE) 是一项针对专家级学术知识与推理能力的测试。
In early testing, our customers reported results consistent with the performance and cost improvements shown above. Here’s what they told us about the new model:
在早期测试中,我们的客户反馈的结果与上述性能和成本改进一致。以下是他们关于新模型的评价:
AsanaHubSpotAlphaSenseBoxRogoCognition
Asana HubSpot AlphaSense Box Rogo Cognition
AsanaHubSpotAlphaSenseBoxRogoCognition
Asana HubSpot AlphaSense Box Rogo Cognition
Quote
引言
“We’re very impressed with Claude Haiku 5.5, particularly its speed. We ran it through our eval suite for AI Teammates, our AI agent product, covering use cases like triaging bugs, setting up projects, and searching large portfolios to surface high-risk or overdue work. Compared with the model we use today, we saw over a 30% reduction in latency for task completions and up to 2.5x faster inference per agent turn. It’s a noticeably snappier experience.”
“我们对 Claude Haiku 5.5 印象深刻,尤其是其速度。我们将其纳入 AI Teammates(我们的 AI 智能体产品)的评估套件中进行测试,涵盖用例包括分类 bug、设置项目以及搜索大型投资组合以凸显高风险或逾期工作。与当前使用的模型相比,任务完成延迟降低了 30% 以上,每个智能体回合的推理速度最高提升了 2.5 倍。体验明显更加敏捷。”
CompanyAsana
公司 Asana
AuthorAaron Vinh, Staff Software Engineer
作者 Aaron Vinh,资深软件工程师
Quote
引言
“At HubSpot, we use simulated portals to evaluate new models on CRM tasks like reporting on deals. We mostly test the smaller, more efficient models, and Claude Haiku 5.5 got the best score we’ve seen on this suite yet, at 92.8% averaged over three runs. One CRM audit task asks models to identify stale but ambiguous records. Across all of the models we tested, Haiku 5.5 was fastest to complete the task, and had the highest hit rate and the lowest false positive rate.”
“在 HubSpot,我们使用模拟门户来评估新模型在 CRM 任务(如交易报告)上的表现。我们主要测试更小、更高效的模型,Claude Haiku 5.5 在该套件中取得了迄今为止最高的分数,三次运行的平均分为 92.8%。其中一项 CRM 审计任务要求模型识别陈旧但模糊的记录。在我们测试的所有模型中,Haiku 5.5 完成任务的速度最快,命中率最高,且误报率最低。”
CompanyHubSpot
公司 HubSpot
AuthorZe’ev Klapow, Distinguished Software Engineer
作者 Ze’ev Klapow,杰出软件工程师
Quote
引言
“Ask in Document is one of our big sources of spend, doing about 8M calls a week in production. It answers very specific questions on top of one or a few documents. We ran 400 queries, and Claude Haiku 5.5 was a statistically significant improvement over Haiku 4.5: 0.84 vs. 0.76.”
“Ask in Document 是我们主要的支出来源之一,在生产环境中每周执行约 800 万次调用。它针对一份或多份文档回答非常具体的问题。我们运行了 400 个查询,Claude Haiku 5.5 相较于 Haiku 4.5 有统计学意义上的显著提升:得分从 0.76 提高到 0.84。”
CompanyAlphaSense
公司 AlphaSense
AuthorDaniel Campos, Distinguished Engineer
作者 Daniel Campos,杰出工程师
Quote
引言
“Our customers use Box AI across large volumes of their enterprise content. With widespread usage comes the need to manage efficiency and cost, and to find the best model to suit the task at hand. In early testing, Claude Haiku 5.5 scored 11 points higher than Haiku 4.5 at about half the latency. We’d put it to use on analytical work that runs at scale, from cost reports to financial summaries and weekly recurring reviews.”
“我们的客户在大量企业内容中使用 Box AI。随着广泛使用,需要管理效率和成本,并找到最适合手头任务的模型。在早期测试中,Claude Haiku 5.5 的得分比 Haiku 4.5 高出 11 分,而延迟约为后者的一半。我们将把它用于大规模运行的分析工作,从成本报告到财务摘要再到每周例行审查。”
CompanyBox
公司 Box
AuthorYashodha Bhavnani, VP of AI Products
作者 Yashodha Bhavnani,AI 产品副总裁
Quote
引言
“The short and high-volume work is where Claude Haiku 5.5 fits for us, like quick lookups, subagents, and summaries. While a bigger model builds the deck, a Haiku 5.5 subagent goes into the 10-K and pulls the segment revenue line the deck needs. It’s accurate enough that we’d trust it there, and fast and cheap enough that we can run it a lot.”
“简短且高容量的工作是 Claude Haiku 5.5 在我们的场景中发挥作用的地方,例如快速查找、子智能体和摘要。当更大的模型构建演示文稿时,Haiku 5.5 子智能体会深入 10-K 表格,提取演示文稿所需的分部收入行。其准确性足以让我们信任它在此类场景中的应用,且速度快、成本低,使我们能够频繁运行它。”
CompanyRogo
公司 Rogo
AuthorAlex Wang, Applied AI
作者 Alex Wang,应用 AI
Quote
引言
“Claude Haiku 5.5 joins the sidekick lineup in Devin Fusion as an excellent option. With Haiku 5.5 as the sidekick, Fusion holds a top-tier FrontierCode score of 66.2 while cutting cost and latency. You can try it today in the Devin CLI with Opus 5.5 as the lead.”
“Claude Haiku 5.5 加入 Devin Fusion 的助手阵容,成为一款出色的选择。以 Haiku 5.5 作为助手,Fusion 在 FrontierCode 评分中位居顶级,得分为 66.2,同时降低了成本和延迟。您可以在今天的 Devin CLI 中尝试使用它,其中 Opus 5.5 作为主导模型。”
CompanyCognition
公司认知
AuthorWalden Yan, Co-Founder & CPO
作者 Walden Yan,联合创始人兼首席产品官
Pricing
定价
The table below shows how Claude Haiku 5.5’s pricing compares to our other models. Haiku 5.5 is especially good value when used for tasks with prompts up to 100,000 tokens, which make up around 90% of requests to our previous Haiku model.
下表展示了 Claude Haiku 5.5 的定价与我们其他模型的对比情况。当用于提示词长度高达 100,000 token 的任务时,Haiku 5.5 具有极高的性价比,这类任务占我们之前 Haiku 模型请求量的约 90%。
| Price per 1 million tokens | Haiku 5.5prompts up to / over 100k | Haiku 4.5 | Sonnet 5.5 |
|---|---|---|---|
| Cache reads | $0.01 / $0.05 | $0.10 | $0.10 |
| Cache writes | $0.125 / $0.625 | $1.25 | $2.50 |
| Input tokens | $0.10 / $0.50 | $1.00 | $2.00 |
| Output tokens | $0.50 / $2.50 | $5.00 | $10.00 |
| 每百万 token 价格 | Haiku 5.5(提示词长度 100k 以内/以上) | Haiku 4.5 | Sonnet 5.5 |
|---|---|---|---|
| 缓存读取 | $0.01 / $0.05 | $0.10 | $0.10 |
| 缓存写入 | $0.125 / $0.625 | $1.25 | $2.50 |
| 输入 token | $0.10 / $0.50 | $1.00 | $2.00 |
| 输出 token | $0.50 / $2.50 | $5.00 | $10.00 |
Safety
安全性
Alignment. Claude Haiku 5.5 shows major improvements across almost all of our alignment evaluations relative to Haiku 4.5. In particular, we found far fewer instances of misaligned behavior, and a lower willingness to cooperate with misuse. The model’s system card describes our evaluation process and results in more detail.
对齐性。与 Haiku 4.5 相比,Claude Haiku 5.5 在我们几乎所有的对齐评估中都显示出显著改进。特别是,我们发现其出现行为偏差的情况大幅减少,且配合不当使用的意愿更低。该模型的系统卡片详细描述了我们评估的过程和结果。
Safeguards. Consistent with its capabilities, Haiku 5.5’s cybersecurity safeguards are more restrictive than Haiku 4.5’s, but somewhat less restrictive than those we’ve applied to other recent models. In cybersecurity, they permit a wider range of defensive tasks than our safeguards for Sonnet 5.5, but they still block penetration testing and other techniques more likely to be used by attackers.
安全护栏。与其能力相一致,Haiku 5.5 的网络安全防护比 Haiku 4.5 更为严格,但比我们应用于其他近期模型的防护稍宽松一些。在网络安全方面,它们允许执行的范围比 Sonnet 5.5 的安全护栏更广泛的防御性任务,但仍会阻止渗透测试及其他更可能被攻击者使用的技术。
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力