Anthropic发布Claude Haiku 5.5:1M上下文
Anthropic Releases Claude Haiku 5.5: A Small Model With 1M Context Priced at $0.10 per Million Input Tokens
Haiku系列首次降价至0.10美元级别,对Agent子任务成本结构有重大影响,建议开发者重新评估高频调用的架构选型。
Anthropic has released Claude Haiku 5.5, its cheapest and fastest small model to date. It targets high-volume work like summaries, compaction, classification and subagent tasks. It keeps a 1M token context window and up to 128K output tokens. Pricing starts at $0.10 per million input tokens and $0.50 per million output tokens. That is 90% below Claude Haiku 4.5 for prompts up to 100K tokens.
Anthropic 发布了 Claude Haiku 5.5,这是迄今为止其最便宜、速度最快的小型模型。它针对摘要、压缩、分类和子代理任务等高吞吐量工作负载。它保留了 1M token 的上下文窗口,并支持最多 128K 的输出 token。定价为每百万输入 token 0.10 美元,每百万输出 token 0.50 美元起。对于不超过 100K token 的提示词,这比 Claude Haiku 4.5 低了 90%。
Is it deployable? Yes, as a hosted API. Haiku 5.5 is generally available on the Claude API, Amazon Bedrock, Google Cloud, Microsoft Foundry and Claude Platform on AWS.
可以部署吗?可以,作为托管 API。Haiku 5.5 已在 Claude API、Amazon Bedrock、Google Cloud、Microsoft Foundry 以及 AWS 上的 Claude Platform 上全面可用。
What Anthropic Shipped
Anthropic 发布了什么
Haiku 5.5 is the first Haiku-class model with an adjustable effort setting. Adaptive thinking is on by default, and the effort parameter defaults to medium.
Haiku 5.5 是首个具有可调节努力设置(effort setting)的 Haiku 级模型。自适应思考(Adaptive thinking)默认开启,努力参数默认为中等。
It takes text and images, outputs text, and has a June 2026 knowledge cutoff. Batch jobs support up to 300K output tokens in beta.
它接受文本和图片输入,输出文本,知识截止日期为 2026 年 6 月。批量作业在测试版中支持最多 300K 输出 token。
It is important to note two key things. Non-default temperature, top_p or top_k values return a 400 error. The new tokenizer also counts the same text as roughly 30% more tokens than Haiku 4.5. The migration guide covers both.
需要注意两个关键点。非默认的 temperature、top_p 或 top_k 值会返回 400 错误。新的分词器也将相同文本计为比 Haiku 4.5 多约 30% 的 token。迁移指南涵盖了这两点。
Pricing: A 2-Tier Structure
定价:双层级结构
Pricing splits at 100K prompt tokens. Up to 100K, input costs $0.10 and output $0.50 per million. Cache reads cost $0.01 and 5-minute cache writes $0.125. Above 100K, rates rise to $0.50 input and $2.50 output.
定价在 100K 提示词 token 处分界。100K 以内,输入成本为每百万 0.10 美元,输出为每百万 0.50 美元。缓存读取成本为 0.01 美元,5 分钟缓存写入成本为 0.125 美元。超过 100K 后,费率上升至输入 0.50 美元,输出 2.50 美元。
Haiku 4.5 charged $1 input and $5 output. Anthropic says about 90% of Haiku 4.5 requests stayed under 100K tokens. After adjusting for the tokenizer, it estimates Haiku 5.5 runs about 75% cheaper on average. Batch processing takes another 50% off.
Haiku 4.5 的收费为输入 1 美元,输出 5 美元。Anthropic 表示,约 90% 的 Haiku 4.5 请求保持在 100K token 以下。经过分词器调整后,它估计 Haiku 5.5 平均运行成本低约 75%。批量处理再降低 50% 的成本。
GPT-6 Luna lists identical short-context rates. Its higher tier starts only above 272K input tokens, at $0.20 and $0.75. For a 150K-token prompt, Luna is cheaper on list price.
GPT-6 Luna 列出了相同的短上下文费率。其较高层级仅在输入 token 超过 272K 时开始适用,费率为 0.20 美元和 0.75 美元。对于 150K token 的提示词,Luna 的列表价格更便宜。
Benchmarks
基准测试
All figures below are Anthropic-reported (see the system card):
以下所有数据均由 Anthropic 报告(参见系统卡片):
- OSWorld 2.1 (offline subset): 72.4%, versus 48.9% for GPT-6 Luna and 15.7% for Haiku 4.5.
- Terminal-Bench 4.0: 39.2%, versus 16.4% for Luna and 0.0% for Haiku 4.5.
- FrontierCode 1.1 (Main): 46.4%, versus 42.4% for Luna.
- Humanity’s Last Exam: 45.9% without tools and 57.4% with tools.
- GDPval-AA v2.1: 1620, versus 1437 for Luna and 735 for Haiku 4.5.
- OSWorld 2.1(离线子集):72.4%,而 GPT-6 Luna 为 48.9%,Haiku 4.5 为 15.7%。
- Terminal-Bench 4.0:39.2%,而 Luna 为 16.4%,Haiku 4.5 为 0.0%。
- FrontierCode 1.1(主集):46.4%,而 Luna 为 42.4%。
- Humanity’s Last Exam:无工具时为 45.9%,有工具时为 57.4%。
- GDPval-AA v2.1:1620,而 Luna 为 1437,Haiku 4.5 为 735。
Sonnet 5.5 still leads every row, including 70.6% on Terminal-Bench 4.0. Anthropic itself recommends Sonnet 5.5 and Opus 5.5 for complex agentic coding.
Sonnet 5.5 仍在每一行中领先,包括在 Terminal-Bench 4.0 上达到 70.6%。Anthropic 本身也推荐 Sonnet 5.5 和 Opus 5.5 用于复杂的智能体编码任务。
Best Use Cases
最佳用例
Three workloads fit Haiku 5.5 best:
有三类工作负载最适合 Haiku 5.5:
- The first is subagent work under Opus 5.5 or Sonnet 5.5. At Rogo, a Haiku 5.5 subagent pulls a 10-K revenue line while a bigger model builds the deck.
- The second is high-volume document Q&A and summarization. AlphaSense tested it on a feature handling about 8M calls a week.
- The third is speed-sensitive work like live customer support and browser use.
- 第一种是在 Opus 5.5 或 Sonnet 5.5 下运行的子代理工作。在 Rogo,一个 Haiku 5.5 子代理负责提取 10-K 表格中的收入行,而更大的模型则负责构建演示文稿。
- 第二种是大批量文档问答和摘要任务。AlphaSense 在一个每周处理约 800 万次调用的功能中测试了该能力。
- 第三种是对速度敏感的工作,如实时客户支持和浏览器使用。
Haiku 5.5 vs Its Closest Competitors
Haiku 5.5 与其最接近的竞争对手对比
| Feature | Claude Haiku 5.5 | GPT-6 Luna | Gemini 3.5 Flash-Lite |
|---|---|---|---|
| Input / output (per 1M) | $0.10 / $0.50 | $0.10 / $0.50 | $0.30 / $2.50 |
| Long-prompt pricing | $0.50 / $2.50 above 100K | $0.20 / $0.75 above 272K | Flat rate |
| Cache read (per 1M) | $0.01 | $0.01 | $0.03 + storage |
| Context window | 1M tokens | 1,050,000 tokens | 1,048,576 tokens |
| Max output | 128K tokens | 128,000 tokens | 65,536 tokens |
| Inputs | Text, images | Text, images | Text, image, video, audio, PDF |
| Reasoning control | Adaptive thinking + effort (default medium) | reasoning.effort none to max (default medium) | Thinking supported |
| Computer use | SDK support in beta | Supported (Responses API) | Supported (Preview) |
| Batch discount | 50% | 50% (Batch and Flex) | 50% |
| Knowledge cutoff | Jun 2026 | May 18, 2026 | Not listed on model page |
| Where to run | Claude API, Bedrock, Google Cloud, Microsoft Foundry, Claude Platform on AWS | OpenAI API | Gemini API |
| 特性 | Claude Haiku 5.5 | GPT-6 Luna | Gemini 3.5 Flash-Lite |
|---|---|---|---|
| 输入/输出(每百万) | $0.10 / $0.50 | $0.10 / $0.50 | $0.30 / $2.50 |
| 长提示定价 | 超过 100K 时为 $0.50 / $2.50 | 超过 272K 时为 $0.20 / $0.75 | 固定费率 |
| 缓存读取(每百万) | $0.01 | $0.01 | $0.03 + 存储费 |
| 上下文窗口 | 1M tokens | 1,050,000 tokens | 1,048,576 tokens |
| 最大输出 | 128K tokens | 128,000 tokens | 65,536 tokens |
| 输入类型 | 文本、图像 | 文本、图像 | 文本、图像、视频、音频、PDF |
| 推理控制 | 自适应思维 + 努力程度(默认中等) | reasoning.effort 从 none 到 max(默认中等) | 支持 Thinking |
| 计算机使用 | SDK 支持(Beta 版) | 支持(Responses API) | 支持(预览版) |
| 批量折扣 | 50% | 50%(Batch 和 Flex) | 50% |
| 知识截止日期 | 2026 年 6 月 | 2026 年 5 月 18 日 | 模型页面上未列出 |
| 运行位置 | Claude API, Bedrock, Google Cloud, Microsoft Foundry, AWS 上的 Claude Platform | OpenAI API | Gemini API |
Sources: Anthropic docs, Anthropic announcement, OpenAI model page, OpenAI pricing, Google model page, Google pricing. Standard-tier list prices, verified October 7, 2026.
来源:Anthropic 文档、Anthropic 公告、OpenAI 模型页面、OpenAI 定价、Google 模型页面、Google 定价。标准层级列表价格,于 2026 年 10 月 7 日核实。
Interactive Explainer
交互式解释器
Key Takeaways
关键要点
- $0.10 input and $0.50 output per 1M tokens for prompts up to 100K tokens.
- 1M context, 128K output, adaptive thinking with an effort setting (default medium).
- 72.4% on OSWorld 2.1 versus 15.7% for Haiku 4.5 (Anthropic-reported).
- Same short-context list price as GPT-6 Luna; Luna is cheaper above 100K tokens.
- Positioned as a subagent under Opus 5.5 and Sonnet 5.5, not a coding lead.
- 对于长达 100K tokens 的提示,每百万 tokens 的输入价格为 $0.10,输出价格为 $0.50。
- 1M 上下文,128K 输出,具备带努力程度设置的自适应思维(默认中等)。
- 在 OSWorld 2.1 上得分为 72.4%,而 Haiku 4.5 为 15.7%(Anthropic 报告数据)。
- 与 GPT-6 Luna 具有相同的短上下文列表价格;但在超过 100K tokens 时,Luna 更便宜。
- 定位为 Opus 5.5 和 Sonnet 5.5 下的子代理,而非代码编写主导模型。
Check out the Technical Details. All credit goes to the researcher of this project. Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well.
查看技术细节。本项目的所有功劳归于其研究人员。此外,欢迎在 Twitter 上关注我们,并别忘了加入我们拥有 15 万+成员的 ML SubReddit 以及订阅我们的新闻通讯。等等!你在 Telegram 上吗?现在你也可以加入我们的 Telegram 群组。
[Sponsored] The web is the one API most agents are missing. Databases, calendars and repos have APIs. The open web mostly doesn’t. The TinyFish MCP server gives any MCP client four tools: TinySearch, TinyFetch (full pages as markdown, JavaScript included), TinyBrowser for logins and forms, and TinyAgent for multi-step jobs. Search and Fetch are free.
[赞助] 网页是大多数智能体缺失的唯一 API。数据库、日历和代码仓库都有 API,而开放的网络大多没有。TinyFish MCP 服务器为任何 MCP 客户端提供四项工具:TinySearch(搜索)、TinyFetch(将完整页面以 Markdown 格式获取,包含 JavaScript)、TinyBrowser(用于登录和表单操作)以及 TinyAgent(用于多步骤任务)。搜索和抓取功能免费。
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力