跳到主内容
@wquguru
精选88Latent Space(RSS)模型发布/更新多源精选 ×7

Anthropic发布Claude Haiku 5.5

[AINews] Claude Haiku 5.5 — better than GPT-6 Luna at the same pricing

原文
发到 X
推荐理由

Haiku系列时隔一年重大更新,直接对标OpenAI竞品且定价策略激进,开发者需关注其分层计费与Token效率变化。

It’s been about a year since Anthropic shipped Haiku 4.5, and with successive launches of Sonnet and Opus and Fable up to 5.5 it was seeming a little forgotten, especially as OpenAI launched Luna 6 alongside Astra and Sol 6.

Anthropic 发布 Haiku 4.5 至今约有一年,随着 Sonnet、Opus 以及 Fable 逐步更新至 5.5 版本,它似乎有些被遗忘,尤其是当 OpenAI 推出 Luna 6 并伴随 Astra 和 Sol 6 一同发布时。

Well, it’s here, and it’s a welcome update. More in the summary below.

不过,它已经来了,这是一次令人欢迎的更新。更多细节见下方摘要。

AI News for 10/06/2026-10/7/2026. We checked 12 subreddits, 544 Twitters and no further Discords. AINews’ website lets you search all past issues. As a reminder, AINews is now a section of Latent Space. You can opt in/out of email frequencies!

2026年10月6日至10月7日的 AI 新闻。我们检查了 12 个 subreddit、544 条推文,且没有进一步的 Discord 动态。AINews 的网站允许你搜索所有过往期号。提醒一下,AINews 现在是 Latent Space 的一个板块。你可以选择订阅或退订邮件频率!

AI Twitter Recap

AI Twitter 回顾

Top Story: Anthropic releases Claude Haiku 5.5

头条新闻:Anthropic 发布 Claude Haiku 5.5

What happened

发生了什么

Anthropic shipped Claude Haiku 5.5, its first Haiku-tier update in about a year. It is priced to match OpenAI’s GPT-6 Luna, and Anthropic cut prices on Sonnet 5.5 and its subscription plans the same day.

Anthropic 发布了 Claude Haiku 5.5,这是其 Haiku 系列约一年来的首次更新。其定价旨在与 OpenAI 的 GPT-6 Luna 持平,Anthropic 在同一天也降低了 Sonnet 5.5 及其订阅计划的价格。

  • Pre-launch signals. @scaling01 posted “happy Haiku 5.5 day” before the announcement. @kimmonismus said the model had already appeared in a Claude Code update and predicted “Luna-pricing.” He then posted the pricing ahead of the official post (@kimmonismus) and confirmed when it went live (@kimmonismus).
  • Official launch. @claudeai and @AnthropicAI called it “the cheapest, fastest, and most capable small model we’ve ever released,” costing about 75% less to run than Haiku 4.5 on average.
  • Availability and intended use. It is live on the Claude Platform and in Claude Code (@ClaudeDevs). Anthropic positions it as a subagent paired with Opus 5.5 or Sonnet 5.5, for high-volume, cost-sensitive work such as summaries, compactions and database queries. @mikeyk described the split as “Opus does the heavy thinking, Haiku does the high-volume work.”
  • Tiered pricing (@ClaudeDevs):
  • Prompt lengthInput / output per 1M tokensCache reads per 1MUnder 100K tokens$0.10 / $0.50$0.01Over 100K tokens$0.50 / $2.50$0.05
  • Sonnet 5.5 price cut. Cache reads were halved from $0.20 to $0.10 per 1M tokens. Anthropic says this makes Sonnet 5.5 about 20% cheaper on most long-running or agentic work (@claudeai).
  • API credits for subscribers. Monthly Claude Platform API credits now come with Max 5x ($100), Max 20x ($200) and Team (up to $500, pooled). They work on any model, including Haiku 5.5, and in third-party harnesses (@ClaudeDevs).
  • Same-day SDK update. Computer-use and browser-use toolsets are now built into the Python and TypeScript Claude SDKs. The SDK runs the action loop and sends clicks and keystrokes to drivers from browser_use, Browserbase, E2B or Daytona, so developers no longer write that loop themselves (@ClaudeDevs, quickstart).
  • Partner rollouts on day one:
  • Cursor: @cursor_ai claims “10x less than Haiku 4.5” on shorter requests and published CursorBench comparisons (@cursor_ai).
  • GitHub Copilot in VS Code: @code reports it “matched Claude Sonnet 5 on many coding tasks while using fewer tokens and steps.”
  • Devin: @cognition reports 58.4% on FrontierCode 1.1, ahead of Sonnet 5 at roughly one-eighth the cost per task, and recommends it as a “sidekick” under an Opus 5.5 lead in Fusion.
  • Arena: added to Agent Arena, Code Arena WebDev, Text, Document and Vision, with scores pending (@arena).
  • OpenDocRouter: added the same day (details below).
  • 发布前信号。@scaling01 在公告发布前发帖称“Happy Haiku 5.5 Day”。@kimmonismus 表示该模型已出现在 Claude Code 的更新中,并预测了其“Luna 式定价”。随后他在官方帖子之前发布了定价信息(@kimmonismus),并在上线后确认了这一点(@kimmonismus)。
  • 正式发布。@claudeai 和 @AnthropicAI 称其为“我们发布过的最便宜、最快且能力最强的小型模型”,平均运行成本比 Haiku 4.5 低约 75%。
  • 可用性与预期用途。它已在 Claude Platform 和 Claude Code(@ClaudeDevs)上线。Anthropic 将其定位为与 Opus 5.5 或 Sonnet 5.5 配合使用的子代理,用于处理高吞吐量、对成本敏感的任务,如摘要生成、压缩和数据库查询。@mikeyk 将这种分工描述为“Opus 负责重型思考,Haiku 负责高吞吐量工作。”
  • 分层定价(@ClaudeDevs):
  • 提示长度输入/输出每百万词元缓存读取每百万词元低于 10 万词元 $0.10 / $0.50 $0.01高于 10 万词元 $0.50 / $2.50 $0.05
  • Sonnet 5.5 降价。缓存读取价格从每百万词元 0.20 美元减半至 0.10 美元。Anthropic 表示这使得 Sonnet 5.5 在大多数长期运行或智能体工作中便宜约 20%(@claudeai)。
  • 订阅者的 API 额度。月度 Claude Platform API 额度现在包含 Max 5x(100 美元)、Max 20x(200 美元)和 Team(最高 500 美元,可共享)。它们适用于任何模型,包括 Haiku 5.5,也可在第三方框架中使用(@ClaudeDevs)。
  • 当日 SDK 更新。Computer-use 和 browser-use 工具集现已内置于 Python 和 TypeScript Claude SDK 中。SDK 运行操作循环,并向来自 browser_use、Browserbase、E2B 或 Daytona 的驱动程序发送点击和键盘输入,因此开发者无需再自行编写该循环(@ClaudeDevs,快速入门)。
  • 合作伙伴首日推出:
  • Cursor:@cursor_ai 声称在较短请求上“比 Haiku 4.5 少用 10 倍”,并发布了 CursorBench 对比结果(@cursor_ai)。
  • VS Code 中的 GitHub Copilot:@code 报告称其在许多编码任务上“与 Claude Sonnet 5 表现相当,同时使用的 token 数和步骤更少。”
  • Devin:@cognition 报告其在 FrontierCode 1.1 上得分为 58.4%,领先于 Sonnet 5,且每项任务的成本约为后者的八分之一,并推荐其作为 Fusion 中 Opus 5.5 主导下的“副手”。
  • Arena:已加入 Agent Arena、Code Arena、WebDev、Text、Document 和 Vision,分数待定(@arena)。
  • OpenDocRouter:同日添加(详情见下文)。

Independent evaluation: Artificial Analysis

独立评估:Artificial Analysis

@ArtificialAnlys published the most detailed third-party numbers (per-eval breakdown, comparison page).

@ArtificialAnlys 发布了最详细的第三方数据(每次评估的细分、对比页面)。

  • Intelligence Index: 43 at max effort, up 26 points from the previous Haiku.
  • Slightly ahead of GLM-5.3 Flash (42), Gemini 3.8 Flash (41) and GPT-6 Luna (38).
  • Comparable to Kimi K3 (44), a 2.8T-parameter open-weights model.
  • Trails Claude Sonnet 5.5 at max effort (56) by 13 points.
  • New controls. This is the first Haiku with Anthropic’s effort settings and adaptive thinking.
  • Token usage is the main caveat.
  • At max effort it uses about 162k output tokens per Index task, roughly 3x GPT-6 Luna at max (~50k).
  • Going from xhigh to max adds 2 points for about 1.8x the tokens.
  • At equal score it is still more verbose: Haiku 5.5 at high effort scores 38 using ~55k tokens, versus Luna at max scoring 38 with ~50k. The gap widens at lower effort settings.
  • Cost figures are provisional. Artificial Analysis does not yet model the 5x price step above 100K tokens. Cost-per-task numbers will follow.
  • AA-Briefcase (private agentic knowledge-work eval): 1578 Elo. That is ahead of Kimi K3 and GLM-5.3, and comparable to Muse Spark 1.3 at max.
  • Terminal-Bench 4.0: 33%, up from 0% for Haiku 4.5.
  • Level with GLM-5.3 Flash.
  • Ahead of Gemini 3.8 Flash (20%) and GPT-6 Luna (13%).
  • Knowledge versus hallucination (AA-Omniscience):
  • ModelAccuracyHallucination rateHaiku 5.536%40%Gemini 3.8 Flash55%55%GPT-6 Luna44%77%
  • Part of Haiku’s lower accuracy comes from being more willing to say it doesn’t know.
  • AutomationBench-AA: 35%, versus 53–60% for Luna, Gemini 3.8 Flash and GLM-5.3 Flash. A pre-release safety bug caused the model to over-refuse. Anthropic is working on a fix, and Artificial Analysis will re-run the eval and expects the score to rise.
  • Specs:
  • 1M-token context, up from 200k for Haiku 4.5.
  • Text and image input, text output.
  • 5-minute cache writes cost $0.125 per 1M tokens ($0.625 above 100K).
  • 智能指数:全力模式下为 43,较上一代 Haiku 提升 26 分。
  • 略高于 GLM-5.3 Flash(42)、Gemini 3.8 Flash(41)和 GPT-6 Luna(38)。
  • 与 Kimi K3(44)相当,后者是一个拥有 2.8T 参数的开放权重模型。
  • 全力模式下落后于 Trails Claude Sonnet 5.5(56),差距为 13 分。
  • 新控制功能。这是首款配备 Anthropic 努力设置和自适应思考功能的 Haiku。
  • 主要注意事项在于 token 使用量。
  • 在全力模式下,每个 Index 任务约使用 162k 输出 token,大致是全力模式下 GPT-6 Luna(约 50k)的 3 倍。
  • 从 xhigh 提升到 max 可增加 2 分,但 token 用量增加约 1.8 倍。
  • 在得分相同的情况下,Haiku 仍然更冗长:Haiku 5.5 在高努力下得分为 38,使用约 55k token;而 Luna 在全力模式下得分为 38,使用约 50k token。在较低的努力设置下,这一差距会进一步扩大。
  • 成本数据为暂定。Artificial Analysis 尚未对超过 100K token 后的 5 倍价格跳跃进行建模。每项任务的成本数据将随后公布。
  • AA-Briefcase(私有代理知识工作评估):1578 Elo。这领先于 Kimi K3 和 GLM-5.3,并与全力模式下的 Muse Spark 1.3 相当。
  • Terminal-Bench 4.0:33%,较 Haiku 4.5 的 0% 有所提升。
  • 与 GLM-5.3 Flash 持平。
  • 领先于 Gemini 3.8 Flash(20%)和 GPT-6 Luna(13%)。
  • 知识 vs 幻觉(AA-Omniscience):
  • 模型准确率 幻觉率 Haiku 5.536%40%Gemini 3.8 Flash55%55%GPT-6 Luna44%77%
  • Haiku 准确率较低的部分原因在于它更愿意承认自己不知道。
  • AutomationBench-AA:35%,而 Luna、Gemini 3.8 Flash 和 GLM-5.3 Flash 为 53–60%。一个预发布安全漏洞导致模型过度拒绝。Anthropic 正在修复该问题,Artificial Analysis 将重新运行评估,预计分数会上升。
  • 规格:
  • 1M-token 上下文,较 Haiku 4.5 的 200k 有所提升。
  • 支持文本和图片输入,输出文本。
  • 5分钟缓存写入成本为每 1M token $0.125(超过 100K 部分为 $0.625)。

Other benchmark claims (mostly vendor or secondhand)

其他基准测试声明(多为厂商提供或二手信息)

  • @ShayneRedford summarized Anthropic’s reported jumps:
  • OSWorld (computer use): 15% → 72%.
  • TerminalBench: 0% → 39%. This differs from Artificial Analysis’s independent 33% on Terminal-Bench 4.0.
  • 10–50% gains in knowledge work and reasoning.
  • Beats Luna on most of these.
  • 1M context with roughly 12k max output tokens.
  • @alexalbert__ (Anthropic) stressed that Haiku 4.5 shipped Oct 15, 2025, so the comparison spans less than a year.
  • @TheRundownAI reported that it beats GPT-6 Luna “across a variety of benchmarks.”
  • Document parsing (independent, on ParseBench):
  • @LoganMarkewich: overall close to Luna, slightly better on tables, worse on chart understanding.
  • @jerryjliu0: about $1.2 per 1,000 pages. Good at tables and reading order for the price; weaker on charts, semantic formatting and bounding boxes.
  • Anecdotal:
  • @simonw wrote pricing notes and ran his pelican-on-a-bicycle test. He says it is “SO MUCH better” than Haiku 4.5, which costs 10x more (comparison).
  • @AI_Screening says Haiku 5.5 “cooked” Luna on a Three.js zebra simulation (single prompt, not systematic).
  • @ShayneRedford 总结了 Anthropic 报告的跃升幅度:
  • OSWorld(计算机使用):15% → 72%。
  • TerminalBench:0% → 39%。这与 Artificial Analysis 在 Terminal-Bench 4.0 上独立的 33% 结果不同。
  • 知识工作和推理能力提升 10–50%。
  • 在大多数此类测试中优于 Luna。
  • 1M 上下文,最大输出 token 数约为 12k。
  • @alexalbert__(Anthropic)强调 Haiku 4.5 于 2025 年 10 月 15 日发布,因此比较跨度不到一年。
  • @TheRundownAI 报道称它在“多种基准测试”中击败了 GPT-6 Luna。
  • 文档解析(独立测试,基于 ParseBench):
  • @LoganMarkewich:整体接近 Luna,表格方面略好,图表理解方面较差。
  • @jerryjliu0:每 1,000 页约 $1.2。价格下表格和阅读顺序表现良好;但在图表、语义格式化和边界框方面较弱。
  • 轶事证据:
  • @simonw 撰写了定价笔记并运行了他的“骑自行车的鹈鹕”测试。他表示 Haiku 5 比 Haiku 4.5 “好得多”,后者成本高 10 倍(对比)。
  • @AI_Screening 表示 Haiku 5.5 在一个 Three.js 斑马模拟中(单次提示,非系统性测试)“碾压”了 Luna。

Opinions and reactions

观点与反应

Bullish

看涨

  • @kimmonismus: “Way better than GPT-6-Luna, close[r] to Sonnet 5.5… Cheap and smart.”
  • @theo likes the tiered pricing: charging a fifth of the price under 100K tokens “makes it really clear what the model is for.” He also found it striking to see an Anthropic model “so far to the left on the cost/intelligence charts” (@theo) and covered the launch on stream (@theo).
  • @draecomino: “Haiku at max effort performs like a frontier model.”
  • @kipperrii argues small, cheap models matter more now that they can do “a ton of useful things.”
  • @scaling01 (”cheap af”) and later (@scaling01): “5.5 models are looking good.”
  • @NotTomBrown (”small but mighty”) and @edwinarbus, both from the Anthropic side, posted celebratory notes.
  • @kimmonismus:‘比 GPT-6-Luna 好得多,更接近 Sonnet 5.5……便宜又聪明。’
  • @theo 喜欢分层定价:在 10 万 token 以下收费仅为五分之一,“这能非常清楚地表明该模型的用途。”他还惊讶地发现 Anthropic 的模型“在成本/智能图表上如此靠左”(@theo),并在直播中报道了此次发布(@theo)。
  • @draecomino:‘Haiku 全力运行时的表现堪比前沿模型。’
  • @kipperrii 认为,既然小型廉价模型现在能完成“大量有用的工作”,它们的重要性更加凸显。
  • @scaling01(“极其便宜”)以及稍后(@scaling01):“5.5 系列模型看起来不错。”
  • 来自 Anthropic 方面的 @NotTomBrown(“小巧但强大”)和 @edwinarbus 都发布了庆祝动态。

Competitive framing

竞争框架

  • The launch is widely read as aimed at OpenAI’s GPT-6 Luna: “rip gpt 6 luna” (@dejavucoder), “time to cook Luna” (@scaling01).
  • @kimmonismus framed the Sonnet cache-read cut as Anthropic pressuring OpenAI. On the API credits he added, “OpenAI: your turn” (@kimmonismus).
  • @teortaxesTex says Anthropic now has “the deepest product lineup of all labs” (Haiku/Sonnet/Opus/Fable plus Mythos) versus OpenAI’s Luna/Sol/Astra. He still thinks Anthropic “cares less about products,” which he reads as a sign of how much slack it has had through 2026.
  • ThursdAI’s @altryne questioned the middle tier: “Sonnet made sense when Opus was expensive.” The show plans to cover Haiku 5.5 (@thursdai_pod).
  • 此次发布被广泛解读为针对 OpenAI 的 GPT-6 Luna:“撕碎 gpt 6 luna”(@dejavucoder),“是时候让 Luna 成熟起来”(@scaling01)。
  • @kimmonismus 将 Sonnet 缓存读取的削减解读为 Anthropic 向 OpenAI 施压。关于 API 积分,他补充道:“OpenAI:轮到你了”(@kimmonismus)。
  • @teortaxesTex 表示,Anthropic 现在拥有“所有实验室中最丰富的产品线”(Haiku/Sonnet/Opus/Fable 加上 Mythos),而 OpenAI 则有 Luna/Sol/Astra。他仍然认为 Anthropic “对产品的关注度较低”,他将此解读为其在 2026 年之前拥有巨大余地的信号。
  • ThursdAI 的 @altryne 对中档产品提出质疑:“当 Opus 昂贵时,Sonnet 是有意义的。”该节目计划报道 Haiku 5.5(@thursdai_pod)。

Caveats (mostly from the data, not loud critics)

注意事项(主要来自数据,而非大声批评者)

  • The headline price may overstate real savings.
  • headline 价格可能夸大了实际节省的费用。

更进一步:量化金融体系

看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力

进入量化体系 →

关联讨论

同一事件的更多信源

相似阅读

关联信息,但可能不是同一事件