Guides

Claude Haiku 5.5: Pricing, Benchmarks, and Migration Changes from Haiku 4.5

11 min read#claude#claude-haiku#claude-api#llm-pricing#model-migration

Who this is forDevelopers who run high-volume summarization, classification, or support tasks through the Claude API and want to know how switching to Haiku 5.5 changes cost and code.

TL;DR: Claude Haiku 5.5, released October 7, 2026, is about 75% cheaper on average than Haiku 4.5, its context grew to 1 million tokens, and you can choose how much it thinks (effort). Anthropic positioned it for high-volume work such as summarization, classification, and subagents, and stated directly that complex agentic coding such as Terminal-Bench 4.0 is better handled by Sonnet 5.5 and Opus 5.5. Moving from Haiku 4.5 requires changing a few parts of your code.

Contents

  1. What Was Released
  2. Pricing: Two Rates Based on 100K Tokens
  3. Performance: Benchmarks and Effort Levels
  4. What to Change When Migrating from Haiku 4.5
  5. Also Announced
  6. Safeguards
  7. What Is Not in the Announcement

On October 7, 2026, Anthropic announced Introducing Claude Haiku 5.5. The information below is based on that announcement and the developer documentation published the same day (Haiku 5.5 overview, What’s New, Migration guide).

1. What Was Released

The announcement describes Haiku 5.5 as the cheapest, fastest, and most capable small model Anthropic has released so far. It lists these uses:

  • Fast, repetitive high-volume work: summarization, compressing long conversations, database lookups, classification
  • Subagents for coding: a companion coding helper that works alongside Opus 5.5 and Sonnet 5.5. A subagent is a helper agent that the main model hands pieces of work to.
  • Speed-sensitive work: real-time customer support, browser use
Item Haiku 4.5 Haiku 5.5
Model ID (Claude API) claude-haiku-4-5-20251001 (alias claude-haiku-4-5) claude-haiku-5-5
Context 200K tokens 1M tokens
Max output 64K tokens 128K tokens
Thinking mode On/off, with a manually set token budget The model adjusts on its own; effort sets intensity (default medium)
Knowledge cutoff February 2025 June 2026
Retirement commitment (Claude API) After October 15, 2026 After October 7, 2027
  • Haiku 5.5 is also available on the same day through Amazon Bedrock, Google Cloud, Microsoft Foundry, and Claude Platform on AWS, in addition to the Claude API. Only Bedrock uses the ID anthropic.claude-haiku-5-5.
  • A footnote in the announcement says Haiku 5.5 is the fastest when each model is compared at its standard speed, but slower than Opus run in Fast Mode.
  • Haiku 5.5 also runs in the claude.ai web, iOS, and Android apps (system prompts documentation). According to the help article, in paid plan chats (Pro, Max, Team, Enterprise) Haiku 5.5’s context is 1M tokens, and the effort selector in the chat interface is also available for Haiku 5.5.
  • The Claude Code model configuration docs say the haiku alias points to Haiku 5.5 on the Anthropic API, and to Haiku 4.5 on Claude Platform on AWS, Bedrock, Google Cloud, and Foundry. Haiku 5.5 is used in Claude Code v2.1.293 and later.

2. Pricing: Two Rates Based on 100K Tokens

Haiku 5.5’s per-token price depends on whether a single request’s prompt is 100K tokens or less, or more than that. The announcement said that about 90% of requests sent to Haiku 4.5 were 100K tokens or less.

Pricing table from Anthropic's announcement, per million tokens. Haiku 5.5 is split into prompts of 100K tokens or less and over 100K tokens: cache read $0.01 and $0.05, cache write $0.125 and $0.625, input $0.10 and $0.50, output $0.50 and $2.50. Haiku 4.5: cache read $0.10, cache write $1.25, input $1, output $5. Sonnet 5.5: cache read $0.10, cache write $2.50, input $2, output $10
In the Haiku 5.5 column, the left side is the rate for prompts of 100K tokens or less, and the right side is the rate for prompts over 100K tokens.
  • According to the announcement’s footnote, requests of 100K tokens or less cost 90% less than on Haiku 4.5, and larger requests cost 50% less.
  • However, Haiku 5.5 uses a new tokenizer (the method for splitting text into tokens), and the documentation explains that the same text is counted as about 30% more tokens than on Haiku 4.5.
  • The announcement’s “about 75% cheaper on average” figure, according to its footnote, also accounts for a small increase in tokens per task and changes in the distribution of requests.
  • A prompt of about 77K tokens on Haiku 4.5 could exceed 100K tokens if you apply the documentation’s roughly 30% increase, which would move it into the tier that costs five times as much.
  • Caching stores and reuses the front part of a frequently used prompt. Storing it costs the cache write rate, and reading it back costs the cache read rate.
  • A 1-hour cache write costs $0.20 for requests of 100K tokens or less and $1 for larger requests. The Batch API gives 50% off input and output (per the Haiku 5.5 overview documentation).

3. Performance: Benchmarks and Effort Levels

The announcement compares Haiku 5.5 with Haiku 4.5 and OpenAI’s GPT-6 Luna, and includes Sonnet 5.5 as a reference.

Benchmark table from Anthropic's announcement, in the order Haiku 5.5, Haiku 4.5, GPT-6 Luna, and reference Sonnet 5.5. GDPval-AA v2.1: 1620, 735, 1437, 1840. AA-Briefcase v1.1: 1578, 614, 1336, 1824. OSWorld 2.1 offline subset: 72.4%, 15.7%, 48.9%, 83.9%. Humanity's Last Exam without tools: 45.9%, 10.2%, not reported, 56.9%; with tools: 57.4%, 18.7%, not reported, 64.5%. Terminal-Bench 4.0: 39.2%, 0.0%, 16.4%, 70.6%. FrontierCode 1.1 Main: 46.4%, not reported, 42.4%, 52.1% at Xhigh. Chartography without tools: 46.4%, 6.4%, 29.1%, 61.6%
The pink cells are Haiku 5.5, and the gray column on the far right is reference Sonnet 5.5.
Benchmark Announcement description
GDPval-AA v2.1 Evaluation on real work tasks across 44 occupations
OSWorld 2.1 Completing long, multi-step tasks by operating a computer
Humanity’s Last Exam Expert-level academic knowledge and reasoning
Terminal-Bench 4.0 Completing complex, multi-step work on the command line
  • The announcement gave the other three benchmarks only a category label: AA-Briefcase is knowledge work, FrontierCode is agentic coding, and Chartography is visual reasoning.
  • On every benchmark with a score, Haiku 5.5 beats Haiku 4.5, and on every benchmark where GPT-6 Luna’s score is published, it beats Luna. It scores lower than Sonnet 5.5 on every benchmark.
  • On Terminal-Bench 4.0, Haiku 5.5 scores 39.2% against Sonnet 5.5’s 70.6%, a large gap. The announcement also says that complex agentic coding such as Terminal-Bench 4.0 is still better handled by Sonnet 5.5 and Opus 5.5, and that Haiku 5.5 fits narrower jobs such as compression, summarization, and subagent work.

Haiku 5.5 is the first Haiku-class model where you can choose effort. The levels are Low, Med, High, Xhigh, and Max (API values low, medium, high, xhigh, max). The announcement graphs scores and costs at each level on three benchmarks.

OSWorld 2.1 offline subset graph from the announcement. The horizontal axis is cost per attempt in dollars (log scale), and the vertical axis is partial score (%). Haiku 5.5's points rise toward the upper right from Low to Med, High, Xhigh, and Max, reaching from the 40% range into the 70% range. GPT-6 Luna sits lower, Sonnet 5.5 sits further upper right, and Haiku 4.5 is a single point in the lower right, in the 10% range
Further right means each attempt costs more, and higher means a higher score. The pink line shows Haiku 5.5's five effort levels.
  • In the graph, Haiku 5.5’s Max level lands in the 70% range and its default Med level in the 50% range. The 72.4% in the table above looks close in height to the graph’s Max point, but the announcement does not say which effort level the table score corresponds to.
  • In the same graph, Haiku 5.5’s Max level sits further left (cheaper) than the leftmost point of Sonnet 5.5.
  • On the Claude API, if you don’t set effort, it runs at medium (per the model overview documentation).

The announcement also includes evaluations from six customers that tested Haiku 5.5 during early testing. The table below covers the four that reported numbers for Haiku 5.5 itself.

  • Cognition reported a score for its own system (Devin Fusion), which uses Haiku 5.5 as a helper model. Rogo described its use case without a number.
  • These are numbers reported by customers, and each measured performance in its own way.
Customer Evaluation in the announcement
Asana AI agent task completion delay cut by more than 30%; agent reasoning per turn up to 2.5x faster (compared with the model currently in use)
HubSpot Average of 92.8% over three runs on a CRM task evaluation; the highest score it has seen on this evaluation, which it has mostly used for small models
AlphaSense 0.84 vs. 0.76 for Haiku 4.5 on 400 document queries
Box 11 points above Haiku 4.5, with latency about half as long

4. What to Change When Migrating from Haiku 4.5

If you call the Messages API directly, there is more to check than the model ID. Some requests that worked on Haiku 4.5 and include certain parameters now return an error (400) on Haiku 5.5. If you use Claude Managed Agents, the docs say that changing the model name is enough.

The table below is the checklist from the migration guide.

What to change Why
Replace the model ID with claude-haiku-5-5 Undated fixed ID, no alias
Recalculate max_tokens and cost estimates The same text uses about 30% more tokens
Replace the budget_tokens thinking setting with adaptive The old setting returns an error
Fix code that reads the first block of the response as the answer The response can start with a thinking block
Remove temperature, top_p, and top_k top_k errors regardless of value; temperature and top_p error unless at default; sending both errors
Remove prefill that ends the last message with an assistant turn Errors even with thinking turned off
Replace the computer use tool with the new toolset On the Claude API and Google Cloud, the old tool returns errors
Replay saved conversations through the account that created them or a linked account From other accounts, thinking blocks are dropped
Do not modify earlier conversation when sending back thinking blocks Errors if earlier content changes (for accounts created before August 31, 2026 UTC, only on requests that include thinking.block_binding.prefix_mismatch_behavior)
Add handling for stop_reason: "refusal" Safety classifiers can refuse requests, and there is no server-side fallback model
  • Thinking block content comes back empty by default. To receive summarized thinking, set display to summarized.
  • If you force tool calls (setting tool_choice to any or a specific tool), the response starts directly with the tool call, with no thinking block.
  • Thinking tokens also count against max_tokens. If you set max_tokens small for Haiku 4.5, the answer can be cut off before it appears. Raise the limit or lower effort.
  • If you use Claude Code, the /claude-api migrate command in the docs applies the changes above to your project and gives you a checklist to verify.

5. Also Announced

Sonnet 5.5 cache reads cut 50%. Starting the same day, Sonnet 5.5’s cache read rate dropped from $0.20 to $0.10 per million tokens. The announcement says Sonnet 5.5 costs about 20% less for most agentic work.

Monthly API credits for Max and Team plans. Max and Team subscribers receive API credits each month for use on Claude Platform. The announcement said distribution starts within the week of the announcement date (October 7).

Amounts and conditions are based on the help article.

Plan Monthly credit
Max 5x $100
Max 20x $200
Team Standard seat $20 per seat
Team Premium seat $100 per seat
Team total cap $500
  • Free, Pro, and Enterprise are not eligible. Discounted Team plans (Nonprofit, Scientists) are eligible for the same amounts as Team.
  • You must have been subscribed for 7 days before receiving credits. Connecting one Claude Console organization in the claude.ai billing settings sends the credits to that organization.
  • The person connecting must be the subscriber for Max, or the Primary Owner or an Owner for Team. For a Console organization, anyone with the Owner, Admin, or Billing role can connect.
  • Credits refill each billing cycle, and unused amounts do not carry over to the next month. They can be used for the Claude API, Batch API, Console Playground, Managed Agents, and Agent SDK.
  • Only one Console organization can be connected. After connecting, it cannot be changed on your own; you have to contact support.
  • Credits cannot be used for Claude Code used in the terminal, IDE, desktop, or web, for extra usage, or for Bedrock, Google Cloud (Vertex AI), or Foundry.

SDK computer use and browser use in beta. Claude’s Python and TypeScript SDKs support computer use and browser use in beta. The announcement says that when you consider speed, performance, and price together, Haiku 5.5 fits this kind of work especially well. Haiku 5.5 can use the browser use tool on the Claude API and Google Cloud, while Haiku 4.5 does not support this tool.

6. Safeguards

  • Alignment: The announcement says Haiku 5.5 improved substantially over Haiku 4.5 on nearly every alignment evaluation, and that it is less prone to misaligned behavior and to cooperating with misuse. The evaluation process is described in the system card.
  • Cybersecurity: It is more restricted than Haiku 4.5 but slightly less restricted than some recent models. It allows a broader range of defensive work than Sonnet 5.5, and it blocks techniques that attackers are likely to use, such as penetration testing.
  • Biology: It has the same safeguards as Sonnet 5, Sonnet 5.5, and Opus 5. Research-oriented biology questions are allowed, and requests with a high potential for harm are restricted.
  • Organizations that need broader biology or cyber work can apply to the Life Sciences Verification Program and the Cyber Verification Program.

7. What Is Not in the Announcement

The announcement, developer documentation, and help articles checked for this post do not say:

  • Whether Haiku 5.5 can be selected on the claude.ai free plan
  • The actual retirement date for Haiku 4.5 (the model deprecations page only has the promise of retirement after October 15, 2026)
  • Separately measured performance or token growth rates for Korean-language work

Frequently asked questions

What is the context window of Claude Haiku 5.5?
Claude Haiku 5.5 has a 1M-token context, up from 200K tokens on Haiku 4.5. Its maximum output is 128K tokens.
Is Claude Haiku 5.5 cheaper than Haiku 4.5?
Yes. Anthropic says Haiku 5.5 is about 75% cheaper on average than Haiku 4.5. Per the announcement's footnote, requests of 100K tokens or less cost 90% less, and larger requests cost 50% less.