ANTM
ModelsAnalysis

Claude Haiku 5.5: Specs, Pricing, Migration Traps and When to Use It

Anthropic's smallest current model costs about a twentieth of Sonnet 5.5 per token, but it has a price cliff at 100,000 tokens and five breaking API changes. What the documentation says, and what it means for routing.

Editorial desk

Published 9 min read

Overlapping translucent waves of fine glowing teal lines sweeping across the bottom of a dark navy frame
Illustration generated with AI (FLUX.1 [schnell] (Black Forest Labs) via Cloudflare Workers AI, Apache 2.0). Prompt and direction by ANTM.

Claude Haiku 5.5 is Anthropic's smallest current model, released on 2026-10-07 with the API ID claude-haiku-5-5. Anthropic positions it for high-volume, latency-sensitive work such as classification, extraction, routing and subagent tasks.[4] This Model Watch reads Anthropic's announcement and documentation to answer the questions that decide whether to adopt it: what it costs, what changed in the API, what the benchmark claims do and do not show, and which jobs should go to Haiku rather than Sonnet or Opus.

All facts below come from Anthropic's pages, read on 2026-10-08. We did not run the model, so nothing here is our own testing; where we calculate or interpret, we say so. Prices and limits change, so confirm them on the linked pages before budgeting.

Quick answer

  • What it is: a small model in the 5.5 family with a 1M-token context window, 128K max output, adaptive thinking and, for the first time in a Haiku, an effort setting.[1][4]
  • What it costs: $0.10 input and $0.50 output per million tokens for prompts up to 100,000 tokens; $0.50 and $2.50 above that.[3]
  • The catch on cost: the same text counts as roughly 30% more tokens than on Haiku 4.5, so the saving is real but smaller than the per-token drop suggests.[5]
  • The catch on migration: five breaking API changes, plus a new refusal stop reason with no server-side fallback.[5][6]
  • Where it fits: narrowly scoped, high-volume work. Anthropic itself says Sonnet 5.5 and Opus 5.5 remain better for complex agentic coding.[1]

Specifications

SpecHaiku 5.5Sonnet 5.5Opus 5.5Fable 5.1
API model IDclaude-haiku-5-5claude-sonnet-5-5claude-opus-5-5claude-fable-5-1
Positioning (Anthropic)High-volume, latency-sensitiveSpeed and intelligenceLong-running agentic coding, knowledge workDemanding reasoning, long-horizon agents
Comparative latencyFastestFastModerateSlower
Context window1M tokens1M tokens1M tokens1M tokens
Max output (sync)128K tokens128K tokens128K tokens128K tokens
ThinkingAdaptiveAdaptiveAdaptive (always on)Adaptive (always on)
Default effortmediumhighmediumhigh
Input / output per MTok$0.10 / $0.50 (to 100k prompt)$2 / $10$4 / $20$10 / $50
Retirement commitmentNot sooner than 2027-10-07Not sooner than 2027-09-28Not sooner than 2027-09-22Not sooner than 2027-09-01

Source: Anthropic models overview and Haiku 5.5 model page, read 2026-10-08.[2][4] Haiku 5.5 accepts text and images and returns text, with a reliable knowledge cutoff of June 2026.[4] The retirement column is a commitment for Anthropic-operated platforms; Amazon Bedrock and Google Cloud set their own dates.[2]

Pricing, including the 100,000-token break

Haiku 5.5 is the one current Claude model with length-tiered pricing. Anthropic's pricing page says Claude 4.6 and later models include the full 1M window at standard pricing, "except Claude Haiku 5.5", which pays higher prices when the prompt exceeds 100,000 tokens.[3]

Price per million tokensPrompt up to 100kPrompt over 100k
Input$0.10$0.50
Output$0.50$2.50
5-minute cache write$0.125$0.625
1-hour cache write$0.20$1.00
Cache read$0.01$0.05
Batch input / output$0.05 / $0.25$0.25 / $1.25

Source: Anthropic pricing page, read 2026-10-08.[3] Anthropic's announcement adds that about 90% of Haiku 4.5 requests fell in the up-to-100k range, and describes Haiku 5.5 as about 90% cheaper per request than Haiku 4.5 up to 100k and 50% cheaper above it, with the tokenizer change accounted for.[1] That is Anthropic's figure for its own traffic mix, not a guarantee for yours.

Worked examples (our arithmetic)

These use the published list prices. The only assumption we add is that Haiku 5.5 counts 30% more tokens than Haiku 4.5 for the same text, which is Anthropic's "approximately" figure and varies by content.[5] We ignore thinking tokens, which are billed as output and which you would need to measure.

Example A: a classification call. A prompt of 2,000 tokens (as counted on Haiku 4.5) with a 200-token answer. On the new tokenizer that is about 2,600 in and 260 out.

ModelCost per requestCost per million requests
Haiku 4.5$0.0030$3,000
Haiku 5.5$0.00039$390
Sonnet 5.5$0.0078$7,800
Opus 5.5$0.0156$15,600

Example B: one long document. A 150,000-token prompt with a 1,000-token answer crosses the price break. Haiku 5.5 costs about $0.0775 per call ($0.075 input at $0.50, $0.0025 output at $2.50). Sonnet 5.5 costs $0.31 ($0.30 plus $0.01). Haiku is still four times cheaper here, but the gap is far narrower than in Example A, so a long-context workload that quietly crosses 100k deserves its own estimate.

Example C: a cached prefix. A 30,000-token cached system prompt plus 500 new input tokens and 300 output tokens. At cache-read prices that is about $0.0005 on Haiku 5.5 and $0.007 on Sonnet 5.5, because Sonnet 5.5 reads cache at $0.10 and Haiku 5.5 at $0.01 per million.[3] Without caching, the same Haiku 5.5 call would cost about $0.0032.

The practical reading: the price gap is largest on short, repetitive, cache-friendly requests, which is exactly the workload Haiku is aimed at. See Tokens, Context Windows and What an AI Request Really Costs for how caching multipliers and re-sent history drive the bill.

What the benchmark table does and does not show

Anthropic's announcement includes a table comparing Haiku 5.5 with Haiku 4.5, OpenAI's GPT-6 Luna and Sonnet 5.5.[1] A selection of rows:

Benchmark (as reported by Anthropic)Haiku 5.5Haiku 4.5GPT-6 LunaSonnet 5.5
GDPval-AA v2.1 (Elo)162073514371840
OSWorld 2.1 (offline subset)72.4%15.7%48.9%83.9%
Humanity's Last Exam, no tools45.9%10.2%not shown56.9%
Terminal-Bench 4.039.2%0.0%16.4%70.6%
FrontierCode 1.1 (Main)46.4%not shown42.4%52.1%

Source: Anthropic announcement, read 2026-10-08.[1] Our interpretation, based on the published numbers:

  • Generational gap: Haiku 5.5 is far ahead of Haiku 4.5 on every row shown, which is consistent with the price drop being a new model class rather than a discount.
  • Tier ordering is preserved: Sonnet 5.5 beats Haiku 5.5 on every row shown, and among the percentage-scored rows the gap is widest on Terminal-Bench 4.0 (70.6% against 39.2%). That matches Anthropic's own guidance that Sonnet and Opus remain better for complex agentic coding.[1]
  • Comparisons with GPT-6 Luna are the vendor's: we did not open OpenAI's pages in this run, so we cannot confirm the competitor figures or whether settings were matched. Anthropic says the full methodology is in the Haiku 5.5 system card, which the announcement links but does not reproduce, and we did not review it. The Sonnet 5.5 FrontierCode figure is marked as run at xhigh effort; we could not confirm the effort setting for the other cells.
  • Customer results are self-reported: the announcement quotes teams such as Box, AlphaSense, HubSpot and Asana reporting latency and accuracy gains on their own evals.[1] They show it can work on real tasks, not that it will on yours.

For how to read claims like these, see How to Read an AI Benchmark Claim Without Being Misled.

What breaks when you migrate from Haiku 4.5

Anthropic's migration guide is clear that this is more than a new model string.[6] The breaking changes:

  1. Manual thinking budgets are rejected. thinking: {"type": "enabled", "budget_tokens": N} returns a 400 error. Use {"type": "adaptive"} and set output_config.effort.[6]
  2. Sampling parameters are rejected. Omit temperature, top_p and top_k; non-default values return a 400.[6]
  3. Assistant prefill is rejected, even with thinking off. End messages with a user turn and use structured outputs or tools for format control.[6]
  4. Computer use moves to a toolset (computer_toolset_20260801) on the Claude API and Google Cloud; the old computer_20250124 tool errors.[6]
  5. Thinking blocks are bound to earlier turns. Changing system, tools or earlier messages and then sending a thinking block back returns a 400, so conversations must be append-only.[6]

Behaviour changes also matter for cost and correctness. Thinking is on by default and counts toward max_tokens, so a limit tuned for Haiku 4.5 can stop after a thinking block and before any text. Responses can begin with thinking blocks, so code must select blocks by type, and thinking text is omitted unless you set display to "summarized".[5][6] Priority Tier is not supported on Haiku 5.5.[6]

Two operational notes from the prompting guide. Haiku 5.5 runs safety classifiers that can return stop_reason: "refusal", and there is no server-side fallback, so clients must handle it; resending the same request usually refuses again.[7] And changing the top-level effort between requests invalidates the prompt cache for the conversation's messages, which interacts badly with the caching savings above unless you use the per-message effort beta.[7]

Effort: the lever that replaces the thinking budget

Haiku 5.5 is the first Haiku with effort levels. Anthropic's guidance:[7]

  • low: cheapest and fastest, for chat, short tool tasks and simple high-volume requests. In long agent prompts the model is more likely to skip a search, stop early or skip a check.
  • medium: the API default and the suggested starting point, including for agentic coding.
  • high: for knowledge work, longer agent tasks and strict instruction following.
  • xhigh and max: only where an eval gain justifies the cost; Anthropic suggests also testing Sonnet 5.5 at those levels.

The practical trade-off is that effort moves cost as well as quality. In Anthropic's own testing, moving from low to medium roughly halved early stopping in a long coding-agent prompt while more than doubling output tokens per attempt.[7] A "cheap" model run at high effort on long outputs can erode the price advantage, which is why a cost test should use your real prompts at the effort you will ship.

Which jobs go where

This is our editorial recommendation, derived from the documented specs, prices and Anthropic's own positioning, not from independent testing.

JobStart withWhy
Classification, routing, tagging, short extractionHaiku 5.5 at low or mediumAnthropic's stated target; lowest price; benchmark-test it against your labels
Summarisation and context compaction in agentsHaiku 5.5Named by Anthropic as a fit; cache reads at $0.01 per million help
Subagent that searches or reads for a lead modelHaiku 5.5, lead on Sonnet 5.5 or Opus 5.5Anthropic describes this pattern; keep the verification instructions from the prompting guide
Prompts regularly above 100k tokensCompare Haiku 5.5 and Sonnet 5.5 on cost per correct answerThe price break narrows the gap
Complex multi-step coding or terminal workSonnet 5.5 or Opus 5.5Anthropic says they remain better; Terminal-Bench gap is the widest of the percentage rows
Customer-facing chatbot with strict rulesTest Haiku 5.5 at high with the system-prompt reinforcement Anthropic suggestsThe guide says instruction-following improves at high[7]
Tasks needing assistant prefill or custom samplingPlan a prompt rewrite firstThese are rejected by the API[6]

For the larger tier comparison, including Fable 5.1 and cache multipliers, see Claude Opus 5.5 vs Sonnet 5.5 vs Fable 5.1; note that article was written before Haiku 5.5 and lists Haiku 4.5 as the small tier. For cross-vendor prices, see GPT-6.1 Sol vs Claude Sonnet 5.5 vs Gemini: API Specs and Pricing Compared, which does not include Haiku 5.5 or other vendors' small models.

A sensible adoption path

  1. Count your real prompts with model set to claude-haiku-5-5; do not reuse Haiku 4.5 counts.[6]
  2. Fix the five breaking changes in a branch and handle refusal.
  3. Build a labelled eval of 100 to 500 real requests and compare Haiku 5.5 at two effort levels against your current model and against Sonnet 5.5.
  4. Compute cost per correct answer, not cost per call, including thinking tokens and any prompts that cross 100k.
  5. Route only the tasks that pass; keep the stronger model as a fallback for low-confidence cases.

What remains uncertain

  • Independent results: the benchmark figures are Anthropic's. We have not seen third-party evaluations of Haiku 5.5, and we ran none.
  • Real token inflation: the 30% tokenizer figure is approximate and depends on content; measure it.
  • Thinking cost: Anthropic does not publish a typical thinking-token overhead per effort level in the pages we read, so Examples A to C understate cost to an unknown degree.
  • Competitor claims: GPT-6 Luna figures come from Anthropic's table; verify against OpenAI's documentation before relying on them.
  • Pages move: the announcement and docs were all published within the last day and may be revised.

If a figure above has changed on Anthropic's pages, trust the page and tell us so we can add a dated correction.

Frequently asked questions

How much does Claude Haiku 5.5 cost?
Anthropic lists $0.10 per million input tokens and $0.50 per million output tokens for prompts up to 100,000 tokens, and $0.50 and $2.50 for prompts over 100,000 tokens. Batch requests are 50% off and cache reads are $0.01 per million tokens in the lower tier. [3][4]
What is the context window of Claude Haiku 5.5?
1M tokens, with up to 128K output tokens on the synchronous Messages API and up to 300K on the Batch API with a beta header. [4]
Can I just change the model ID from Haiku 4.5 to Haiku 5.5?
Not safely. Anthropic's migration guide lists breaking changes: manual extended thinking, non-default sampling parameters and assistant prefill now return errors, and thinking blocks are invalidated if earlier turns change. [5][6]
Is Haiku 5.5 better than Sonnet 5.5?
Not on Anthropic's own published benchmark table, where Sonnet 5.5 scores higher in every row shown. Haiku 5.5 is positioned for narrowly scoped, high-volume and latency-sensitive work at a much lower price. [1]

The ANTM newsletter

The signal, not the noise.

Sourced AI coverage in your inbox. Double opt-in, unsubscribe in one click.

Referenced sources

  1. 1.
    Introducing Claude Haiku 5.5(opens in a new tab)

    AnthropicCompanyPublished Oct 7, 2026Accessed Oct 3, 2026

  2. 2.
    Models overview(opens in a new tab)

    AnthropicPrimary sourcePublished Oct 8, 2026Accessed Oct 3, 2026

  3. 3.
    Pricing(opens in a new tab)

    AnthropicPrimary sourcePublished Oct 8, 2026Accessed Oct 3, 2026

  4. 4.
    Claude Haiku 5.5 model overview(opens in a new tab)

    AnthropicPrimary sourcePublished Oct 7, 2026Accessed Oct 3, 2026

  5. 5.
    What's new in Claude Haiku 5.5(opens in a new tab)

    AnthropicPrimary sourcePublished Oct 7, 2026Accessed Oct 3, 2026

  6. 6.
    Claude Haiku 5.5 migration guide(opens in a new tab)

    AnthropicPrimary sourcePublished Oct 7, 2026Accessed Oct 3, 2026

  7. 7.
    Prompting Claude Haiku 5.5(opens in a new tab)

    AnthropicPrimary sourcePublished Oct 7, 2026Accessed Oct 3, 2026

ANTM Editorial

Editorial desk

The editorial desk at AI's Next Top Model. Every article is sourced to primary documents and approved by an editor before publication. See the editorial policy for how we work.