Claude Haiku 5.5 is Anthropic's smallest current model, released on 2026-10-07 with the API ID claude-haiku-5-5. Anthropic positions it for high-volume, latency-sensitive work such as classification, extraction, routing and subagent tasks.[4] This Model Watch reads Anthropic's announcement and documentation to answer the questions that decide whether to adopt it: what it costs, what changed in the API, what the benchmark claims do and do not show, and which jobs should go to Haiku rather than Sonnet or Opus.
All facts below come from Anthropic's pages, read on 2026-10-08. We did not run the model, so nothing here is our own testing; where we calculate or interpret, we say so. Prices and limits change, so confirm them on the linked pages before budgeting.
Quick answer
- What it is: a small model in the 5.5 family with a 1M-token context window, 128K max output, adaptive thinking and, for the first time in a Haiku, an
effortsetting.[1][4] - What it costs: $0.10 input and $0.50 output per million tokens for prompts up to 100,000 tokens; $0.50 and $2.50 above that.[3]
- The catch on cost: the same text counts as roughly 30% more tokens than on Haiku 4.5, so the saving is real but smaller than the per-token drop suggests.[5]
- The catch on migration: five breaking API changes, plus a new refusal stop reason with no server-side fallback.[5][6]
- Where it fits: narrowly scoped, high-volume work. Anthropic itself says Sonnet 5.5 and Opus 5.5 remain better for complex agentic coding.[1]
Specifications
| Spec | Haiku 5.5 | Sonnet 5.5 | Opus 5.5 | Fable 5.1 |
|---|---|---|---|---|
| API model ID | claude-haiku-5-5 | claude-sonnet-5-5 | claude-opus-5-5 | claude-fable-5-1 |
| Positioning (Anthropic) | High-volume, latency-sensitive | Speed and intelligence | Long-running agentic coding, knowledge work | Demanding reasoning, long-horizon agents |
| Comparative latency | Fastest | Fast | Moderate | Slower |
| Context window | 1M tokens | 1M tokens | 1M tokens | 1M tokens |
| Max output (sync) | 128K tokens | 128K tokens | 128K tokens | 128K tokens |
| Thinking | Adaptive | Adaptive | Adaptive (always on) | Adaptive (always on) |
| Default effort | medium | high | medium | high |
| Input / output per MTok | $0.10 / $0.50 (to 100k prompt) | $2 / $10 | $4 / $20 | $10 / $50 |
| Retirement commitment | Not sooner than 2027-10-07 | Not sooner than 2027-09-28 | Not sooner than 2027-09-22 | Not sooner than 2027-09-01 |
Source: Anthropic models overview and Haiku 5.5 model page, read 2026-10-08.[2][4] Haiku 5.5 accepts text and images and returns text, with a reliable knowledge cutoff of June 2026.[4] The retirement column is a commitment for Anthropic-operated platforms; Amazon Bedrock and Google Cloud set their own dates.[2]
Pricing, including the 100,000-token break
Haiku 5.5 is the one current Claude model with length-tiered pricing. Anthropic's pricing page says Claude 4.6 and later models include the full 1M window at standard pricing, "except Claude Haiku 5.5", which pays higher prices when the prompt exceeds 100,000 tokens.[3]
| Price per million tokens | Prompt up to 100k | Prompt over 100k |
|---|---|---|
| Input | $0.10 | $0.50 |
| Output | $0.50 | $2.50 |
| 5-minute cache write | $0.125 | $0.625 |
| 1-hour cache write | $0.20 | $1.00 |
| Cache read | $0.01 | $0.05 |
| Batch input / output | $0.05 / $0.25 | $0.25 / $1.25 |
Source: Anthropic pricing page, read 2026-10-08.[3] Anthropic's announcement adds that about 90% of Haiku 4.5 requests fell in the up-to-100k range, and describes Haiku 5.5 as about 90% cheaper per request than Haiku 4.5 up to 100k and 50% cheaper above it, with the tokenizer change accounted for.[1] That is Anthropic's figure for its own traffic mix, not a guarantee for yours.
Worked examples (our arithmetic)
These use the published list prices. The only assumption we add is that Haiku 5.5 counts 30% more tokens than Haiku 4.5 for the same text, which is Anthropic's "approximately" figure and varies by content.[5] We ignore thinking tokens, which are billed as output and which you would need to measure.
Example A: a classification call. A prompt of 2,000 tokens (as counted on Haiku 4.5) with a 200-token answer. On the new tokenizer that is about 2,600 in and 260 out.
| Model | Cost per request | Cost per million requests |
|---|---|---|
| Haiku 4.5 | $0.0030 | $3,000 |
| Haiku 5.5 | $0.00039 | $390 |
| Sonnet 5.5 | $0.0078 | $7,800 |
| Opus 5.5 | $0.0156 | $15,600 |
Example B: one long document. A 150,000-token prompt with a 1,000-token answer crosses the price break. Haiku 5.5 costs about $0.0775 per call ($0.075 input at $0.50, $0.0025 output at $2.50). Sonnet 5.5 costs $0.31 ($0.30 plus $0.01). Haiku is still four times cheaper here, but the gap is far narrower than in Example A, so a long-context workload that quietly crosses 100k deserves its own estimate.
Example C: a cached prefix. A 30,000-token cached system prompt plus 500 new input tokens and 300 output tokens. At cache-read prices that is about $0.0005 on Haiku 5.5 and $0.007 on Sonnet 5.5, because Sonnet 5.5 reads cache at $0.10 and Haiku 5.5 at $0.01 per million.[3] Without caching, the same Haiku 5.5 call would cost about $0.0032.
The practical reading: the price gap is largest on short, repetitive, cache-friendly requests, which is exactly the workload Haiku is aimed at. See Tokens, Context Windows and What an AI Request Really Costs for how caching multipliers and re-sent history drive the bill.
What the benchmark table does and does not show
Anthropic's announcement includes a table comparing Haiku 5.5 with Haiku 4.5, OpenAI's GPT-6 Luna and Sonnet 5.5.[1] A selection of rows:
| Benchmark (as reported by Anthropic) | Haiku 5.5 | Haiku 4.5 | GPT-6 Luna | Sonnet 5.5 |
|---|---|---|---|---|
| GDPval-AA v2.1 (Elo) | 1620 | 735 | 1437 | 1840 |
| OSWorld 2.1 (offline subset) | 72.4% | 15.7% | 48.9% | 83.9% |
| Humanity's Last Exam, no tools | 45.9% | 10.2% | not shown | 56.9% |
| Terminal-Bench 4.0 | 39.2% | 0.0% | 16.4% | 70.6% |
| FrontierCode 1.1 (Main) | 46.4% | not shown | 42.4% | 52.1% |
Source: Anthropic announcement, read 2026-10-08.[1] Our interpretation, based on the published numbers:
- Generational gap: Haiku 5.5 is far ahead of Haiku 4.5 on every row shown, which is consistent with the price drop being a new model class rather than a discount.
- Tier ordering is preserved: Sonnet 5.5 beats Haiku 5.5 on every row shown, and among the percentage-scored rows the gap is widest on Terminal-Bench 4.0 (70.6% against 39.2%). That matches Anthropic's own guidance that Sonnet and Opus remain better for complex agentic coding.[1]
- Comparisons with GPT-6 Luna are the vendor's: we did not open OpenAI's pages in this run, so we cannot confirm the competitor figures or whether settings were matched. Anthropic says the full methodology is in the Haiku 5.5 system card, which the announcement links but does not reproduce, and we did not review it. The Sonnet 5.5 FrontierCode figure is marked as run at xhigh effort; we could not confirm the effort setting for the other cells.
- Customer results are self-reported: the announcement quotes teams such as Box, AlphaSense, HubSpot and Asana reporting latency and accuracy gains on their own evals.[1] They show it can work on real tasks, not that it will on yours.
For how to read claims like these, see How to Read an AI Benchmark Claim Without Being Misled.
What breaks when you migrate from Haiku 4.5
Anthropic's migration guide is clear that this is more than a new model string.[6] The breaking changes:
- Manual thinking budgets are rejected.
thinking: {"type": "enabled", "budget_tokens": N}returns a 400 error. Use{"type": "adaptive"}and setoutput_config.effort.[6] - Sampling parameters are rejected. Omit
temperature,top_pandtop_k; non-default values return a 400.[6] - Assistant prefill is rejected, even with thinking off. End
messageswith a user turn and use structured outputs or tools for format control.[6] - Computer use moves to a toolset (
computer_toolset_20260801) on the Claude API and Google Cloud; the oldcomputer_20250124tool errors.[6] - Thinking blocks are bound to earlier turns. Changing
system,toolsor earlier messages and then sending a thinking block back returns a 400, so conversations must be append-only.[6]
Behaviour changes also matter for cost and correctness. Thinking is on by default and counts toward max_tokens, so a limit tuned for Haiku 4.5 can stop after a thinking block and before any text. Responses can begin with thinking blocks, so code must select blocks by type, and thinking text is omitted unless you set display to "summarized".[5][6] Priority Tier is not supported on Haiku 5.5.[6]
Two operational notes from the prompting guide. Haiku 5.5 runs safety classifiers that can return stop_reason: "refusal", and there is no server-side fallback, so clients must handle it; resending the same request usually refuses again.[7] And changing the top-level effort between requests invalidates the prompt cache for the conversation's messages, which interacts badly with the caching savings above unless you use the per-message effort beta.[7]
Effort: the lever that replaces the thinking budget
Haiku 5.5 is the first Haiku with effort levels. Anthropic's guidance:[7]
low: cheapest and fastest, for chat, short tool tasks and simple high-volume requests. In long agent prompts the model is more likely to skip a search, stop early or skip a check.medium: the API default and the suggested starting point, including for agentic coding.high: for knowledge work, longer agent tasks and strict instruction following.xhighandmax: only where an eval gain justifies the cost; Anthropic suggests also testing Sonnet 5.5 at those levels.
The practical trade-off is that effort moves cost as well as quality. In Anthropic's own testing, moving from low to medium roughly halved early stopping in a long coding-agent prompt while more than doubling output tokens per attempt.[7] A "cheap" model run at high effort on long outputs can erode the price advantage, which is why a cost test should use your real prompts at the effort you will ship.
Which jobs go where
This is our editorial recommendation, derived from the documented specs, prices and Anthropic's own positioning, not from independent testing.
| Job | Start with | Why |
|---|---|---|
| Classification, routing, tagging, short extraction | Haiku 5.5 at low or medium | Anthropic's stated target; lowest price; benchmark-test it against your labels |
| Summarisation and context compaction in agents | Haiku 5.5 | Named by Anthropic as a fit; cache reads at $0.01 per million help |
| Subagent that searches or reads for a lead model | Haiku 5.5, lead on Sonnet 5.5 or Opus 5.5 | Anthropic describes this pattern; keep the verification instructions from the prompting guide |
| Prompts regularly above 100k tokens | Compare Haiku 5.5 and Sonnet 5.5 on cost per correct answer | The price break narrows the gap |
| Complex multi-step coding or terminal work | Sonnet 5.5 or Opus 5.5 | Anthropic says they remain better; Terminal-Bench gap is the widest of the percentage rows |
| Customer-facing chatbot with strict rules | Test Haiku 5.5 at high with the system-prompt reinforcement Anthropic suggests | The guide says instruction-following improves at high[7] |
| Tasks needing assistant prefill or custom sampling | Plan a prompt rewrite first | These are rejected by the API[6] |
For the larger tier comparison, including Fable 5.1 and cache multipliers, see Claude Opus 5.5 vs Sonnet 5.5 vs Fable 5.1; note that article was written before Haiku 5.5 and lists Haiku 4.5 as the small tier. For cross-vendor prices, see GPT-6.1 Sol vs Claude Sonnet 5.5 vs Gemini: API Specs and Pricing Compared, which does not include Haiku 5.5 or other vendors' small models.
A sensible adoption path
- Count your real prompts with
modelset toclaude-haiku-5-5; do not reuse Haiku 4.5 counts.[6] - Fix the five breaking changes in a branch and handle
refusal. - Build a labelled eval of 100 to 500 real requests and compare Haiku 5.5 at two effort levels against your current model and against Sonnet 5.5.
- Compute cost per correct answer, not cost per call, including thinking tokens and any prompts that cross 100k.
- Route only the tasks that pass; keep the stronger model as a fallback for low-confidence cases.
What remains uncertain
- Independent results: the benchmark figures are Anthropic's. We have not seen third-party evaluations of Haiku 5.5, and we ran none.
- Real token inflation: the 30% tokenizer figure is approximate and depends on content; measure it.
- Thinking cost: Anthropic does not publish a typical thinking-token overhead per effort level in the pages we read, so Examples A to C understate cost to an unknown degree.
- Competitor claims: GPT-6 Luna figures come from Anthropic's table; verify against OpenAI's documentation before relying on them.
- Pages move: the announcement and docs were all published within the last day and may be revised.
If a figure above has changed on Anthropic's pages, trust the page and tell us so we can add a dated correction.




