ANTM
ModelsGuide

Claude Opus 5.5 vs Sonnet 5.5 vs Fable 5.1: Which Tier for Which Job?

Anthropic now sells four current tiers at four price points. What the official pages actually document, what the benchmark tables do and do not let you compare, and how to choose.

Editorial desk

Published 7 min read

Four translucent teal glass columns of increasing height standing on a dark grid, with faint coral specks inside
Illustration generated with AI (FLUX.1 [schnell] (Black Forest Labs) via Cloudflare Workers AI, Apache 2.0). Prompt and direction by ANTM.

Anthropic currently lists four generally available Claude tiers, and the gap between the cheapest and the most expensive is a factor of ten on list price. If you are choosing between Claude Opus 5.5 and Sonnet 5.5, or wondering whether Fable 5.1 is worth its premium, the official pages answer part of the question: specifications, prices and Anthropic's own benchmark tables. They do not answer the part that matters most, which is how each tier behaves on your workload. This guide separates what is documented from what is interpretation, and ends with a routing rule you can test.

Everything below was read from Anthropic's own pages on 2026-10-03. Prices and specs change, so check the linked pages before committing a budget.

Quick answer

  • Start with Sonnet 5.5 if cost matters. It lists at $2 / $10 per million input / output tokens and Anthropic describes it as "the best combination of speed and intelligence."[4] Note that Anthropic's own models overview says to "start with Claude Opus 5.5 for most workloads."[4] Our Sonnet-first default is a cost-driven view, not Anthropic's guidance, so test both.
  • Move to Opus 5.5 where your own evaluation shows Sonnet falling short, particularly on long-running agentic coding and knowledge work.[4]
  • Reserve Fable 5.1 for "demanding reasoning and long-horizon agentic work," or when evals on Opus 5.5 at higher effort still fall short.[4]
  • Use Haiku 4.5 for high-volume simple tasks, but note its retirement commitment is the earliest of the four (see the table).[4]

The documented specifications

The table comes from Anthropic's models overview and pricing page, both read on 2026-10-03.[4][5]

SpecFable 5.1Opus 5.5Sonnet 5.5Haiku 4.5
Anthropic's descriptionDemanding reasoning and long-horizon agentic workLong-running agentic coding and knowledge workBest combination of speed and intelligenceFastest, near-frontier intelligence
Input / output per million tokens$10 / $50$4 / $20$2 / $10$1 / $5
Cache read per million input tokens$0.25$0.20$0.20$0.10
Batch API input / output per million$5 / $25$2 / $10$1 / $5$0.50 / $2.50
Context window1M tokens1M tokens1M tokens200K tokens
Max output (synchronous)128K tokens128K tokens128K tokens64K tokens
Comparative latencySlowerModerateFastFastest
Reliable knowledge cutoffJun 2026Jun 2026Jun 2026Feb 2025
API model IDclaude-fable-5-1claude-opus-5-5claude-sonnet-5-5claude-haiku-4-5-20251001
Retirement (Anthropic-operated platforms)Not sooner than Sep 1, 2027Not sooner than Sep 22, 2027Not sooner than Sep 28, 2027Not sooner than Oct 15, 2026

Three details are easy to miss.

Caching is priced differently per tier. The standard cache-read multiplier is 0.1x the input price, but Anthropic's pricing page lists 0.05x for Opus 5.5 and 0.025x for Fable 5.1.[5] That is why Opus 5.5 and Sonnet 5.5 share the same $0.20 cache-read price despite a 2x difference in input price. For an agent that re-reads a large, stable prompt on every turn, the cached rate matters more than the headline input price.

Long context is not a surcharge tier. The pricing page says Claude 4.6 and later models include the full 1M-token window at standard pricing, so a 900k-token request is billed at the same per-token rate as a 9k-token request.[5] Capacity and reliability are different questions; see Long Context Windows: Capacity Is Not the Same as Reliability.

Modifiers stack. Anthropic lists a 50% Batch API discount, a 1.1x multiplier for US-only inference (inference_geo: "us") on Claude 4.6 and later, and fast mode for Opus 5.5 at $8 / $40 per million tokens. The pricing page says caching multipliers stack with the others, and that fast mode is not available with the Batch API.[5]

What the benchmark tables do and do not tell you

Anthropic's Sonnet 5.5 launch page includes a table with Sonnet 5.5, Sonnet 5, Opus 5.5 and a competitor column. These are vendor-reported results; some rows carry footnotes on the page that qualify the setup, so read the page before quoting them.[1] A selection of the Sonnet 5.5 and Opus 5.5 columns:

Benchmark (as labelled by Anthropic)Sonnet 5.5Opus 5.5
Terminal-Bench 4.0 (agentic coding)70.6%66.4% (footnoted)
FrontierCode 1.1 Main (agentic coding)46.2% (Max effort)54.4%
CursorBench 4.0 (agentic coding)55.5%57.8%
GDPval-AA v2.1 (knowledge work, Elo)18441846
AA-Briefcase v1.1 (knowledge work, Elo)18111822
Humanity's Last Exam (with tools)64.5%67.7%
OSWorld 2.1 (computer use, partial)80.1%81.8%
Chartography (visual chart recognition, no tools)61.6%64.4%

Source: Anthropic's Sonnet 5.5 page, read 2026-10-03.[1] The Opus 5.5 page reports the same Opus figures for the rows it covers.[2]

Our interpretation, which is analysis and not a measured result:

  1. The two tiers are close on several rows and separated on others. The knowledge-work Elo scores differ by 2 and 11 points; the larger gaps are on FrontierCode (8.2 points) and Humanity's Last Exam (3.2 points). If your work looks like the rows where they are close, the half-price tier is a strong default. If it looks like FrontierCode, test Opus.
  2. Do not read a single row as a ranking. Sonnet 5.5 is shown above Opus 5.5 on Terminal-Bench 4.0 while Opus leads on the other agentic coding rows, and the Opus figure carries a footnote. A benchmark is a snapshot of a particular harness; see How to Read an AI Benchmark Claim Without Being Misled.
  3. Fable 5.1 cannot be placed on this table. Its launch page reports different benchmark versions (for example CursorBench 3.2.0 and OSWorld 2.0 strict) against Fable 5 and Opus 5.[3] Numbers from different versions are not comparable, so we cannot say from published material how Fable 5.1 compares with Opus 5.5 or Sonnet 5.5 on the same task. Anthropic's own guidance is to move to Fable when evals on Opus 5.5 at higher effort still fall short.[4]

On cost per task, Anthropic says Sonnet 5.5 is priced the same as Sonnet 5 "but it typically needs far fewer tokens to do the same work," and that "in our testing, it costs up to 30% less per task than its predecessor."[1] That is a vendor claim about its own testing against the previous version, not a comparison with Opus.

A worked cost example

The following is arithmetic on the list prices above, not a measurement. Assume a task that sends 50,000 input tokens and produces 10,000 output tokens, with no caching and no batch discount. We assume equal token counts across models, which will not hold in practice because Anthropic says token usage differs by model.[1]

TierInput costOutput costTotal per taskTotal per 1,000 tasks
Fable 5.1$0.50$0.50$1.00$1,000
Opus 5.5$0.20$0.20$0.40$400
Sonnet 5.5$0.10$0.10$0.20$200
Haiku 4.5$0.05$0.05$0.10$100

Now suppose 40,000 of the 50,000 input tokens are cache reads on Opus 5.5. Those cost 40,000 x $0.20 / 1,000,000 = $0.008 instead of $0.16, so the task falls from $0.40 to about $0.248 before the first request's cache write. Running the same task through the Batch API at 50% would halve input and output again. The lesson is that prompt structure and routing often move cost more than the choice between adjacent tiers.

Note also that Anthropic says Claude 4.7 and later models use a newer tokenizer that produces approximately 30% more tokens for the same text compared with earlier models.[5] When comparing against a prior-generation model, count tokens on both sides rather than comparing per-token prices.

Choosing by workload

The mapping below is our recommendation, built from the documented descriptions and prices. It is a starting hypothesis to confirm with your own tests.

  • High-volume classification, extraction, routing: Haiku 4.5 or Sonnet 5.5, whichever passes your accuracy bar at the lower cost. Plan for Haiku 4.5's earliest retirement date.[4]
  • Everyday coding assistance and general business work: Sonnet 5.5 first if you are cost-sensitive. It is the fast tier and the knowledge-work scores are close to Opus on Anthropic's table.[1] Anthropic's own default recommendation is Opus 5.5 for most workloads,[4] so weigh the 2x price difference against your own results.
  • Long-running agentic coding, large refactors, multi-hour agent runs: Opus 5.5, which Anthropic positions for long-running agentic coding and knowledge work, and which leads on FrontierCode and CursorBench in the table above.[2][4]
  • The hardest reasoning and research-style tasks: try Opus 5.5 at higher effort, then Fable 5.1 if it still falls short, following Anthropic's guidance.[4]
  • Overnight, non-interactive jobs: use the Batch API at half price on whichever tier passes your evals.[5]
  • Latency-sensitive interactive features: Anthropic lists Sonnet 5.5 as "Fast" and Opus 5.5 as "Moderate"; Opus 5.5 also has a research-preview fast mode at $8 / $40.[4][5]

For agent workloads, tier choice interacts with safety design; see How AI Agents Work, and How to Secure Them.

How to test it yourself

  1. Collect 30 to 50 real tasks with a pass/fail rubric you wrote before looking at outputs.
  2. Run each tier with the same prompt and the same effort setting. Anthropic's default effort differs by model (high for Fable 5.1 and Sonnet 5.5, medium for Opus 5.5), so set it explicitly or you are comparing settings as well as models.[4]
  3. Record pass rate, total input and output tokens, and wall-clock time.
  4. Compute cost per passed task, not cost per call.
  5. Pick the cheapest tier that clears your bar, and route only the failures to the next tier up.

What remains uncertain

  • Footnotes on the benchmark table qualify some figures, and we have not reproduced any result; every score here is vendor-reported.
  • There is no published like-for-like comparison of Fable 5.1 against Opus 5.5 or Sonnet 5.5 on the same benchmark versions in the pages we read.
  • This is a single-vendor comparison. A cross-vendor guide needs the same level of primary-source reading for each lab and will be handled separately.
  • Prices, retirement dates and model lineups change; this article reflects pages read on 2026-10-03.

Frequently asked questions

Which Claude model is cheapest for production?
By list price, Haiku 4.5 at $1 per million input tokens and $5 per million output tokens, then Sonnet 5.5 at $2 and $10. Cost per task also depends on how many tokens a model needs, which Anthropic says is lower for Sonnet 5.5 than for Sonnet 5. [2][4]
Is Sonnet 5.5 as good as Opus 5.5?
Not on every row of Anthropic's own table. The two are close on GDPval-AA and AA-Briefcase and OSWorld, and Opus 5.5 is ahead on FrontierCode, CursorBench, Humanity's Last Exam and Chartography. Whether the gap matters is a question for your own evaluation. [2]
Do the models have the same context window?
Fable 5.1, Opus 5.5 and Sonnet 5.5 list 1M tokens and Haiku 4.5 lists 200K. Anthropic's pricing page says Claude 4.6 and later models include the full 1M window at standard pricing. [4][5]
Can I use the Batch API to cut costs?
Yes for asynchronous work. Anthropic lists a 50% discount on input and output tokens, and fast mode is not available with the Batch API. [5]

The ANTM newsletter

The signal, not the noise.

Sourced AI coverage in your inbox. Double opt-in, unsubscribe in one click.

Referenced sources

  1. 1.
    Introducing Claude Sonnet 5.5(opens in a new tab)

    AnthropicCompanyPublished Sep 28, 2026Accessed Oct 3, 2026

  2. 2.
    Introducing Claude Opus 5.5(opens in a new tab)

    AnthropicCompanyPublished Sep 22, 2026Accessed Oct 3, 2026

  3. 3.
  4. 4.
    Models overview(opens in a new tab)

    AnthropicPrimary sourcePublished Oct 3, 2026Accessed Oct 3, 2026

  5. 5.
    Pricing(opens in a new tab)

    AnthropicPrimary sourcePublished Oct 3, 2026Accessed Oct 3, 2026

ANTM Editorial

Editorial desk

The editorial desk at AI's Next Top Model. Every article is sourced to primary documents and approved by an editor before publication. See the editorial policy for how we work.