ANTM
ModelsGuide

GPT-6.1 Sol vs Claude Sonnet 5.5 vs Gemini: API Specs and Pricing Compared

Three labs now sell a model at $2 per million input tokens. The headline prices match; the long-prompt rules, cache mechanics and output limits do not. A comparison built only from the vendors' own pages.

Editorial desk

Published 7 min read

A perspective grid of fine dotted lines in the lower part of a dark navy frame, crossed by a thin coral horizontal threshold line with a short vertical column of coral and teal dots rising from the grid
Illustration generated with AI (FLUX.1 [schnell] (Black Forest Labs) via Cloudflare Workers AI, Apache 2.0). Prompt and direction by ANTM.

Three labs now sell a capable model at the same headline input price. OpenAI's GPT-6.1 Sol and Anthropic's Claude Sonnet 5.5 each list $2 per million input tokens and $10 per million output tokens, and Google's Gemini 3.1 Pro Preview lists $2 input and $12 output for prompts up to 200k tokens.[1][5][6] If the comparison stopped at that row, the choice would be a coin flip. It does not stop there: surcharges for long prompts, cache write and storage fees, output limits and preview status all differ, and they decide what a real workload costs.

Everything below was read from the vendors' own documentation on 2026-10-04. This is a comparison of documented specifications and list prices only. We did not run benchmarks, and we do not claim any of these models gives better answers than another. Prices and limits change; check the linked pages before budgeting.

Quick answer

  • Under about 200K tokens per request, the list prices of Sol and Sonnet 5.5 are identical, and Gemini 3.1 Pro Preview is 20% higher on output ($12 against $10).[1][5][6] Differences then come from caching, tokens-per-task and quality, not list price.
  • Past 272K input tokens the three diverge. OpenAI bills Sol at 2x input and 1.5x output; Anthropic says Sonnet 5.5 stays at the standard rate through its 1M window; Gemini 3.1 Pro Preview moves to $4 input and $18 output above 200k.[1][5][6]
  • Gemini 3.8 Flash is the low-price option at $0.75 input and $3.75 output per million tokens, but that is an introductory rate through 2026-12-31, after which Google lists $1.50 and $7.50.[6]
  • For very long outputs, Sol and Sonnet 5.5 allow 128K output tokens; both Gemini models list 65,536.[1][4][7][8]

Why these models

GPT-6.1 Sol was released on 2026-09-29 for "complex coding and professional work," according to OpenAI's changelog, after GPT-6 Astra (2026-09-03) and GPT-6 Sol and Luna (2026-09-22).[3] OpenAI's models page positions Sol as "near-Astra performance for complex work at a lower cost," between Astra at $10 / $50 and Luna at $0.10 / $0.50.[2] Claude Sonnet 5.5 is Anthropic's mid tier, described as "the best combination of speed and intelligence."[4] On Google's side, Gemini 3.1 Pro Preview is the Pro-class model on the pricing page, and Gemini 3.8 Flash is described by Google as engineered for long-horizon software engineering and autonomous agents.[9]

These are the mid-priced offerings each lab documents today. They are not strict equivalents; the table shows where they differ.

The documented specifications

SpecGPT-6.1 SolClaude Sonnet 5.5Gemini 3.1 Pro PreviewGemini 3.8 Flash
Input / output per million tokens$2 / $10$2 / $10$2 / $12 (prompts up to 200k); $4 / $18 above$0.75 / $3.75 through 2026-12-31; $1.50 / $7.50 from 2027-01-01
Cached input read, per million$0.10$0.20$0.20 (up to 200k); $0.40 above$0.075 (the page lists no separate 2027 cache-read price)
Cache write$2.50 per million$2.50 (5 min) or $4 (1 hour) per millionStorage fee $4.50 per million per hourStorage fee $0.50 per million per hour
Batch discountNot read on the pages we opened50%50%50%
Context window1,050,000 tokens (max input 922,000)1M tokens1,048,576 tokens1,048,576 tokens
Max output128,000 tokens128K tokens65,536 tokens65,536 tokens
Input modalitiesText, imageText, imageText, image, video, audio, PDFText, image, video, audio, PDF
OutputTextTextTextText
Knowledge cutoff2026-04-30Jun 2026 (reliable)Not stated on the pageNot stated on the page
StatusReleased 2026-09-29Retirement not sooner than 2027-09-28Preview; page updated Feb 2026Listed as stable; page updated Sep 2026
Model IDgpt-6.1-solclaude-sonnet-5-5gemini-3.1-pro-previewgemini-3.8-flash

Sources: OpenAI's Sol model page and changelog,[1][3] Anthropic's models overview and pricing page,[4][5] Google's pricing page and the two Gemini model pages.[6][7][8]

Two reading notes. First, "cache write" is not the same mechanism everywhere: OpenAI lists a per-token write price, Anthropic lists write multipliers of 1.25x (5 minutes) and 2x (1 hour) of the input price, and Google lists a per-hour storage fee for cached tokens instead.[1][5][6] Second, Anthropic's overview says all current Claude models accept text and image input and produce text.[4]

The long-prompt rules are the real difference

All four models advertise a window of about one million tokens. What they charge for using it differs:

  • GPT-6.1 Sol: OpenAI's page states that requests above 272K tokens use 2x input and cache pricing and 1.5x output pricing. The maximum input is listed as 922,000 tokens, below the 1.05M total window, which presumably reserves room for output; OpenAI's page does not say so in those words, so treat that as our reading.[1]
  • Claude Sonnet 5.5: Anthropic's pricing page says Claude 4.6 and later models include the full 1M window at standard pricing, and that a 900k-token request is billed at the same per-token rate as a 9k-token request.[5]
  • Gemini 3.1 Pro Preview: input rises from $2 to $4 and output from $12 to $18 for prompts above 200k tokens.[6]
  • Gemini 3.8 Flash: the pricing summary we read lists a single rate with no long-prompt tier. We read this through a page summary and recommend confirming it on Google's page before relying on it for large prompts.[6]

Capacity is not reliability; a window you can afford to fill may still not retrieve well. See Long Context Windows: Capacity Is Not the Same as Reliability.

Worked cost examples

This is arithmetic on the list prices above, not a measurement. It assumes equal token counts across models, no batch discount and no caching.

Example 1: a typical agent-style request, 100,000 input tokens and 10,000 output tokens.

ModelInput costOutput costTotal
GPT-6.1 Sol$0.20$0.10$0.30
Claude Sonnet 5.5$0.20$0.10$0.30
Gemini 3.1 Pro Preview$0.20$0.12$0.32
Gemini 3.8 Flash (through 2026-12-31)$0.075$0.0375$0.1125
Gemini 3.8 Flash (from 2027-01-01)$0.15$0.075$0.225

Example 2: one long request, 300,000 input tokens and 10,000 output tokens.

ModelInput costOutput costTotal
GPT-6.1 Sol (2x input, 1.5x output above 272K)$1.20$0.15$1.35
Claude Sonnet 5.5$0.60$0.10$0.70
Gemini 3.1 Pro Preview (above 200k)$1.20$0.18$1.38
Gemini 3.8 Flash (through 2026-12-31, single rate as read)$0.225$0.0375$0.2625

Our interpretation: for workloads that routinely cross 200K to 272K tokens, the headline price is a poor guide, and Sonnet 5.5's flat rate roughly halves the cost of the second example relative to Sol and Gemini 3.1 Pro Preview. For workloads that stay short, the two $2 / $10 models cost the same and the decision moves to caching and quality.

Caching changes the ranking again. For an agent that re-sends the same large prefix every turn, the cache read price matters more than the input price. Sol's $0.10 read is half of Sonnet 5.5's $0.20, while Gemini 3.1 Pro Preview's $0.20 read comes with a separate hourly storage fee of $4.50 per million cached tokens.[1][5][6] Whether Sol's lower read price outweighs its $2.50 write price depends on how many times each prefix is reused. For background on why re-sent history dominates agent bills, see Tokens, Context Windows and What an AI Request Really Costs.

Tokenizers differ. The same text becomes different token counts on different vendors' tokenizers, and Anthropic says Claude 4.7 and later models produce about 30% more tokens for the same text than earlier ones.[5] A per-token price comparison across vendors therefore needs a recount of your real prompts on each API.

Choosing by workload

This mapping is our recommendation from the documented numbers. It is a hypothesis to confirm with your own tests.

  • Short-to-medium prompts, price-sensitive: Gemini 3.8 Flash is the cheapest listed option; test whether it clears your accuracy bar, and plan for the 2027-01-01 price doubling.[6]
  • Long documents or whole-repository prompts above 272K tokens: Sonnet 5.5's flat pricing is the most predictable of the three mid-tier options on list price.[5]
  • Heavy prompt reuse with a large stable prefix: compare Sol's $0.10 cache reads against your reuse count; it is the lowest read price among the $2-input models.[1]
  • Video, audio or PDF input as a native modality: the Gemini models list these as inputs; Sol and Sonnet 5.5 list text and image.[1][4][7]
  • Very long single outputs: Sol and Sonnet 5.5 allow 128K output tokens against 65,536 for the Gemini models.[1][4][7]
  • Production stability: Gemini 3.1 Pro Preview carries a preview label, so confirm Google's terms before committing a production dependency.[8][9]

What this comparison cannot tell you

It cannot tell you which model solves your task more reliably, which is usually the larger cost: a cheaper model that needs three attempts is not cheaper. Vendor benchmark tables use different harnesses and versions, so we left them out; see How to Read an AI Benchmark Claim Without Being Misled. For the Anthropic tiers beyond Sonnet, see Claude Opus 5.5 vs Sonnet 5.5 vs Fable 5.1: Which Tier for Which Job?.

A fair test is 30 to 50 of your own tasks, scored against a rubric written before you look at outputs, with total tokens, retries and latency recorded per model.

What remains uncertain

  • We did not open OpenAI's or Google's launch announcements; OpenAI's announcement page returned an access error. The release date comes from OpenAI's changelog.
  • We could not confirm a batch discount for Sol, a knowledge cutoff for either Gemini model, or whether Gemini 3.8 Flash has a long-prompt tier beyond what the pricing summary showed.
  • Google's Pro-class model is still labelled a preview, and its page was last updated in February 2026; a newer Pro-class model may be announced without this page changing.

Prices and limits read on 2026-10-04 and re-checked against the OpenAI, Anthropic and Google pages on 2026-10-07. Updated: 2026-10-07.

Frequently asked questions

Which is cheaper, GPT-6.1 Sol or Claude Sonnet 5.5?
At list prices they are identical for prompts under 272K tokens: $2 per million input and $10 per million output. Above 272K input tokens OpenAI applies 2x input and 1.5x output pricing to GPT-6.1 Sol, while Anthropic's pricing page says Sonnet 5.5 is billed at the standard rate across its full 1M window. Cache reads are $0.10 per million on Sol and $0.20 on Sonnet. [1][5]
Does Gemini 3.8 Flash cost less than GPT-6.1 Sol?
On the list prices Google publishes, yes: $0.75 input and $3.75 output per million tokens through 2026-12-31, rising to $1.50 and $7.50 from 2027-01-01. The pricing and model pages say nothing about benchmark quality, so a lower price does not tell you whether it passes your task. [6]
Is Gemini 3.1 Pro Preview a production model?
Google labels it a preview version, and its model page shows a latest update of February 2026. Preview models can change before becoming stable, so check Google's terms for the stability you need. [8][9]

The ANTM newsletter

The signal, not the noise.

Sourced AI coverage in your inbox. Double opt-in, unsubscribe in one click.

Referenced sources

  1. 1.
    GPT-6.1 Sol model page(opens in a new tab)

    OpenAIPrimary sourcePublished Oct 4, 2026Accessed Oct 3, 2026

  2. 2.
    OpenAI API models(opens in a new tab)

    OpenAIPrimary sourcePublished Oct 4, 2026Accessed Oct 3, 2026

  3. 3.
    OpenAI API changelog(opens in a new tab)

    OpenAIPrimary sourcePublished Oct 4, 2026Accessed Oct 3, 2026

  4. 4.
    Models overview(opens in a new tab)

    AnthropicPrimary sourcePublished Oct 4, 2026Accessed Oct 3, 2026

  5. 5.
    Pricing(opens in a new tab)

    AnthropicPrimary sourcePublished Oct 4, 2026Accessed Oct 3, 2026

  6. 6.
    Gemini API pricing(opens in a new tab)

    GooglePrimary sourcePublished Oct 4, 2026Accessed Oct 3, 2026

  7. 7.
    Gemini 3.8 Flash(opens in a new tab)

    GooglePrimary sourcePublished Oct 4, 2026Accessed Oct 3, 2026

  8. 8.
    Gemini 3.1 Pro Preview(opens in a new tab)

    GooglePrimary sourcePublished Oct 4, 2026Accessed Oct 3, 2026

  9. 9.
    Gemini models(opens in a new tab)

    GooglePrimary sourcePublished Oct 4, 2026Accessed Oct 3, 2026

ANTM Editorial

Editorial desk

The editorial desk at AI's Next Top Model. Every article is sourced to primary documents and approved by an editor before publication. See the editorial policy for how we work.