A token is the unit a language model reads, writes and bills in, and almost every limit and price in an AI API is expressed in them. The practical questions are narrower than "what is a token": what fills a context window, which tokens you pay for, and why the same task can cost very different amounts depending on how it is run. This guide answers those from the vendors' own documentation, then works through one arithmetic example that explains why agent workloads are expensive.
Everything below was read from Anthropic's and OpenAI's documentation on 2026-10-05. Most of the detail is Anthropic's because its pages document billing mechanics in the most depth; where a rule is vendor-specific we say so. Prices and limits change, so check the linked pages before budgeting.
Quick answer
- A token is a chunk of text, not a word. OpenAI's rule of thumb is about 4 characters or 0.75 words of English per token; Anthropic's pricing page gives the same figures.[3][5]
- The context window holds everything in a request and the answer: system prompt, messages, tool results, images, documents, tool definitions and the model's output, including thinking.[2]
- You pay separately for input and output, and output is priced at 5x input on every tier in Anthropic's current price table.[3]
- In agents and long chats, the history is sent again each turn. Caching can cut that cost sharply, but only for prompts above a minimum length and only while the prefix is unchanged.[2][4]
- Never reuse token counts across model generations. Count with the model you will actually call.[1]
What a token is
OpenAI's documentation describes tokens as "chunks" of text that "represent commonly occurring sequences of characters." Its example is that " tokenization" splits into " token" and "ization", while a short common word such as " the" is a single token.[5] The consequence is that token counts depend on the text. Rare words, code, other languages and unusual formatting usually cost more tokens per character than plain English prose, which is why the vendors call their 4-characters figure a rough rule only.[3][5]
The tokenizer is also part of the model. Anthropic says Claude 4.7 and later models use a newer tokenizer that produces approximately 30 percent more tokens for the same text, depending on content and workload, and that Fable 5.1 shares it.[1] The same prompt can therefore be a bigger bill on a newer model even if the per-token price is unchanged. Anthropic's advice is to recount prompts against the model you plan to use.[1]
What fills the context window
Anthropic defines the context window as all the text a model can reference when generating a response, "including the response itself," and describes it as working memory, distinct from training data.[2] Its list of what counts is specific: the system prompt, every message (including tool results, images and documents), your tool definitions, and the output for the turn including extended thinking.[2] OpenAI states the same constraint for text models: prompt and generated output combined must not exceed the model's maximum context length.[5]
Some details that surprise people:
- Thinking counts twice in different ways. Anthropic says thinking tokens are a subset of
max_tokens, are billed as output tokens, and count toward rate limits. Whether thinking blocks from earlier turns stay in context depends on the model; on Opus 4.5 and later Opus models, Sonnet 4.6 and later Sonnet models and Fable 5.1, the API keeps them by default, so they are billed again as input on later turns.[2] - Tool definitions are not free. When you supply tools the API adds a tool-use system prompt: 286 tokens for Opus 5.5 and Sonnet 5.5 at the default tool choice, on top of the tokens in your tool names, descriptions and schemas.[3] In Anthropic's token-counting example, a request with one small weather tool counted 403 input tokens.[1]
- Images and documents are tokens too. Anthropic's examples count an image request at 1,028 input tokens and a PDF request at 2,188; these are the documentation's sample inputs, not typical sizes.[1] Its pricing page estimates a 10 kB web page at about 2,500 tokens and a 500 kB research PDF at about 125,000 when fetched into context.[3]
- Overflow behaves differently by case. If the input alone exceeds the window, the API returns a 400 error. On Claude 4.5 models and newer, a request whose input plus
max_tokensexceeds the window is accepted, and generation stops withstop_reason: "model_context_window_exceeded"if it hits the limit.[2]
A bigger window is not a free upgrade. Anthropic's own page says "more context isn't automatically better": as token count grows, accuracy and recall degrade, which it calls context rot.[2] See Long Context Windows: Capacity Is Not the Same as Reliability for the research on that.
What you are billed for
| Item | Counts toward context window | How it is billed (Anthropic) |
|---|---|---|
| New input tokens | Yes | Base input price |
| Output tokens, including thinking | Yes | Output price (5x input on current tiers) |
| Cache write, 5-minute | Yes | 1.25x base input |
| Cache write, 1-hour | Yes | 2x base input |
| Cache read | Yes | 0.1x base input (0.05x Opus 5.5, 0.025x Fable 5.1) |
| Batch API requests | Yes | 50% off input and output |
| Server tools such as web search | Results count as input | Extra usage fee (web search: $10 per 1,000 searches) |
Sources: Anthropic's context-window and pricing pages, read 2026-10-05.[2][3] The multipliers stack with each other, and long-context requests on the 1M-token models are billed at the same per-token rate as short ones.[3]
Why agent loops cost more than they look
A chat or agent turn is stateless from the API's point of view: the conversation so far is sent again with each request. Anthropic's context-window page describes this directly, with each turn's input containing all previous history plus the new message.[2] So cost grows with the sum of every prefix, not just the final length.
The following is our arithmetic on Anthropic's list prices for Sonnet 5.5 ($2 per million input tokens, $10 per million output tokens, $2.50 per million for a 5-minute cache write, $0.20 per million for a cache read).[3] It is a model of a workload, not a measurement.
Assumptions: a 10-turn tool-using loop. A fixed 20,000-token prefix (system prompt plus tool definitions). Each turn adds 2,500 tokens to the history (2,000 of new tool results and user text, plus 500 tokens of the previous answer). Each turn produces 500 output tokens. A cache breakpoint is set at the end of the history each turn, and no cache entry expires between turns.
| Quantity | No caching | With caching |
|---|---|---|
| Total input tokens processed over 10 turns | 312,500 | 312,500 |
| of which cache writes | 0 | 42,500 |
| of which cache reads | 0 | 270,000 |
| Input cost | $0.625 | about $0.160 |
| Output cost (5,000 tokens) | $0.050 | $0.050 |
| Total | $0.675 | about $0.210 |
The input arithmetic is 20,000 + 2,500 x (n - 1) tokens on turn n, summed over 10 turns. Without caching that is 312,500 x $2 / 1M = $0.625. With caching, the 42,500 written tokens cost 42,500 x $2.50 / 1M = $0.106 and the 270,000 read tokens cost 270,000 x $0.20 / 1M = $0.054.
Two observations, labelled as our interpretation. First, without caching, input is about 93% of the total ($0.625 of $0.675), so the usual instinct to shorten answers addresses the smaller part of the bill. Second, caching cuts this example by roughly two thirds, but only because the 20,000-token prefix stayed identical for every turn. If something earlier in the prompt changes, the entries after it are invalidated and you pay write prices again.
The rules that decide whether caching works
Anthropic's caching page sets out conditions that are easy to miss:[4]
- Minimum length. The minimum cacheable prompt is 512 tokens for Sonnet 5.5, Opus 5.5 and Fable 5.1, 1,024 for Sonnet 5 and Sonnet 4.6, and 4,096 for Haiku 4.5. Shorter prompts are processed without caching and without an error, so a missing discount does not announce itself.
- Order matters. Cache prefixes are built in the order tools, then system, then messages. A change at one level invalidates that level and everything after it; changing tool definitions invalidates all three.
- Lifetime. The default is 5 minutes, refreshed free each time the entry is read. A 1-hour option costs 2x base input for the write. The lifetime is measured from the start of the request, so a four-minute streamed response leaves about a minute for the next request to reuse it.
- Break-even. Per the pricing page, a 5-minute write pays off after one cache read and a 1-hour write after two.[3]
Caching also does nothing for the context window: cached prefixes still occupy it.[2]
Cost levers, in the order we would try them
- Measure first. Anthropic's token counting endpoint accepts the same inputs as the Messages API, including system prompts, tools, images and PDFs, is free, and returns an estimate that can differ slightly from the billed count.[1] It rejects some inputs the Messages API accepts, such as server tools and URL-sourced images, and does not apply caching.[1]
- Stabilise the prefix. Put static content (tools, system prompt, reference documents) first and variable content last, so caches survive.[4]
- Trim what you re-send. Anthropic documents server-side compaction and context editing, including clearing old tool results, for long conversations.[2]
- Batch what can wait. The Batch API is 50% off input and output for asynchronous work.[3]
- Pick the tier by cost per passed task, not per call. See Claude Opus 5.5 vs Sonnet 5.5 vs Fable 5.1 for a routing approach.
A minimal count, following Anthropic's documented Python call:
import anthropic
client = anthropic.Anthropic() # reads ANTHROPIC_API_KEY from the environment
response = client.messages.count_tokens(
model="claude-sonnet-5-5",
system="You are a scientist",
messages=[{"role": "user", "content": "Hello, Claude"}],
)
print(response.json())Run it with the model you plan to deploy, then again with the model you are migrating from; the difference is your tokenizer effect.[1]
What remains uncertain
- The worked example is a model with stated assumptions. Real loops vary in turn count, tool-result size and cache hit rate, and we did not run this workload.
- Most mechanics here are Anthropic-specific. OpenAI's concepts page, which we read, covers what a token is and the context-length constraint but not caching or billing multipliers, so we make no claims about OpenAI's pricing rules. Check each vendor's own pricing page.
- Vendor documents change. Minimum cache lengths, multipliers and tokenizers differ by model and have shifted between generations.




