A context window is the amount of text a model can consider in one request. Windows have grown a great deal, and a tempting conclusion follows: skip the retrieval system and put everything in the prompt. Sometimes that is exactly right. But the size of the window answers only one question, how much you can send. It does not answer how well the model will use what you sent.
What the research found
Two widely cited studies address that second question directly. Both examined models available at the time, so they show what to test for, not what is true of every current model.
Position matters
In "Lost in the Middle," Nelson Liu and colleagues studied how models use long inputs. They reported that "performance is often highest when relevant information occurs at the beginning or end of the input context, and significantly degrades when models must access relevant information in the middle of long contexts." The degradation appeared even in models explicitly built for long contexts. The paper was posted in July 2023 and published in the journal Transactions of the Association for Computational Linguistics.[1]
Claimed size is not usable size
The RULER benchmark, from Cheng-Ping Hsieh and colleagues, tested a different weakness. The common "needle in a haystack" test hides one fact in a long document and asks for it. RULER found that "despite achieving nearly perfect accuracy in the vanilla NIAH test, almost all models exhibit large performance drops as the context length increases." It added tasks such as multi-hop tracing and aggregation, and reported that although the tested models "all claim context sizes of 32K tokens or greater, only half of them can maintain satisfactory performance at the length of 32K."[2]
Why this matters for cost and design
Our interpretation of these findings, plus how providers price usage:
- Accuracy can depend on layout. If your task is sensitive to where facts sit, ordering your prompt carefully, for example putting the question and key instructions near the end, may matter as much as the model choice.
- Every extra token has a price. Providers charge for input tokens, and processing a long prompt takes longer. Prompt caching can reduce the cost when you reuse the same prefix.
- More text can mean more distraction. Extra documents are also extra material that can conflict with or crowd out the relevant passage.
A test you can run in an afternoon
You do not need a research benchmark to find out whether this affects you:
- Take a realistic long input from your own work.
- Insert a known fact at several depths: near the start, the middle and the end.
- Ask a question whose answer depends on that fact, and repeat across depths and several prompts.
- Plot accuracy by depth and by total length. If accuracy changes with position, your task is position-sensitive at that length.
Run it at the length you plan to use, not the maximum the model accepts.
When long context is the right choice, and when retrieval is
| Situation | Better fit | Reason |
|---|---|---|
| One-off analysis of a bounded set of documents | Long context | No index to build, and cross-document reasoning is the goal |
| Summarising or checking consistency across one long document | Long context | The task needs the whole text |
| A large collection that exceeds any window | Retrieval | The material does not fit |
| Data that changes daily | Retrieval | You update documents instead of re-sending everything |
| High request volume | Retrieval | Sending only relevant passages keeps cost per request down |
The two approaches also combine: retrieve broadly, then place the best passages in a generous window. Anthropic's guidance on this trade-off is that a knowledge base under about 200,000 tokens can often go straight into the prompt, a figure from September 2024 that depends on the model and its pricing, as we cover in RAG vs fine-tuning.[3]
What to watch
The most useful thing a model provider can publish is recall measured across the full window, with the prompts and fact positions disclosed. Until then, your own position test is the evidence that counts. For how to read such claims critically, see how to read an AI benchmark claim.



