ANTM
ModelsAnalysis

Long Context Windows: Capacity Is Not the Same as Reliability

A model that accepts a very long prompt has not necessarily learned to use all of it. What the research shows, and a test you can run on your own task.

Editorial desk

Published 3 min read

A bright cyan horizon line with fine dotted grid lines receding across a dark plane, flecked with small coral points
Illustration generated with AI (FLUX.1 [schnell] (Black Forest Labs) via Cloudflare Workers AI, Apache 2.0). Prompt and direction by ANTM.

A context window is the amount of text a model can consider in one request. Windows have grown a great deal, and a tempting conclusion follows: skip the retrieval system and put everything in the prompt. Sometimes that is exactly right. But the size of the window answers only one question, how much you can send. It does not answer how well the model will use what you sent.

What the research found

Two widely cited studies address that second question directly. Both examined models available at the time, so they show what to test for, not what is true of every current model.

Position matters

In "Lost in the Middle," Nelson Liu and colleagues studied how models use long inputs. They reported that "performance is often highest when relevant information occurs at the beginning or end of the input context, and significantly degrades when models must access relevant information in the middle of long contexts." The degradation appeared even in models explicitly built for long contexts. The paper was posted in July 2023 and published in the journal Transactions of the Association for Computational Linguistics.[1]

Claimed size is not usable size

The RULER benchmark, from Cheng-Ping Hsieh and colleagues, tested a different weakness. The common "needle in a haystack" test hides one fact in a long document and asks for it. RULER found that "despite achieving nearly perfect accuracy in the vanilla NIAH test, almost all models exhibit large performance drops as the context length increases." It added tasks such as multi-hop tracing and aggregation, and reported that although the tested models "all claim context sizes of 32K tokens or greater, only half of them can maintain satisfactory performance at the length of 32K."[2]

Why this matters for cost and design

Our interpretation of these findings, plus how providers price usage:

  • Accuracy can depend on layout. If your task is sensitive to where facts sit, ordering your prompt carefully, for example putting the question and key instructions near the end, may matter as much as the model choice.
  • Every extra token has a price. Providers charge for input tokens, and processing a long prompt takes longer. Prompt caching can reduce the cost when you reuse the same prefix.
  • More text can mean more distraction. Extra documents are also extra material that can conflict with or crowd out the relevant passage.

A test you can run in an afternoon

You do not need a research benchmark to find out whether this affects you:

  1. Take a realistic long input from your own work.
  2. Insert a known fact at several depths: near the start, the middle and the end.
  3. Ask a question whose answer depends on that fact, and repeat across depths and several prompts.
  4. Plot accuracy by depth and by total length. If accuracy changes with position, your task is position-sensitive at that length.

Run it at the length you plan to use, not the maximum the model accepts.

When long context is the right choice, and when retrieval is

SituationBetter fitReason
One-off analysis of a bounded set of documentsLong contextNo index to build, and cross-document reasoning is the goal
Summarising or checking consistency across one long documentLong contextThe task needs the whole text
A large collection that exceeds any windowRetrievalThe material does not fit
Data that changes dailyRetrievalYou update documents instead of re-sending everything
High request volumeRetrievalSending only relevant passages keeps cost per request down

The two approaches also combine: retrieve broadly, then place the best passages in a generous window. Anthropic's guidance on this trade-off is that a knowledge base under about 200,000 tokens can often go straight into the prompt, a figure from September 2024 that depends on the model and its pricing, as we cover in RAG vs fine-tuning.[3]

What to watch

The most useful thing a model provider can publish is recall measured across the full window, with the prompts and fact positions disclosed. Until then, your own position test is the evidence that counts. For how to read such claims critically, see how to read an AI benchmark claim.

The ANTM newsletter

The signal, not the noise.

Sourced AI coverage in your inbox. Double opt-in, unsubscribe in one click.

Referenced sources

  1. 1.
    Lost in the Middle: How Language Models Use Long Contexts (Liu et al.)(opens in a new tab)

    arXiv / TACLResearchPublished Jul 6, 2023Accessed Oct 3, 2026

  2. 2.
  3. 3.
    Introducing Contextual Retrieval(opens in a new tab)

    AnthropicCompanyPublished Sep 19, 2024Accessed Oct 3, 2026

ANTM Editorial

Editorial desk

The editorial desk at AI's Next Top Model. Every article is sourced to primary documents and approved by an editor before publication. See the editorial policy for how we work.