ANTM

RAG vs Fine-Tuning: When Should You Use Each?

Retrieval-augmented generation and fine-tuning solve different problems. A guide to choosing between them, and to the cheaper option people overlook.

Editorial desk

Published 3 min read

Translucent document cards drifting from the top left toward flowing teal contour lines, with a faceted teal crystal on the right, on a dark navy background
Illustration generated with AI (FLUX.1 [schnell] (Black Forest Labs) via Cloudflare Workers AI, Apache 2.0). Prompt and direction by ANTM.

Retrieval-augmented generation (RAG) and fine-tuning are often presented as rivals. They are better understood as answers to different questions. RAG decides what information the model sees when it answers. Fine-tuning decides how the model behaves by default. If you pick the one that matches your actual problem, the choice is usually clear.

What each approach changes

The term RAG comes from a 2020 paper by Patrick Lewis and colleagues, which explored "a general-purpose fine-tuning recipe for retrieval-augmented generation (RAG) -- models which combine pre-trained parametric and non-parametric memory for language generation." The parametric memory is the pre-trained model. The non-parametric memory, in that paper, was a dense vector index of Wikipedia. The authors reported that their RAG models generated "more specific, diverse and factual language than a state-of-the-art parametric-only seq2seq baseline."[1]

In practice today, RAG usually means: retrieve relevant passages from your own documents at question time, put them in the prompt, and ask the model to answer from them. Fine-tuning means: continue training the model on examples so it behaves differently without being told each time.

FactorRAGFine-tuning
ChangesThe information available at answer timeThe model's default behaviour
Good forPrivate, large or frequently changing knowledge; citations to sourcesConsistent format, tone, classification rules, domain phrasing
Weak atStyle and behaviour that must hold without long instructionsKeeping facts current; showing where an answer came from
UpdatingEdit or add documentsPrepare data and train again

The table is our summary of the trade-offs, not a measured benchmark. The right balance depends on your data and your model.

A decision guide

Ask these questions in order:

  1. Does the model fail because it lacks information? Missing company policies, recent events or private documents point to retrieval. A model cannot reliably recite what it was never given.
  2. Does it fail because it behaves inconsistently? Wrong output format, tone drift or a task it misunderstands point to better instructions and examples first, and to fine-tuning if those are not enough.
  3. Does the information change often? Frequent change favours retrieval, because you update a document instead of retraining a model.
  4. Do you need to show sources? Retrieval can return the passage an answer came from, which makes answers checkable.

The option people skip: put it all in the prompt

Before you build a retrieval pipeline, check whether you need one. Anthropic's guidance on retrieval says that "if your knowledge base is smaller than 200,000 tokens (about 500 pages of material), you can just include the entire knowledge base in the prompt" with no retrieval step at all. It adds that prompt caching makes this approach faster and cheaper, with latency reduced by more than 2x and costs by up to 90%.[2]

That figure dates from September 2024 and describes Anthropic's models and pricing at the time, so check current context limits and caching terms for the model you use. Longer prompts also bring their own problems, which our analysis of long context windows covers.

If you do build retrieval, measure it

Retrieval is where many RAG systems quietly fail. If the right passage is not retrieved, the model cannot use it. Anthropic described one cause: "traditional RAG solutions remove context when encoding information," so a chunk may lose the context that made it meaningful. Its fix, called Contextual Retrieval, prepends explanatory context to each chunk before embedding and indexing. In Anthropic's tests the top-20-chunk retrieval failure rate fell as follows:[2]

TechniqueFailure rateReduction
Baseline5.7%n/a
Contextual embeddings3.7%35%
Contextual embeddings plus BM252.9%49%
Contextual embeddings, BM25 and reranking1.9%67%

These are results from Anthropic's own evaluation datasets, so treat them as evidence that retrieval design matters, not as a number you will reproduce on your data. Build a small set of real questions with known answers, and measure retrieval and answer quality separately.

Where to go next

Retrieval is one half of building reliable AI features. The other half is deciding how a model uses tools, which we cover in what Model Context Protocol is, and how to judge a claim that one system beats another in how to read an AI benchmark claim.

Frequently asked questions

Does RAG stop a model from making things up?
Not by itself. It gives the model relevant text to work from, which helps, but the model can still misread or ignore it. Measure faithfulness to the retrieved passages instead of assuming it.
Can fine-tuning teach a model new facts?
It can shift what a model reproduces, but it is a poor way to keep facts current and hard to audit. For facts that change, retrieval at question time is the usual choice.
Can I use both?
Yes. A common pattern is a lightly tuned model for format and tone, with retrieval for the facts. Add each only when an evaluation shows the need.

The ANTM newsletter

The signal, not the noise.

Sourced AI coverage in your inbox. Double opt-in, unsubscribe in one click.

Referenced sources

  1. 1.
  2. 2.
    Introducing Contextual Retrieval(opens in a new tab)

    AnthropicCompanyPublished Sep 19, 2024Accessed Oct 3, 2026

ANTM Editorial

Editorial desk

The editorial desk at AI's Next Top Model. Every article is sourced to primary documents and approved by an editor before publication. See the editorial policy for how we work.