Retrieval-augmented generation (RAG) and fine-tuning are often presented as rivals. They are better understood as answers to different questions. RAG decides what information the model sees when it answers. Fine-tuning decides how the model behaves by default. If you pick the one that matches your actual problem, the choice is usually clear.
What each approach changes
The term RAG comes from a 2020 paper by Patrick Lewis and colleagues, which explored "a general-purpose fine-tuning recipe for retrieval-augmented generation (RAG) -- models which combine pre-trained parametric and non-parametric memory for language generation." The parametric memory is the pre-trained model. The non-parametric memory, in that paper, was a dense vector index of Wikipedia. The authors reported that their RAG models generated "more specific, diverse and factual language than a state-of-the-art parametric-only seq2seq baseline."[1]
In practice today, RAG usually means: retrieve relevant passages from your own documents at question time, put them in the prompt, and ask the model to answer from them. Fine-tuning means: continue training the model on examples so it behaves differently without being told each time.
| Factor | RAG | Fine-tuning |
|---|---|---|
| Changes | The information available at answer time | The model's default behaviour |
| Good for | Private, large or frequently changing knowledge; citations to sources | Consistent format, tone, classification rules, domain phrasing |
| Weak at | Style and behaviour that must hold without long instructions | Keeping facts current; showing where an answer came from |
| Updating | Edit or add documents | Prepare data and train again |
The table is our summary of the trade-offs, not a measured benchmark. The right balance depends on your data and your model.
A decision guide
Ask these questions in order:
- Does the model fail because it lacks information? Missing company policies, recent events or private documents point to retrieval. A model cannot reliably recite what it was never given.
- Does it fail because it behaves inconsistently? Wrong output format, tone drift or a task it misunderstands point to better instructions and examples first, and to fine-tuning if those are not enough.
- Does the information change often? Frequent change favours retrieval, because you update a document instead of retraining a model.
- Do you need to show sources? Retrieval can return the passage an answer came from, which makes answers checkable.
The option people skip: put it all in the prompt
Before you build a retrieval pipeline, check whether you need one. Anthropic's guidance on retrieval says that "if your knowledge base is smaller than 200,000 tokens (about 500 pages of material), you can just include the entire knowledge base in the prompt" with no retrieval step at all. It adds that prompt caching makes this approach faster and cheaper, with latency reduced by more than 2x and costs by up to 90%.[2]
That figure dates from September 2024 and describes Anthropic's models and pricing at the time, so check current context limits and caching terms for the model you use. Longer prompts also bring their own problems, which our analysis of long context windows covers.
If you do build retrieval, measure it
Retrieval is where many RAG systems quietly fail. If the right passage is not retrieved, the model cannot use it. Anthropic described one cause: "traditional RAG solutions remove context when encoding information," so a chunk may lose the context that made it meaningful. Its fix, called Contextual Retrieval, prepends explanatory context to each chunk before embedding and indexing. In Anthropic's tests the top-20-chunk retrieval failure rate fell as follows:[2]
| Technique | Failure rate | Reduction |
|---|---|---|
| Baseline | 5.7% | n/a |
| Contextual embeddings | 3.7% | 35% |
| Contextual embeddings plus BM25 | 2.9% | 49% |
| Contextual embeddings, BM25 and reranking | 1.9% | 67% |
These are results from Anthropic's own evaluation datasets, so treat them as evidence that retrieval design matters, not as a number you will reproduce on your data. Build a small set of real questions with known answers, and measure retrieval and answer quality separately.
Where to go next
Retrieval is one half of building reliable AI features. The other half is deciding how a model uses tools, which we cover in what Model Context Protocol is, and how to judge a claim that one system beats another in how to read an AI benchmark claim.




