When a retrieval-augmented generation (RAG) system answers wrongly, the instinct is to blame the model or rewrite the prompt. Often the problem happened earlier: the passage that contained the answer was never retrieved, so the model had nothing correct to work from. This guide explains how RAG retrieval works, stage by stage, so that you know which part of the pipeline to inspect first. For the separate question of whether you need RAG at all, see our guide to RAG vs fine-tuning.
The two halves of RAG
The original RAG paper, by Patrick Lewis and colleagues (2020), describes models that "combine pre-trained parametric and non-parametric memory for language generation". In that paper the generator was a pre-trained seq2seq model, the non-parametric memory was a dense vector index of Wikipedia, and a pre-trained neural retriever connected them.[1]
That split still describes every RAG system. Retrieval finds candidate passages for a query. Generation writes an answer using them. The halves fail differently, and they should be measured separately: a retrieval failure means the evidence was absent, a generation failure means the model had the evidence and misused it. Tuning the prompt cannot fix the first kind.
Stage 1: chunking
Documents are split into smaller pieces, because retrieving a whole manual for every question is wasteful and imprecise. Anthropic describes the standard process as breaking documents into chunks, converting the chunks to vector embeddings, and storing them in a searchable database. It also describes chunks as typically a few hundred tokens.[2]
Chunking has a known cost. Anthropic gives an example of a chunk reading "The company's revenue grew by 3% over the previous quarter." On its own, it does not say which company or which quarter, so a query naming the company may not find it.[2] The practical lesson is that a chunk should carry enough context to be understood alone. Common ways to do that are keeping titles and section headings with each chunk, or, as Anthropic proposes, generating a short context note for every chunk (see stage 4).
Stage 2: embeddings and similarity search
An embedding model turns text into a vector so that texts with similar meaning land close together. Retrieval then embeds the query and finds the nearest document vectors. Anthropic's documentation describes embeddings as numerical representations of text that enable measuring semantic similarity.[3]
Some details from that documentation are worth knowing because they affect results:
- Anthropic does not offer its own embedding model. Its documentation points to Voyage AI as one provider, and tells readers to assess several vendors.[3]
- Queries and documents are embedded differently. For retrieval, Voyage AI's
input_typeparameter should be set toqueryordocument; the documentation says this prepends a different instruction to the text before embedding and advises not to omit it.[3] - Similarity function. Voyage embeddings are normalised to length 1, so dot product and cosine similarity give the same ranking, and the dot product is faster to compute.[3]
- Storage trade-offs. The documentation lists quantisation options that reduce storage, memory and cost by 4x (8-bit integers) or 32x (single bits), and notes that float output gives the highest accuracy. It also describes shortening vectors from some models by keeping the leading dimensions and normalising again.[3]
The minimal retrieval step from that documentation looks like this (Python, voyageai and numpy; check the current package version and model names before copying):
import numpy as np
import voyageai
vo = voyageai.Client() # reads VOYAGE_API_KEY from the environment
doc_embds = vo.embed(documents, model="voyage-4", input_type="document").embeddings
query_embd = vo.embed([query], model="voyage-4", input_type="query").embeddings[0]
similarities = np.dot(doc_embds, query_embd)
best = documents[int(np.argmax(similarities))]This is the whole idea in six lines. Production systems replace the in-memory array with a vector index, but the logic is the same. Model names change often, so treat voyage-4 as the example in the documentation as of 2026-10-06, not a recommendation.
Stage 3: lexical search and hybrid retrieval
Embeddings match meaning, which is also their weakness: an exact string can get lost. Anthropic's example is an error code, "Error code TS-999". An embedding search may return generic error-handling passages, while BM25, a ranking function based on lexical matching, looks for the specific text string.[2]
The academic evidence points the same way. The BEIR benchmark evaluated retrieval models zero-shot across many tasks and reported that "BM25 is a robust baseline and re-ranking and late-interaction-based models on average achieve the best zero-shot performances", while dense retrieval models "often underperform other approaches" in generalisation.[4] That result is from 2021 and the models it tested have since been superseded, so read it as evidence that lexical search deserves a place in the pipeline, not as a ranking of today's embedding models.
Hybrid retrieval runs both searches and merges the candidate lists. Anthropic's tests combined embeddings with BM25 for exactly this reason.[2]
Stage 4: contextual chunks and reranking
Anthropic's Contextual Retrieval adds a short, chunk-specific explanation of where the chunk sits in the document before it is embedded and indexed for BM25. Anthropic reports the following results, measured as a reduction in the rate of failed retrievals (relevant chunks missing from the top results):[2]
| Technique | Reported reduction in retrieval failures |
|---|---|
| Contextual embeddings alone | 35% |
| Contextual embeddings plus contextual BM25 (top-20 chunks) | 49% |
| The above plus reranking | 67% |
Anthropic reports the tests across codebases, fiction, arXiv papers and scientific papers, and says passing 20 chunks to the model worked better than 5 or 10. It also names the trade-offs: the one-time cost of generating context was $1.02 per million document tokens in its setup, using prompt caching, and reranking adds latency.[2] These are the vendor's own measurements on its own datasets and 2024 pricing. Your documents may behave differently, so run the comparison on your own data.
Reranking is a second pass that scores each retrieved chunk against the query and keeps only the best. It lets the first stage cast a wide net cheaply and the second stage be selective. Voyage AI offers reranker models for this step, listed in the same documentation.[3]
Stage 5: what the model does with the passages
Retrieval ends when passages are placed in the prompt, but their placement matters. The "Lost in the Middle" study found that model performance "can degrade significantly" when relevant information moves, and is often highest when it sits at the beginning or end of the input and worst in the middle.[5] Sending 20 chunks therefore has a cost, and order and count are worth testing. Our article on long context windows covers this in more depth, and tokens, context windows and cost explains what extra chunks cost per request.
Where to look first when RAG fails
Our reading of the sources above suggests this order of investigation. It is editorial guidance, not a benchmarked procedure:
- Was the right passage in the corpus at all? If not, no pipeline change helps.
- Was it retrieved? Log the top results for failing queries. If the passage is absent, inspect chunking, then embeddings, then add lexical search if the query contains identifiers or exact terms.
- Was it ranked high enough? If it is retrieved but buried, try reranking or fewer, better chunks.
- Did the model use it? Only now is the prompt or the model the suspect. Check passage order and count.
Build a small set of real questions with the known correct passage for each, and track retrieval hit rate separately from answer quality. That one habit separates retrieval problems from generation problems, and it is what lets you tell whether any change in the table above helps your data.
What this article did not test
We did not run these techniques ourselves. Figures come from the cited sources and describe their setups and dates. The numbers for contextual retrieval are the vendor's own; the BEIR result is from 2021. Treat both as directional, and verify against your own corpus before committing to a design.




