ANTM

How RAG Retrieval Works: Embeddings, Chunking, Hybrid Search and Reranking

Retrieval-augmented generation is only as good as the passages it finds. A walk through the retrieval pipeline, stage by stage, and what each stage can and cannot fix.

Editorial desk

Published 6 min read

A glowing teal wireframe funnel narrowing to a point above a field of scattered teal dots, with a few small coral dots at the right, on a near-black background
Illustration generated with AI (FLUX.1 [schnell] (Black Forest Labs) via Cloudflare Workers AI, Apache 2.0). Prompt and direction by ANTM.

When a retrieval-augmented generation (RAG) system answers wrongly, the instinct is to blame the model or rewrite the prompt. Often the problem happened earlier: the passage that contained the answer was never retrieved, so the model had nothing correct to work from. This guide explains how RAG retrieval works, stage by stage, so that you know which part of the pipeline to inspect first. For the separate question of whether you need RAG at all, see our guide to RAG vs fine-tuning.

The two halves of RAG

The original RAG paper, by Patrick Lewis and colleagues (2020), describes models that "combine pre-trained parametric and non-parametric memory for language generation". In that paper the generator was a pre-trained seq2seq model, the non-parametric memory was a dense vector index of Wikipedia, and a pre-trained neural retriever connected them.[1]

That split still describes every RAG system. Retrieval finds candidate passages for a query. Generation writes an answer using them. The halves fail differently, and they should be measured separately: a retrieval failure means the evidence was absent, a generation failure means the model had the evidence and misused it. Tuning the prompt cannot fix the first kind.

Stage 1: chunking

Documents are split into smaller pieces, because retrieving a whole manual for every question is wasteful and imprecise. Anthropic describes the standard process as breaking documents into chunks, converting the chunks to vector embeddings, and storing them in a searchable database. It also describes chunks as typically a few hundred tokens.[2]

Chunking has a known cost. Anthropic gives an example of a chunk reading "The company's revenue grew by 3% over the previous quarter." On its own, it does not say which company or which quarter, so a query naming the company may not find it.[2] The practical lesson is that a chunk should carry enough context to be understood alone. Common ways to do that are keeping titles and section headings with each chunk, or, as Anthropic proposes, generating a short context note for every chunk (see stage 4).

An embedding model turns text into a vector so that texts with similar meaning land close together. Retrieval then embeds the query and finds the nearest document vectors. Anthropic's documentation describes embeddings as numerical representations of text that enable measuring semantic similarity.[3]

Some details from that documentation are worth knowing because they affect results:

  • Anthropic does not offer its own embedding model. Its documentation points to Voyage AI as one provider, and tells readers to assess several vendors.[3]
  • Queries and documents are embedded differently. For retrieval, Voyage AI's input_type parameter should be set to query or document; the documentation says this prepends a different instruction to the text before embedding and advises not to omit it.[3]
  • Similarity function. Voyage embeddings are normalised to length 1, so dot product and cosine similarity give the same ranking, and the dot product is faster to compute.[3]
  • Storage trade-offs. The documentation lists quantisation options that reduce storage, memory and cost by 4x (8-bit integers) or 32x (single bits), and notes that float output gives the highest accuracy. It also describes shortening vectors from some models by keeping the leading dimensions and normalising again.[3]

The minimal retrieval step from that documentation looks like this (Python, voyageai and numpy; check the current package version and model names before copying):

python
import numpy as np
import voyageai

vo = voyageai.Client()  # reads VOYAGE_API_KEY from the environment

doc_embds = vo.embed(documents, model="voyage-4", input_type="document").embeddings
query_embd = vo.embed([query], model="voyage-4", input_type="query").embeddings[0]

similarities = np.dot(doc_embds, query_embd)
best = documents[int(np.argmax(similarities))]

This is the whole idea in six lines. Production systems replace the in-memory array with a vector index, but the logic is the same. Model names change often, so treat voyage-4 as the example in the documentation as of 2026-10-06, not a recommendation.

Stage 3: lexical search and hybrid retrieval

Embeddings match meaning, which is also their weakness: an exact string can get lost. Anthropic's example is an error code, "Error code TS-999". An embedding search may return generic error-handling passages, while BM25, a ranking function based on lexical matching, looks for the specific text string.[2]

The academic evidence points the same way. The BEIR benchmark evaluated retrieval models zero-shot across many tasks and reported that "BM25 is a robust baseline and re-ranking and late-interaction-based models on average achieve the best zero-shot performances", while dense retrieval models "often underperform other approaches" in generalisation.[4] That result is from 2021 and the models it tested have since been superseded, so read it as evidence that lexical search deserves a place in the pipeline, not as a ranking of today's embedding models.

Hybrid retrieval runs both searches and merges the candidate lists. Anthropic's tests combined embeddings with BM25 for exactly this reason.[2]

Stage 4: contextual chunks and reranking

Anthropic's Contextual Retrieval adds a short, chunk-specific explanation of where the chunk sits in the document before it is embedded and indexed for BM25. Anthropic reports the following results, measured as a reduction in the rate of failed retrievals (relevant chunks missing from the top results):[2]

TechniqueReported reduction in retrieval failures
Contextual embeddings alone35%
Contextual embeddings plus contextual BM25 (top-20 chunks)49%
The above plus reranking67%

Anthropic reports the tests across codebases, fiction, arXiv papers and scientific papers, and says passing 20 chunks to the model worked better than 5 or 10. It also names the trade-offs: the one-time cost of generating context was $1.02 per million document tokens in its setup, using prompt caching, and reranking adds latency.[2] These are the vendor's own measurements on its own datasets and 2024 pricing. Your documents may behave differently, so run the comparison on your own data.

Reranking is a second pass that scores each retrieved chunk against the query and keeps only the best. It lets the first stage cast a wide net cheaply and the second stage be selective. Voyage AI offers reranker models for this step, listed in the same documentation.[3]

Stage 5: what the model does with the passages

Retrieval ends when passages are placed in the prompt, but their placement matters. The "Lost in the Middle" study found that model performance "can degrade significantly" when relevant information moves, and is often highest when it sits at the beginning or end of the input and worst in the middle.[5] Sending 20 chunks therefore has a cost, and order and count are worth testing. Our article on long context windows covers this in more depth, and tokens, context windows and cost explains what extra chunks cost per request.

Where to look first when RAG fails

Our reading of the sources above suggests this order of investigation. It is editorial guidance, not a benchmarked procedure:

  1. Was the right passage in the corpus at all? If not, no pipeline change helps.
  2. Was it retrieved? Log the top results for failing queries. If the passage is absent, inspect chunking, then embeddings, then add lexical search if the query contains identifiers or exact terms.
  3. Was it ranked high enough? If it is retrieved but buried, try reranking or fewer, better chunks.
  4. Did the model use it? Only now is the prompt or the model the suspect. Check passage order and count.

Build a small set of real questions with the known correct passage for each, and track retrieval hit rate separately from answer quality. That one habit separates retrieval problems from generation problems, and it is what lets you tell whether any change in the table above helps your data.

What this article did not test

We did not run these techniques ourselves. Figures come from the cited sources and describe their setups and dates. The numbers for contextual retrieval are the vendor's own; the BEIR result is from 2021. Treat both as directional, and verify against your own corpus before committing to a design.

Frequently asked questions

What is the difference between embeddings search and BM25?
Embedding search compares vectors that capture semantic meaning, so it can match a question to a passage that uses different words. BM25 is a lexical ranking function that looks for specific words or strings, which makes it strong for identifiers and exact terms.[2]
Do I need a vector database for RAG?
Not for a small corpus. Anthropic notes that a knowledge base under about 200,000 tokens can be placed directly in the prompt with no retrieval step. Check that first.[2]
Why does the same text get a different embedding as a query and as a document?
Some embedding services take an input type for retrieval. Voyage AI documents that setting it to query or document prepends a different instruction before embedding, which can improve retrieval quality.[3]
What is reranking?
A second stage that scores the retrieved chunks against the query and keeps only the best, so the model receives fewer, more relevant passages.[2]

The ANTM newsletter

The signal, not the noise.

Sourced AI coverage in your inbox. Double opt-in, unsubscribe in one click.

Referenced sources

  1. 1.
  2. 2.
    Introducing Contextual Retrieval(opens in a new tab)

    AnthropicCompanyPublished Sep 19, 2024Accessed Oct 3, 2026

  3. 3.
    Embeddings (Claude API documentation)(opens in a new tab)

    AnthropicPrimary sourcePublished Oct 6, 2026Accessed Oct 3, 2026

  4. 4.
  5. 5.
    Lost in the Middle: How Language Models Use Long Contexts(opens in a new tab)

    arXivResearchPublished Jul 6, 2023Accessed Oct 3, 2026

ANTM Editorial

Editorial desk

The editorial desk at AI's Next Top Model. Every article is sourced to primary documents and approved by an editor before publication. See the editorial policy for how we work.