Ask a language model about a person it has barely seen in training and it will often answer anyway, fluently and with the wrong details. That behaviour is called hallucination, and the useful question for anyone building with these systems is not when it will disappear but why it happens and what reduces it. This guide covers why AI models hallucinate, what the research says about the incentives behind it, and which mitigations have documented support. For a related problem, where the right evidence is simply never shown to the model, see how RAG retrieval works.
What counts as a hallucination
Anthropic's documentation defines the term plainly: language models "can sometimes generate text that is factually incorrect or inconsistent with the given context."[2] That definition covers two different failures worth separating:
- Unsupported facts: the model states something false about the world, such as an invented detail about a person.
- Context violations: the model is given a document and says something the document does not support.
The first is mostly a knowledge problem. The second is a grounding problem, and it is the one you can engineer against most directly.
Why they happen: the research explanation
In "Why Language Models Hallucinate" (September 2025), Adam Tauman Kalai, Ofir Nachum, Santosh Vempala and Edwin Zhang argue that language models generate false information because training and evaluation procedures reward guessing over acknowledging uncertainty. They describe hallucinations as stemming from "errors in binary classification": when a model cannot distinguish incorrect statements from facts, it can produce plausible falsehoods.[1]
Two parts of the argument matter for practice.
Origin in pretraining. The paper's claim is that hallucinations are caused by errors made during pretraining, not only by later stages. Facts that appear very rarely in training data are the hardest case, because the model has little basis for telling a true statement from a plausible one.[1]
Incentives in evaluation. The authors argue the problem persists because benchmarks reward confident answers, leaving models "optimized to be good test-takers". Under a scoring rule that gives credit for a lucky guess and nothing for saying "I don't know", guessing is the better strategy. Their proposed remedy is socio-technical: change how existing benchmarks are scored so uncertainty is not penalised, rather than adding yet another hallucination test.[1]
Our interpretation: this is why you should be sceptical of any single accuracy number. A model tuned to maximise a benchmark that never rewards abstention has little incentive to abstain in your application either. Our guide to reading a benchmark claim covers what else to check.
Bigger is not automatically more truthful
The TruthfulQA benchmark (Lin, Hilton and Evans, 2021) contains 817 questions across 38 categories, designed to elicit common misconceptions. In the authors' tests, the best model was truthful on 58% of questions, against 94% for humans. They also reported that larger models were generally less truthful, and concluded that "scaling up models alone is less promising for improving truthfulness" than fine-tuning approaches that go beyond imitating web text.[3]
Caveats: those results are from 2021 and tested GPT-3, GPT-Neo/J, GPT-2 and a T5-based model, all far older than current systems. Do not read the 58% as a statement about any model you use today. The durable lesson is narrower: model size is not a safeguard on its own.
What reduces hallucination: documented techniques
Anthropic's guide groups its recommendations into basic and advanced techniques. The table summarises them. The "What it costs you" column is our reading, not the documentation's.
| Technique | What the documentation says | What it costs you |
|---|---|---|
| Allow "I don't know" | Explicitly give the model permission to admit uncertainty; described as a simple technique that can drastically reduce false information.[2] | Some answers become refusals, so measure how often valid questions are declined. |
| Quote first | For long documents (over 20k tokens), ask for word-for-word quotes before the task, then base the analysis on them.[2] | Longer outputs and an extra step in the prompt. |
| Cite and retract | Have the model cite a supporting quote for each claim; if it cannot find one, it must retract the claim.[2] | A verification pass; quotes must still be checked against the source. |
| Restrict external knowledge | Instruct the model to use only the provided documents, not general knowledge.[2] | Fewer answers beyond the documents, which is often the intent. |
| Best-of-N comparison | Run the same prompt several times; inconsistent outputs can indicate hallucination.[2] | Multiplies request cost by N. |
| Iterative refinement | Feed outputs back and ask the model to verify or expand earlier statements.[2] | More latency and more tokens. |
| Review reasoning | Use extended thinking with summarised output and read the thinking blocks when an answer looks wrong.[2] | Human attention, and only after a suspicious answer. |
The same page carries the important caveat: these techniques "significantly reduce hallucinations" but "don't eliminate them entirely", and critical information should always be validated, especially for high-stakes decisions.[2]
Retrieval helps, with a condition
Retrieval-augmented generation is the architectural version of the grounding advice above. In the original RAG paper, the authors report that RAG models "generate more specific, diverse and factual language" than a parametric-only baseline, using a dense vector index of Wikipedia as non-parametric memory.[4] That finding supports grounding in retrieved text as a way to improve factual output.
The condition is that retrieval has to find the right passage. If it does not, the model is again answering from its parameters, and the failure looks identical to a hallucination from the outside. Measure retrieval and generation separately, as described in our RAG retrieval guide.
A practical checklist
- Decide what a wrong answer costs. Low-stakes drafting tolerates errors. Legal, medical, financial or code-execution paths need verification by something other than the same model.
- Ground the answer. Put the source text in the prompt, or retrieve it, and instruct the model to rely only on it.
- Permit abstention, then test it. Add an explicit "say you do not know" instruction and run a set of questions whose answers are absent from the documents. Count how often the model abstains correctly.
- Require evidence. Ask for quotes or citations per claim and check a sample against the source by hand.
- Build your own evaluation. Because public benchmarks may reward guessing, a small test set drawn from your real task will tell you more than a leaderboard position. Label it as an internal test when you report it.
- Keep a human in the loop where it matters. The documentation's own advice is to validate critical information.[2]
What remains uncertain
This article rests on one theoretical paper, one vendor guidance page, one 2021 benchmark and the original RAG paper. We have not run our own hallucination tests, so nothing here is a measured comparison between current models. The 2025 paper offers an explanation and a proposed change to evaluation practice; whether benchmark scoring changes widely is something to watch rather than assume. Mitigation effectiveness also varies by task, so the techniques in the table are starting points to test, not guarantees.




