Every model launch comes with charts. A bar for the new model is taller than a bar for the old one, and the gap is presented as the news. Benchmarks are genuinely useful, since without them there would be no shared way to compare systems. But a score is a compressed summary of a complicated measurement, and the compression hides what you need to know to trust it.
Two benchmarks, two different questions
Looking at how two well-known benchmarks work shows why a number cannot be read without its context.
SWE-bench: real software issues
SWE-bench, introduced in October 2023 by Carlos Jimenez and colleagues, asks whether language models can resolve real-world GitHub issues. The paper describes "2,294 software engineering problems drawn from real GitHub issues and corresponding pull requests across 12 popular Python repositories."[1] At the time of the paper, the authors reported that "the best-performing model, Claude 2, is able to solve a mere 1.96% of the issues."[1]
Two lessons follow. The benchmark tests a specific skill, editing a codebase to resolve an issue, in a specific language ecosystem. And a score is tied to its time: that 1.96% was a snapshot of late 2023 and says nothing about current systems. Always ask when a number was measured and with which model version.
Chatbot Arena: human preference
Chatbot Arena, described by Wei-Lin Chiang and colleagues in March 2024, works differently. It uses "a pairwise comparison approach and leverages input from a diverse user base through crowdsourcing." The authors reported more than 240,000 votes at the time, and said they confirmed that "crowdsourced questions are sufficiently diverse and discriminating and that the crowdsourced human votes are in good agreement with those of expert raters."[2]
That is a different kind of evidence. It measures which answer people preferred, on the prompts people chose to submit. It is valuable for general chat quality, and it tells you less about a narrow professional task such as extracting fields from your contracts.
The contamination problem
Benchmarks are public, and models are trained on very large amounts of public text. A survey by Cheng Xu and colleagues defines benchmark data contamination as occurring "when language models inadvertently incorporate evaluation benchmark information from their training data, leading to inaccurate or unreliable performance" during assessment.[3] If a model has effectively seen the test, a high score may reflect memorisation instead of capability. Some newer benchmarks respond by using fresh questions or keeping their test sets private, and it is a reason to be cautious about scores on old, widely circulated test sets.
Seven questions before you trust a score
- Who ran it? A vendor's own evaluation, an independent group, or a community leaderboard. Self-reported numbers are not wrong by default, but they are not neutral.
- Which model version, and when? Models are updated, and a score applies to a specific version at a specific time.
- What exactly was the setup? Prompts, number of attempts, tool access, allowed reasoning time and scoring rules can each change a result.
- Is the comparison like-for-like? A chart that compares one model with extra attempts or tools against another without them is not measuring the same thing.
- How large is the test, and how large is the gap? Differences smaller than the noise in the test are not evidence of a real difference.
- Could the model have seen the test? Consider contamination, especially on older public benchmarks.
- Does the task resemble yours? A coding score says little about summarising legal text, and a chat-preference rank says little about extraction accuracy.
Build the benchmark that matters
The most trustworthy evidence for your decision is a small test built from your own work: collect real inputs with the outputs you would accept, run the candidate models on them with identical prompts, and score the results the way you would judge them in practice. Include the hard and unusual cases, because they are where systems differ. Keep part of the set aside so you do not tune against it, and re-run it whenever you change a model or prompt.
This is also how you test retrieval and long-context behaviour, which we cover in RAG vs fine-tuning and long context windows.



