ANTM
ModelsGuide

How to Read an AI Benchmark Claim Without Being Misled

A single benchmark number hides who ran the test, on what, and what the model had already seen. Seven questions that make a claim checkable.

Editorial desk

Published 3 min read

A horizontal bar under a circular magnifying lens that shows its fine texture, above a ruler of plain tick marks
Illustration generated with AI (FLUX.1 [schnell] (Black Forest Labs) via Cloudflare Workers AI, Apache 2.0). Prompt and direction by ANTM.

Every model launch comes with charts. A bar for the new model is taller than a bar for the old one, and the gap is presented as the news. Benchmarks are genuinely useful, since without them there would be no shared way to compare systems. But a score is a compressed summary of a complicated measurement, and the compression hides what you need to know to trust it.

Two benchmarks, two different questions

Looking at how two well-known benchmarks work shows why a number cannot be read without its context.

SWE-bench: real software issues

SWE-bench, introduced in October 2023 by Carlos Jimenez and colleagues, asks whether language models can resolve real-world GitHub issues. The paper describes "2,294 software engineering problems drawn from real GitHub issues and corresponding pull requests across 12 popular Python repositories."[1] At the time of the paper, the authors reported that "the best-performing model, Claude 2, is able to solve a mere 1.96% of the issues."[1]

Two lessons follow. The benchmark tests a specific skill, editing a codebase to resolve an issue, in a specific language ecosystem. And a score is tied to its time: that 1.96% was a snapshot of late 2023 and says nothing about current systems. Always ask when a number was measured and with which model version.

Chatbot Arena: human preference

Chatbot Arena, described by Wei-Lin Chiang and colleagues in March 2024, works differently. It uses "a pairwise comparison approach and leverages input from a diverse user base through crowdsourcing." The authors reported more than 240,000 votes at the time, and said they confirmed that "crowdsourced questions are sufficiently diverse and discriminating and that the crowdsourced human votes are in good agreement with those of expert raters."[2]

That is a different kind of evidence. It measures which answer people preferred, on the prompts people chose to submit. It is valuable for general chat quality, and it tells you less about a narrow professional task such as extracting fields from your contracts.

The contamination problem

Benchmarks are public, and models are trained on very large amounts of public text. A survey by Cheng Xu and colleagues defines benchmark data contamination as occurring "when language models inadvertently incorporate evaluation benchmark information from their training data, leading to inaccurate or unreliable performance" during assessment.[3] If a model has effectively seen the test, a high score may reflect memorisation instead of capability. Some newer benchmarks respond by using fresh questions or keeping their test sets private, and it is a reason to be cautious about scores on old, widely circulated test sets.

Seven questions before you trust a score

  1. Who ran it? A vendor's own evaluation, an independent group, or a community leaderboard. Self-reported numbers are not wrong by default, but they are not neutral.
  2. Which model version, and when? Models are updated, and a score applies to a specific version at a specific time.
  3. What exactly was the setup? Prompts, number of attempts, tool access, allowed reasoning time and scoring rules can each change a result.
  4. Is the comparison like-for-like? A chart that compares one model with extra attempts or tools against another without them is not measuring the same thing.
  5. How large is the test, and how large is the gap? Differences smaller than the noise in the test are not evidence of a real difference.
  6. Could the model have seen the test? Consider contamination, especially on older public benchmarks.
  7. Does the task resemble yours? A coding score says little about summarising legal text, and a chat-preference rank says little about extraction accuracy.

Build the benchmark that matters

The most trustworthy evidence for your decision is a small test built from your own work: collect real inputs with the outputs you would accept, run the candidate models on them with identical prompts, and score the results the way you would judge them in practice. Include the hard and unusual cases, because they are where systems differ. Keep part of the set aside so you do not tune against it, and re-run it whenever you change a model or prompt.

This is also how you test retrieval and long-context behaviour, which we cover in RAG vs fine-tuning and long context windows.

Frequently asked questions

Are benchmark scores useless?
No. They are useful for comparing systems on a defined task and for spotting progress. They are weak evidence about how a model will do on your task, so use them to shortlist and your own evaluation to decide.
What is data contamination?
It happens when evaluation benchmark information appears in a model's training data, which can lead to inaccurate or unreliable performance measurements. [3]
Should I trust a leaderboard rank?
Check what it measures. A rank from human preference votes tells you what raters preferred in their prompts, not how a model performs on your workflow.

The ANTM newsletter

The signal, not the noise.

Sourced AI coverage in your inbox. Double opt-in, unsubscribe in one click.

Referenced sources

  1. 1.
  2. 2.
  3. 3.

ANTM Editorial

Editorial desk

The editorial desk at AI's Next Top Model. Every article is sourced to primary documents and approved by an editor before publication. See the editorial policy for how we work.