
Models
How to Read an AI Benchmark Claim Without Being Misled
A single benchmark number hides who ran the test, on what, and what the model had already seen. Seven questions that make a claim checkable.
Guide3 min read
Topic
Papers and results that change what is possible, read closely and explained plainly.
2 stories
A single benchmark number hides who ran the test, on what, and what the model had already seen. Seven questions that make a claim checkable.
Guide3 min read

A model that accepts a very long prompt has not necessarily learned to use all of it. What the research shows, and a test you can run on your own task.
Analysis3 min read