Why AI benchmarks keep lying to you
Every model launch arrives with a table of numbers, and every table shows the new model winning. You’ve probably noticed that the winner changes depending on whose slide deck you’re reading, which is the first clue that something is wrong with how we measure this stuff.
The problem isn’t that labs lie. It’s that the benchmarks themselves have quietly stopped measuring what everyone assumes they measure.
Contamination: the test is already in the training data
Start with contamination, because it’s the least understood and the most damaging. A benchmark works only if the model hasn’t seen the answers. But these models are trained on scraped internet text, and benchmarks live on the internet.
So the test leaks into the training data. And the numbers on that are not small. Empirical audits have found leakage ranging from 1% to 45% across popular question-answering benchmarks, with contamination growing over time as more of the internet gets scraped.
Even Meta’s own LLaMA-2 report found that over 16% of MMLU samples were contaminated. That’s a lab checking its own homework and admitting a sixth of one of the most-cited benchmarks was compromised.
A model that memorised the answer and a model that worked it out score identically. The benchmark cannot tell them apart.
Why decontamination does not fix it
You’d think decontamination would fix this, and labs do run it. Stripping out test items that appear verbatim in training data. The trouble is that verbatim matching is a weak filter.
Paraphrased or translated benchmark items slip straight through while still inflating scores. Leakage even crosses language boundaries, which means a question memorised in Arabic can lift an English score without any surface-overlap detector noticing.
That’s the part worth sitting with. The detection tools are looking for exact text, and the contamination has already moved past exact text.
It isn’t confined to text either. A systematic analysis of multimodal models found both images and text leaking into training sets, so the same problem now applies to anything claiming to score well on vision tasks.
Optimisation pressure, without anyone cheating
Which brings us to the second problem, and it’s structural rather than accidental. Benchmarks are public, fixed, and enormously commercially valuable to score well on. Anything with those three properties gets optimised for.
Nobody has to cheat for this to distort results. Choosing training data that resembles the test, tuning until the number goes up, running the eval many times and reporting the best. All of it is normal engineering, and all of it widens the gap between benchmark performance and real-world usefulness.
Then there’s selection. A lab runs dozens of evaluations and publishes the flattering ones. That’s not fraud, it’s marketing, but it means a launch table tells you where a model is strong and almost nothing about where it’s weak.
What a composite AI benchmark score hides
You can see the compression problem in any composite score. When Grok 4.6 landed at 61 on the Artificial Analysis Intelligence Index, level with GPT-5.6 Sol Max. That single figure was averaging across wildly different capabilities. Dig into the components and the same model scored 15.8% on a legal reasoning benchmark where a rival managed 2.5%.
Two models “tied” on the index while being six times apart on a task someone might actually care about. The tie is real and the tie is useless.
So what survives all this? Benchmarks that resist contamination by construction. Held-out sets that were never published, evaluations that refresh their questions, and continuously evolving benchmarks that swap items faster than they can leak. Those numbers mean considerably more than a score on a fixed public set.
Replace AI benchmarks with your own ten test cases
The practical move, though, is to stop outsourcing the judgement. Build ten test cases from your own work, the messy ones, the ambiguous ones. And run every model you’re considering against them. It takes an afternoon and it tells you more than every launch table combined.
Because that’s the thing the leaderboard can’t do. It can’t know what you need the model for, and it was never really trying to.
Read the next announcement with that in mind. Ask what was measured, whether the set could have leaked, and how many evaluations went unmentioned. Our piece on whether scaling has stalled runs into the same wall. A lot of that argument rests on benchmark numbers nobody can independently verify.
Get the daily rundown
One email each weekday with the AI news that matters, every claim linked to its primary source.
Free, one email each weekday, unsubscribe in one click. We never sell or share your address.
