Models & Research

Stanford AI Index: benchmarks are now saturating within months

SWE-bench Verified went from 60% to near 100% in a single year. Humanity’s Last Exam, a test built specifically to be hard for machines, moved 30 percentage points over the same period.

Both figures come from Stanford’s 2026 AI Index, and together they describe a measurement problem rather than only a capability one.

The Stanford AI Index shows benchmarks expiring in months

The Index puts it plainly: evaluations meant to stand for years are being saturated in months, which compresses the window in which any benchmark is useful for tracking progress.

A test everyone passes measures nothing. The industry keeps building harder exams and keeps running out of headroom faster than it can write them.

That has a practical consequence for anyone reading a launch announcement. A model claiming a strong score on a benchmark released two years ago is telling you very little, because the ceiling was reached before it shipped.

It also makes comparison between models harder rather than easier. When several systems all sit above 95%, the remaining differences are within the noise of how the test was run.

The agent numbers moved most

BenchmarkMove in a yearWhat it tests
SWE-bench Verified60% to near 100%Resolving real software issues
Humanity’s Last Exam+30 pointsExpert-level questions across fields
OSWorld~12% to 66.3%Agents operating a computer
RLBench89.4% successRobotic manipulation, in simulation
Real household tasks12% successRobots, in an actual home
Source: Stanford AI Index 2026. The last two rows are the same capability, measured twice.

OSWorld is the number to sit with. An agent driving an operating system went from barely functional to within six points of human performance, and that is the capability every enterprise product is now being built on.

It’s also the gap between a benchmark and a deployment. Our piece on why long-horizon agents fail covers what happens when those single-task scores get chained into a fifty-step run.

Who writes the next exam

Building a benchmark that lasts is now a research problem in itself, and the incentives around it are awkward.

A hard, well-designed evaluation takes months of expert work and is obsolete within a year of publication. Meanwhile the labs whose models it grades are the best-resourced organisations able to build one.

Publishing it is what kills it. Once the questions are public they end up in a crawl, then in a corpus, and the next generation has effectively seen the exam.

Which pushes serious evaluation toward held-out private sets, and those cannot be independently checked. Everyone ends up trusting a number they cannot verify, from a party with an interest in it.

Why the jump happened when it did

A 40-point move on a coding benchmark in twelve months is not the shape of ordinary scaling, and the cause is reasonably well understood.

Models started spending compute at answer time. Google Research showed in Chain-of-Thought Prompting that working through steps beat a fine-tuned model on maths, and reinforcement learning later turned that into trained behaviour.

DeepSeek’s R1 work reported “self-reflection, verification, and dynamic strategy adaptation” emerging where answers could be checked mechanically, and every benchmark in that table has a checkable answer.

Which is the qualifier the headline numbers drop. These gains concentrate in domains with ground truth, and our explainer on what thinking compute buys covers where they stop.

What saturation hides

Contamination is the obvious worry when scores climb this fast. A benchmark published online eventually appears in training data, and a model that has seen the answers is not solving the problem.

Nobody outside the labs can rule that out, because nobody outside the labs knows what went into the corpus. That is the same transparency gap our piece on why benchmarks keep lying works through.

The robotics rows are the useful corrective. Simulated manipulation is at 89.4% and real household tasks sit at 12%, which is the clearest available illustration that a controlled test and the world are different problems.

Language benchmarks have the same gap; it’s just harder to see, because there’s no kitchen floor to drop the plate on.

Long-context claims deserve the same scepticism. Stanford’s Lost in the Middle found accuracy degrades sharply for information in the middle of a long input, so a stated window size is a capacity figure rather than a usable one.

And where humans score the output, presentation gets rewarded alongside substance. OpenAI’s InstructGPT work found a 1.3B tuned model preferred over 175B GPT-3, which should temper how much weight anyone puts on preference-based rankings.

The gap between a score and a deployment

Take SWE-bench Verified at close to 100% and ask the obvious question: why is anyone still employing software engineers?

Because the benchmark is a well-specified task with a test suite attached. Someone wrote the issue clearly, the repository builds, and success is a passing test.

Real work arrives as a vague complaint about a report being wrong, in a codebase nobody has documented, where the fix has to not break three things nobody remembered to test.

None of that is a criticism of the benchmark. It measures what it says it measures, and the error is in reading a solved benchmark as a solved job.

The same reasoning applies to OSWorld at 66.3%. Operating a computer for one task is genuinely useful, and a third of attempts still failing is a very different product from one that works.

What replaces the leaderboard

Held-out evaluations that nobody publishes are the honest answer, and they’re exactly what a public leaderboard cannot be.

For a buying decision, your own tasks beat any of this. Ten real examples from your workload, run against two models, tells you more than a table of scores that every vendor optimised toward.

The Index’s own summary of the year is worth reading alongside the charts, since its twelve takeaways put the capability numbers next to the deployment ones rather than in isolation.

The clearest illustration of that gap is physical, with robots scoring 89% in simulation and 12% in a real house.

Watch for evaluations that measure cost and reliability alongside accuracy, since that’s where the Index says competition has moved, and it’s the part no current benchmark reports.

Get the daily rundown

One email each weekday with the AI news that matters, every claim linked to its primary source.

Free, one email each weekday, unsubscribe in one click. We never sell or share your address.

Rundowns AI Desk

The Rundowns AI desk covers artificial intelligence research, tools, business and policy. Every factual claim we publish links to the primary source it came from, so readers can check it themselves.

Leave a Reply

Your email address will not be published. Required fields are marked *