Models & Research

Stanford: robots score 89% in simulation and 12% in a real house

Robotic manipulation now succeeds 89.4% of the time on RLBench, a simulated benchmark. In real household tasks, robots succeed 12% of the time.

Both numbers are from Stanford’s 2026 AI Index, and the distance between them is the most honest thing published about AI progress this year.

What simulation leaves out

A simulator gives a policy clean state: known object positions, consistent lighting, no friction it wasn’t modelled for, and a reset button when something goes wrong.

A kitchen gives it none of those. The mug is where somebody left it, the light changes through the day, and a dropped plate does not reset.

Nine out of ten in the lab and one in eight in a house is not a tuning problem. It is a statement about which problem was actually solved.

The pattern is not unique to robots. It’s the same gap that separates benchmark scores from production behaviour in language models, and physical failure just makes it impossible to talk around.

Driving is the counter-example

One physical domain did cross into deployment, and how it did so is instructive.

SystemScaleConstraint
Waymo~450,000 weekly trips, five US citiesMapped territory only
Apollo Go11 million driverless rides, up 175% year on yearDefined operating zones
Household robots12% task successAnything, anywhere in a home
Source: Stanford AI Index 2026. The successful cases narrowed the world first.

Roads are enormously more constrained than homes. Lanes are marked, behaviour is codified, other drivers mostly follow rules, and the operator gets to choose which streets are in scope and survey them first.

That’s the transferable lesson. Autonomy works where the environment can be bounded, and a household is defined by not being bounded.

What the AI robotics benchmark was built to measure

RLBench was never meant to stand in for a house. James and colleagues introduced it as a robot learning benchmark and learning environment, a controlled suite for comparing algorithms.

That is a legitimate and useful thing to build. The distortion comes later, when a score meant for comparing methods gets quoted as a statement about capability in the world.

The field has been trying to close the gap with scale and with language models. Google’s RT-2 work put vision-language models in the control loop so a robot could act on instructions it had never been trained on.

Data sharing is the other route, with Open X-Embodiment pooling demonstrations across many robot types to build the corpus no single lab could collect. Both help, and neither has yet produced a household number anyone wants to print.

The economics are different too

Software capability ships at close to zero marginal cost. A better model is a deployment, and every user gets it the same afternoon.

Robotics has to manufacture something. A better policy still needs hardware built, shipped, installed and maintained, and the hardware wears out.

That difference explains why the two fields feel like they are moving at incomparable speeds even when the underlying research is related. One compounds instantly and the other compounds at the pace of factories.

It also explains the deployment pattern in the table. Both driverless services scaled by concentrating a large fleet in a small mapped area, which is the only way the fixed costs work out.

Why the software gains do not carry over

Language models improved on a supply of text that already existed. Robotics has no equivalent corpus, because interaction data has to be generated by robots doing things.

Simulation is the obvious workaround and it inherits the modeller’s assumptions, which is precisely the gap the 12% figure measures.

The data-supply problem is showing up on the language side too. Villalobos and colleagues projected models being trained on datasets matching the stock of public human text between 2026 and 2032, and our piece on synthetic data covers why generating your way out is harder than it sounds.

Failure cost differs too. A wrong sentence gets edited; a wrong grasp breaks something, and no amount of model quality changes that asymmetry.

Why this matters beyond robotics

The 89% and the 12% are the same capability measured under two different assumptions about the world, and that framing transfers directly to software.

An agent scoring well on a benchmark is operating in a version of the world someone prepared: the task is stated clearly, the tools respond, the data is where it should be.

Deployed, it meets an API that times out, a document in a format nobody anticipated, and an instruction written by someone who assumed context the agent does not have.

The failure is quieter than a dropped plate, which is exactly what makes it harder to manage. A robot that fails is obvious; an agent that fails writes a confident paragraph.

So the useful habit from robotics is to always ask what the environment was, not only what the score was. It is the question that separates a demo from a product in both fields.

How to read an AI robotics benchmark demo

Ask whether the environment was prepared, how many attempts the clip represents, whether anything in the scene moved between takes, and whether the operator had a hand near the stop button.

Ask for a success rate over a hundred trials in a space the team didn’t arrange. Almost nobody publishes that, and the silence is the finding, the same way it is with benchmark reporting elsewhere.

The Index’s twelve takeaways are a useful counterweight to demo season generally, because they place the physical results beside the software ones instead of reporting each in its own press cycle.

One more thing worth asking is what happens when the robot fails. A system that stops safely is deployable at a much lower success rate than one that keeps going, and almost no demo covers this.

The number to watch over the next year is the household figure. Simulation scores have room to climb without meaning anything; 12% moving to 40% would be the first real evidence the gap is closing.

Get the daily rundown

One email each weekday with the AI news that matters, every claim linked to its primary source.

Free, one email each weekday, unsubscribe in one click. We never sell or share your address.

Rundowns AI Desk

The Rundowns AI desk covers artificial intelligence research, tools, business and policy. Every factual claim we publish links to the primary source it came from, so readers can check it themselves.

Leave a Reply

Your email address will not be published. Required fields are marked *