Models & Research

Anthropic, xAI and Google now sit within 79 Elo points of each other

As of March 2026 six labs sat within 79 Elo points of each other at the top of the Arena ratings: Anthropic on 1,503, xAI on 1,495, Google on 1,494, OpenAI on 1,481, Alibaba on 1,449 and DeepSeek on 1,424.

Stanford’s 2026 AI Index reads that cluster as competitive pressure moving toward cost, reliability and domain-specific performance. That reading is worth taking seriously, because it describes a different market from the one everyone still talks about.

Two of those six publish weights

Alibaba and DeepSeek are in the top tier with open-weight families, which was not true two years ago and undermines the assumption that closed labs hold a durable capability lead. A third of the frontier is now downloadable.

When six labs are within 5% of each other, the question stops being which model is best and becomes which one you can afford to run at volume.

Researchers have started asking the follow-up question directly, testing whether open-weight models compete on financial text comprehension, which is the sort of narrow professional domain where the gap was supposed to persist.

Releases have kept coming through 2026. Meta published a 30B agent model for consumer hardware and Thinking Machines shipped Inkling, both from organisations with the option of staying closed.

What frontier model benchmarks look like at parity

When one lab leadsWhen six are level
Buyers pay for capabilityBuyers pay for throughput
Launches claim a new frontierLaunches claim nine benchmark wins
Releases every several monthsReleases every three weeks
Switching is expensiveSwitching is a base URL
Parity changes the shape of the market more than it changes the models.

The cadence row is observable right now. Google shipped an entry-level model three weeks after its predecessor, which is a schedule set by rivals rather than by research.

Price is where parity shows up hardest. Grok 4.6 matched GPT-5.6 Sol at a fifth the output price, and a five-times spread on equivalent quality is only sustainable while buyers are not paying attention.

The release cadence is the tell

Watch what the labs do rather than what they say, and the parity claim gets easier to believe.

Meta published a 30-billion-parameter model built to run agents on a single consumer GPU, framed by TechCrunch as an open version of its strongest closed system.

Thinking Machines put out open weights within weeks of that, per SiliconANGLE. Companies protect a lead when they have one, and give away weights when the lead is thin enough that distribution is worth more.

The Index’s summary of the year makes the same point from the data side, with capability converging while deployment and cost diverge.

What parity does to the labs

Convergence at the top is comfortable for buyers and difficult for the companies doing the training.

Training costs did not fall because six labs reached the same place. Each of them still paid for its own frontier run, and now competes on price with five others who did the same.

That is the structure of a commodity market arriving on top of a research cost base, which is an uncomfortable combination and the reason our piece on training economics matters more now than a year ago.

Two escapes exist. Differentiate on something other than the base model, which is what the enterprise and agent products are attempting, or get far enough ahead that parity breaks.

Nothing in the current data suggests the second is happening, and the three-week release cadence suggests the first is where the effort is going.

Elo is a weak instrument, and it still tells you this

Arena ratings come from humans picking a preferred response, so they reward answers that read well and say nothing about factual accuracy.

That bias is the same one OpenAI documented in InstructGPT, where a 1.3B tuned model was preferred over 175B GPT-3. Preference measures presentation as much as substance.

None of which changes the clustering. Even a crude instrument showing six labs inside 79 points is evidence that no one holds a decisive lead on the work most people actually send a model, and a crude instrument is what buyers use anyway.

Where the differences that survive actually live

Parity on a general rating does not mean the models are interchangeable, and the places they still diverge are unusually practical.

Long-context behaviour is one. Two models can advertise the same window and behave very differently at 200,000 tokens, since capacity and usable attention are separate properties.

Instruction-following under pressure is another. A model that holds a strict output format across a thousand calls is worth more to a pipeline than one that scores two points higher and drifts.

Then there are the unglamorous ones that decide deployments: rate limits, regional availability, how long a version stays supported, and whether the provider trains on your inputs by default.

None of those appear in an Elo rating, and all of them will change your experience more than 79 points of preference score will.

What frontier model benchmarks mean for your stack

Build for portability. If four vendors can serve your workload, a hard dependency on one of them is a cost you’re choosing rather than one you’re forced into, and abstracting the call site takes an afternoon.

Then compete the providers on your own tasks rather than on published tables, since the differences that survive at parity are the ones specific to your domain, and nobody publishes a benchmark for your domain.

Keep an evaluation set of your own, too. At parity the published tables cannot distinguish between the options, and ten of your own tasks can.

And treat the leader board position as the least durable fact about any of them. At 79 points across six labs, today’s order is next month’s noise, and an architecture that assumes otherwise ages badly.

Watch for consolidation as the counter-signal. If labs start being acquired rather than raising independently, that says the market decided parity was unprofitable, which our piece on who keeps the money sets out.

Get the daily rundown

One email each weekday with the AI news that matters, every claim linked to its primary source.

Free, one email each weekday, unsubscribe in one click. We never sell or share your address.

Rundowns AI Desk

The Rundowns AI desk covers artificial intelligence research, tools, business and policy. Every factual claim we publish links to the primary source it came from, so readers can check it themselves.

Leave a Reply

Your email address will not be published. Required fields are marked *