Grok 4.6 matches GPT-5.6 Sol on benchmarks at a fifth the output price
Grok 4.6, from Elon Musk’s AI company xAI, landed this week scoring 61 on the Artificial Analysis Intelligence Index. That’s up from 56 for Grok 4.5, and level with GPT-5.6 Sol Max. The number that matters, though, is the price tag underneath it.
The model costs $2 per million input tokens and $6 per million output, unchanged from the previous release. GPT-5.6 Sol runs $5 and $30 for the same volumes, per Artificial Analysis figures. That’s roughly 60% cheaper on input and five times cheaper on output, for a comparable index score.
| Benchmark | Grok 4.6 |
|---|---|
| AA Intelligence Index | 61 (was 56) |
| GDPVal-AA v2 | 1753 |
| CursorBench v3.2 | 69.9% |
| DeepSWE v1.1 | 65.9% |
| Harvey LAB (Vals) | 15.8% |
| Price per 1M tokens | $2 in / $6 out |
The legal reasoning result is the outlier. On Harvey LAB, Grok 4.6 scored 15.8% against 2.5% for GPT-5.6 Sol Max and 11.3% for Fable 5 Max. That gap is wide enough to say more about task fit than general capability.
It sits behind Claude Opus 5 and Fable 5, which hold the top two spots on the same index. So this isn’t a claim on the frontier. It’s a claim on the price-to-intelligence curve, and on that basis it’s a strong one.
The engineering detail is worth noting. MarkTechPost reports this as a post-training upgrade, not a larger base model.
The foundation was held constant. The work went into a longer supplemental training run, regenerated supervised fine-tuning trajectories, and reinforcement learning in agentic environments. Context extends to 500K tokens.
That’s the same pattern showing up across the industry: gains extracted after pre-training rather than from a bigger run. We covered why that shift is happening in our piece on whether scaling laws have stalled.
Reported weaknesses cluster around terminal use, with knowledge work and legal reasoning the strongest areas, according to a breakdown of the release. Agentic coding scores are solid without leading.
Read benchmark tables from any lab with the usual caution. They’re selected by the company shipping the model, and index scores compress very different capabilities into one figure. The pricing, though, is externally checkable and doesn’t move with interpretation.
For anyone running volume workloads, that’s the whole story. A five-fold difference in output token cost changes which model you can afford to put in a loop. That decision rarely comes down to the last few points of an index score. Comparable coverage reached the same read.
Get the daily rundown
One email each weekday with the AI news that matters, every claim linked to its primary source.
Free, one email each weekday, unsubscribe in one click. We never sell or share your address.
