AI model pricing compared: same benchmark, five times the bill
If two models score the same on a capability index and one costs five times more per token, the interesting question isn’t which is better. It’s what the expensive one is charging for.
That situation is live right now. Grok 4.6 landed at 61 on the Artificial Analysis Intelligence Index, level with GPT-5.6 Sol Max. The price gap between them is not level at all.
| Model | Input / 1M | Output / 1M | AA Index |
|---|---|---|---|
| Grok 4.6 | $2 | $6 | 61 |
| GPT-5.6 Sol | $5 | $30 | 61 (Sol Max) |
| Muse Glimmer (30B, open) | self-hosted. Hardware only | not indexed | |
Why output tokens decide AI model pricing
Output tokens are where this bites. Input is what you send; output is what the model generates, and it’s typically priced three to five times higher because generating is the expensive part.
So a five-fold difference on output isn’t a rounding error in your bill. It’s the difference between an agent you can afford to let run in a loop and one you can’t.
Work it through. An agent making 50 calls per task, generating 2,000 output tokens each, burns 100,000 output tokens per task.
At $6 per million that’s 60 cents. At $30 it’s $3. Run a thousand of those a day and you’re choosing between $600 and $3,000.
Same task, same index score, five times the invoice.
Nobody picks a model on the index once the agent is running. They pick on what the loop costs.
What the index tie conceals about AI model pricing
But the index tie is doing a lot of concealing, and this is where a simple price comparison gets you into trouble. A composite score averages across very different tasks.
Coverage of the release broke the components out. On Harvey LAB, a legal reasoning evaluation, Grok 4.6 scored 15.8% against 2.5% for GPT-5.6 Sol Max and 11.3% for Fable 5 Max. Those three models are nowhere near each other on that task, despite what the headline number suggests. The published figures only make sense read component by component.
Which means the honest version of “cheaper and just as good” is narrower: cheaper, and comparable on average, with real differences underneath that may or may not touch your workload.
The self-hosted third column
Then there’s the third column, which changes the arithmetic entirely. Meta’s Muse Glimmer is a 30-billion-parameter open-weight model under Apache 2.0, sized to run on consumer hardware.
Self-hosting turns a per-token cost into a fixed hardware cost. That’s a terrible deal at low volume and a very good one at high volume, and the crossover point is the only number that matters when you’re deciding.
Roughly: if you’re spending more per month on tokens than a suitable GPU costs to rent, self-hosting starts making sense. Provided you can absorb the operational burden, which most small teams underestimate.
Meta’s own framing leans that way too. TechCrunch read Glimmer as a bet on models running on your own device rather than behind an API.
What we verified, and what we could not
One caveat on all of this, stated plainly. We verified Grok and GPT-5.6 Sol pricing against published Artificial Analysis figures. We could not independently verify current list pricing for Claude Opus 5 or Fable 5, which sit above both on the same index, so they’re absent from the table rather than estimated.
Price gaps within a provider matter as much as gaps between them, and our comparison of flagship against entry tier covers how often the cheap model is genuinely sufficient.
The takeaway isn’t that one model wins. It’s that the index and the invoice have come apart, and only one of them shows up on your card at the end of the month. Our piece on benchmark scores covers why the index deserves less trust than it gets.
Get the daily rundown
One email each weekday with the AI news that matters, every claim linked to its primary source.
Free, one email each weekday, unsubscribe in one click. We never sell or share your address.
