Comparisons

LLM pricing tracked: output runs $0.25 to $450 per million tokens

Across the ten vendor pricing pages we read on 5 October 2026, published output prices run from $0.25 to $450 per million tokens. Both ends of that range sit inside OpenAI’s own table. Every figure below came from the page the vendor publishes it on, read on that date, and the link under each table goes to the page we read, so you don’t have to take our word for any cell.

That 1,800-fold spread is the first thing worth understanding about LLM pricing, because it isn’t a spread between cheap models and expensive ones. It’s mostly a spread between service tiers, context lengths and cache states on the same model. GPT-6 Astra is listed at four output prices depending on how fast you want the answer, and each of those four has a second, higher rate once your prompt gets long. That’s eight published output prices for one model. A quoted number without those qualifiers is close to meaningless.

This page is a list-price tracker, re-checked monthly. It’s not a benchmark comparison: we did that separately when two models tied on a capability index with five times the bill between them. Here the only question is what the vendors say they charge.

What the frontier models cost per million tokens today

These are standard-tier, short-context, undiscounted list prices in US dollars. Cached input is the rate for reading a prompt the provider has already processed, which is the cheapest token in every table and the one we come back to below.

ModelInput / 1MCached input / 1MOutput / 1M
GPT-6 Astra (OpenAI)$10.00$1.00$50.00
Claude Fable 5.1 (Anthropic)$10.00$0.25$50.00
Claude Opus 5.5$4.00$0.20$20.00
Gemini 3.1 Pro Preview (Google)$2.00$0.20$12.00
GPT-6.1 Sol$2.00$0.10$10.00
Claude Sonnet 5.5$2.00$0.20$10.00
Grok 4.7 (xAI)$2.00$0.50$6.00
GLM-5.3 (Z.ai)$1.40$0.26$4.40
DeepSeek-V4-Pro, peak$1.32$0.044$3.96
Claude Haiku 4.5$1.00$0.10$5.00
Gemini 3.8 Flash, promotional$0.75$0.075$3.75
DeepSeek-flash, peak$0.30$0.006$1.20
GLM-5.3-Flash$0.15$0.03$0.50
GPT-6 Luna$0.10$0.01$0.50
List prices per 1M tokens, USD, read 5 October 2026 from OpenAI, Anthropic, Google, xAI, Z.ai and DeepSeek. Gemini and Claude cached-input rates are from Vertex AI and the Claude docs. DeepSeek rows show peak rates; off-peak is half.

Two things stand out. The top of the table is a tie: OpenAI and Anthropic both price their top-priced model at exactly $10 in and $50 out, to the cent. Below that, the ordering is not what the marketing suggests. Anthropic’s Opus 5.5 is described on its own pricing page as the “daily driver for agentic coding and enterprise work” rather than the frontier model, and it costs $4 and $20. Google’s top-priced Gemini text model, 3.1 Pro Preview, undercuts it at $2 and $12, and Grok 4.7 undercuts that again on output at $6.

The Chinese labs sit a further step down. Z.ai lists GLM-5.3 at $1.40 and $4.40, which puts Claude Fable 5.1’s output at roughly eleven times GLM-5.3’s. That gap is why we bothered to check whether the cheaper model actually loses on the work. DeepSeek goes lower still, and its cheapest input token, a cache hit on deepseek-flash in off-peak hours, is listed at $0.003 per million.

The sticker price is one of five prices for the same model

OpenAI’s pricing page carries five service tiers, and the same model appears in all of them. Batch and Flex both run at half of standard. Fast mode, which the page notes was renamed from Priority processing on 30 July 2026, doubles it. Ultrafast multiplies it by six, and on the day we read the page only GPT-6 Astra was listed in that tier at all.

GPT-6 Astra tierInput, short ctxOutput, short ctxInput, long ctxOutput, long ctx
Batch and Flex$5.00$25.00$10.00$37.50
Standard$10.00$50.00$20.00$75.00
Fast$20.00$100.00$40.00$150.00
Ultrafast$60.00$300.00$120.00$450.00
One model, eight prices. Per 1M tokens, USD, from OpenAI’s pricing page, read 5 October 2026. Short context is 272K input tokens or fewer.

So the difference between the cheapest and most expensive way to buy the identical model is 18-fold on output. That’s our arithmetic on their table, and it dwarfs most of the gaps between vendors in the first table. If you’ve read a “GPT-6 Astra costs $50 per million” line anywhere, including our own coverage of the launch, that was the standard short-context rate and nothing else.

Anthropic’s structure is simpler but has the same shape. The Batch API takes 50% off input and output, so Fable 5.1 drops to $5 and $25. Fast mode on Opus 5.5 costs $8 and $40, exactly double standard, and the docs state it can’t be combined with the Batch API. xAI is the outlier here: its own model page lists Grok 4.7 with “Batch API: Not supported”, so there’s no discounted asynchronous tier to reach for.

Long context is a cliff, not a slope

Cache discounts shrink what a long prompt costs to re-read, but they don’t touch what it costs the first time. Every major vendor now charges more for long prompts, and they all draw the line in a different place. OpenAI splits at 272K input tokens, above which input doubles and output rises by half. Google’s Gemini 3.1 Pro Preview splits at 200K, going from $2 and $12 to $4 and $18. Grok 4.7 also splits at 200K, with both input and output doubling.

The xAI version is the harshest, and its Grok 4.7 documentation is admirably blunt about why. The rule isn’t that tokens past the threshold cost more. It’s that crossing the threshold reprices the whole request.

Requests whose prompt reaches 200k tokens are billed at the higher rate for all tokens in the request.

xAI, Grok 4.7 model documentation, read 5 October 2026

Work that through. A 199,999-token prompt to Grok 4.7 bills input at $2 per million, so about 40 cents. One token more and the same prompt bills at $4 per million, so about 80 cents. The marginal token costs you roughly 40 cents, which is our calculation from the two rates xAI publishes. That’s the kind of edge that makes a million-token context window a budgeting problem rather than a free upgrade.

Cache reads are where the discount actually moved this year

Prompt caching lets a provider charge less for input it has already processed, and it’s the line that has fallen fastest. Anthropic’s docs set the default cache hit at 10% of the standard input price, then carve out two exceptions: 5% on Opus 5.5, and 2.5% on Fable 5.1 and Mythos 5.1. In dollars that’s $0.20 and $0.25 per million. So Fable 5.1 charges only 25% more than Opus 5.5 to re-read a cached prompt, even though its fresh input costs 2.5 times as much, which is a different ranking from the one the headline rates give you.

Writing to the cache costs extra, and the docs give the break-even directly: 1.25 times base input for a five-minute cache, 2 times for an hour, which pays off after one read at five minutes or two reads at an hour. That matters because the multipliers compound. Anthropic states that caching stacks with the Batch API discount and with data residency, and OpenRouter’s models endpoint lists Opus 5.5 batch cache reads at $0.10 per million, which is what the stacked rule produces. So an Opus 5.5 input token can land at a fortieth of its standard rate, but only because two modifiers multiplied on it at once.

Not every vendor has moved as far. Grok 4.7’s cached input is $0.50 against $2.00 fresh, a 25% rate rather than 10%, and Z.ai’s GLM-5.3 cache hit at $0.26 against $1.40 works out near 19%. DeepSeek is at the other extreme: a cache hit on deepseek-flash at peak is $0.006 against $0.30 for a miss, which is 2% of the standard rate.

A 20% price cut that can cost you 4% more

Here’s the footnote that undermines every price table on the internet, including the two above. Anthropic’s pricing docs carry it under “Additional notes”, and it has nothing to do with dollars.

Claude 4.7 and later models and Claude Mythos Preview use a newer tokenizer that contributes to their improved performance on a wide range of tasks. This tokenizer produces approximately 30% more tokens for the same text.

Anthropic pricing documentation, read 5 October 2026

A token is the unit a model chops text into before it reads it, and the chopping rule is the vendor’s choice, not a law of nature. We’ve written up what tokenisation does to a model’s view of text; what it does to an invoice is simpler. If the same English sentence becomes 30% more tokens, a 30% price cut buys you nothing.

Run the two cases from Anthropic’s own table. Opus 4.6 uses the previous tokenizer and lists at $25 per million output tokens. Opus 4.7 lists at the same $25 and uses the newer one, so a given piece of text costs about $32.50 to generate, roughly 30% more at an unchanged sticker price. Opus 5.5 lists at $20, and 1.3 million tokens at that rate is $26. Against Opus 4.6’s $25, that’s a headline cut of 20% landing as a bill about 4% higher on identical text. Those are our calculations from the rates and the 30% figure Anthropic publishes.

The counter-case is real, though, and it’s the stronger argument. Price per token isn’t price per task. A model that reasons more efficiently emits fewer output tokens for the same answer, so it can be cheaper per job while costing more per thousand words. Anthropic made exactly that claim in September, saying Opus 5.5 costs 40% less to run than Opus 5 at default settings, a figure we separated from the 20% sticker cut at the time. Both readings can be true at once, which is precisely why a price table can’t settle the question on its own, and why metering your own workload beats reading ours.

Data residency costs 10% almost everywhere, and the clock matters at DeepSeek

Three of the Western vendors have settled on the same number for keeping inference inside a region. OpenAI’s footnotes say regional-processing and FedRAMP endpoints carry a 10% uplift for models released on or after 5 March 2026. Anthropic charges a 1.1x multiplier on every token category, input, output, cache writes and cache reads, when you pin US-only inference on Claude 4.6 or later. Google’s Vertex table prints both rates side by side: Gemini 3.8 Flash is $0.75 and $3.75 on a global endpoint, $0.825 and $4.125 on a non-global one.

Resellers, by contrast, mostly don’t mark up. We resolved the pricing feed behind the Amazon Bedrock pricing page and found Claude Opus 5.5 at $4 input and $20 output in all 33 regions it lists, which matches Anthropic’s first-party price to the cent. OpenRouter’s models endpoint returns the same list prices for GPT-6 Astra, Fable 5.1, Opus 5.5, Grok 4.7 and Gemini 3.1 Pro Preview. The markup, where it exists, is in the endpoint type rather than the storefront, which also means the hosted-versus-self-hosted decision isn’t about finding a cheaper reseller.

DeepSeek prices on a dimension nobody else in the table uses: the time of day. Its docs define peak hours as 01:00 to 04:00 and 06:00 to 10:00 UTC, Monday to Friday, excluding Chinese public holidays, and every other hour is off-peak at half the rate. Weekends are off-peak in full. For batch work with no deadline, that’s a 50% discount available by scheduling, on top of a model that already lists under $4 per million output tokens.

Then there are the meters that aren’t tokens at all

A per-token table misses a growing share of the bill, because agentic features bill on their own units. That is not a rounding error, and it is not in any of the tables above. Anthropic’s docs list web search at $10 per 1,000 searches on top of token costs, code execution at $0.05 per container-hour after 50 free hours a day per organisation, and Managed Agents at $0.08 per session-hour of active runtime. The docs also note that the Batch API discount doesn’t apply to Managed Agents sessions, so one of the two levers you’d reach for isn’t there.

The others publish their own. Z.ai charges $0.01 per web search use. xAI’s table prices text-to-speech at $15 per million characters, note the unit, and speech-to-text at $0.10 an hour over REST or $0.20 streaming. None of that appears in a per-million-tokens comparison, and for an agent that searches on most turns it can be the larger line.

The changelog: every dated price move we can point at

Vendors rarely announce a price change as a price change, so the dates below come from footnotes and from tables that print a future rate next to the current one. Two of these haven’t happened yet, which is the useful part, because a scheduled increase is the one thing in pricing you can plan around.

DateWhat changed, or is scheduled to
5 March 2026OpenAI: models released on or after this date carry a 10% uplift on regional-processing and FedRAMP endpoints
30 July 2026OpenAI renames Priority processing to Fast mode, the 2x tier
22 September 2026Opus 5.5 arrives at $4 and $20, 20% under Opus 5, with cache reads cut to $0.20; GPT-6 Sol lists at $2 and $10 against GPT-5.6 Sol’s $4 and $20
21 November 2026OpenAI’s footnote guarantees GPT-5.6 Sol promotional pricing only “at least through” this date
31 December 2026Promotional pricing on Gemini 3.6, 3.7 and 3.8 Flash ends
1 January 2027Google’s own table lists those three Flash models at $1.50 and $7.50, double the promotional rate
Compiled from vendor footnotes and forward-dated rate tables, read 5 October 2026. Sources: OpenAI pricing, Vertex AI pricing, Claude pricing docs.

The Google row is the one to sit with. Three consecutive Flash generations are all listed at the same $0.75 and $3.75, and all three are scheduled to move to $1.50 and $7.50 on the same day. So a model priced at about a thirteenth of the Claude and GPT flagships on output has a published doubling with a date on it, and anyone who built a unit-economics model on the promotional number has a deadline rather than a price.

What these numbers can’t tell you

Three limits, stated plainly. These are list prices, and the large buyers aren’t paying them: Anthropic’s docs describe negotiated discounts applied as fewer Consumption Units metered on AWS, and the size of those discounts isn’t published anywhere we could read. Second, the tables are a snapshot of one date, and two of the rows above are models that were first listed in September. Third, and most important, nothing here measures tokens per task, which is the term that actually decides your invoice and the one no vendor page contains.

What would change the conclusion? A vendor publishing cost per completed task on a fixed workload would make this page mostly redundant, and none of them does. Short of that, the number we’d watch is the cache read line. Anthropic’s documented default is 10% of the input price and its top model is now at 2.5%, and for agentic work that replays a long prompt on every turn, that line moves the bill more than the headline rate does. We re-check every page in this piece monthly and date the table when we do.

Get the daily rundown

One email each weekday with the AI news that matters, every claim linked to its primary source.

Free, one email each weekday, unsubscribe in one click. We never sell or share your address.

Leave a Reply

Your email address will not be published. Required fields are marked *