Generative AI platforms compared: API, marketplace or open weights
Search for generative AI platforms and you get a list that puts ChatGPT next to Amazon Bedrock next to Hugging Face. They read as competitors for the same money. They don’t. Four different products share the word, and choosing the wrong category costs more than choosing the wrong vendor inside one. You’re picking between four things. There’s a consumer assistant sold per seat, and a first-party model API sold per token. There’s a cloud marketplace that resells other labs’ models on an invoice you already have. And there’s an open-weight hub where nobody meters your tokens, because you’re renting the hardware instead.
What follows prices all four on one workload, from vendor pages we opened on 18 August 2026. Every rate is quoted from the seller’s own page, and the monthly totals are our arithmetic on those rates. The short answer is that the API tier is where the price competition actually happens. The cloud marketplaces charge roughly what the labs charge, but they run a model generation or two behind. The open-weight route moves your bill from tokens to GPUs and legal review.
Four products share the word “platform”
The categories differ by what you’re actually buying, not by which company owns them. Most of the big names sell in two or three of these categories at once, which is why the comparison articles blur them together.
| Category | Examples | What you buy | Billing unit | Who it suits |
|---|---|---|---|---|
| Consumer assistant | ChatGPT, Claude apps, Gemini app | A chat window and a subscription | Per seat, per month | People doing work by hand |
| First-party model API | OpenAI Platform, Claude Platform, Gemini API, xAI, Mistral | The newest models, directly from the lab | Per million tokens | Anyone building a product |
| Cloud marketplace | Amazon Bedrock, Microsoft Foundry, Vertex AI | Many labs’ models on one existing cloud contract | Per million tokens, or per VM hour | Teams whose procurement is already done |
| Open-weight hub or self-host | Hugging Face, Mistral open weights, Llama | Weights you can run yourself | Compute time, plus licence terms | Data-residency and high-volume cases |
The seat tier is the one people mean when they say generative AI tool, and it’s priced on a completely different logic. Anthropic’s pricing page lists Pro at “$20 if billed monthly” and Max at “From $100 Per month”. A Team standard seat is “$25 if billed monthly” for teams of 2 to 150, and a premium seat is “$125 if billed monthly”. That $25 seat buys roughly 12.5 million input tokens a month at the Claude Sonnet 5 API rate, which is our arithmetic and not a figure the vendor publishes. The catch is that a seat is a person, so none of that capacity reaches your product.
First-party APIs are where the newest model lands first
If you’re building rather than typing, the API tier is the default, and the prices are public. OpenAI’s pricing documentation lists gpt-5.6-sol at $5.00 per million input tokens and $30.00 per million output, gpt-5.6-terra at $2.00 and $12.00, and gpt-5.6-luna at $0.20 and $1.20. Past the long-context threshold the input price doubles and the output price rises by half, so sol becomes $10.00 and $45.00.
Anthropic’s pricing page is flatter. Claude Opus 5 sits at $5 per million input tokens and $25 output. Sonnet 5 is $2 and $10, and Haiku 4.5 is $1 and $5. The page also states that the full 1M token context window is included at standard pricing for Claude 4.6 and later, so “a 900k-token request is billed at the same per-token rate as a 9k-token request”. The same page records that Sonnet 5’s introductory $2 and $10 rate is now the standard price. An increase to $3 and $15 had been scheduled for 1 September 2026, and the page says it “will not occur”.
Google prices lower still, with a deadline attached. The Gemini API pricing page puts Gemini 3.7 Flash at “$0.75 through December 31, 2026” for input and “$3.75 through December 31, 2026” for output, rising to $1.50 and $7.50 on 1 January 2027. That’s a doubling written into the page a reader can see today, which is more honest than most and worth budgeting against. Elsewhere, xAI’s model documentation puts grok-4.6 at $2.00 and $6.00 below 200k tokens, doubling to $4.00 and $12.00 at or above it. Mistral’s pricing page says “Mistral Large costs $0.5 /M tokens in and $1.5 /M tokens out”.
A price per token is only comparable if a token means the same thing at both vendors, and it doesn’t.
Rundowns AI
That’s not a rhetorical point. Anthropic’s own pricing page carries the warning in a footnote, and it undercuts every league table of per-token prices you’ve read, including ours.
Claude 4.7 and later models and Claude Mythos Preview use a newer tokenizer that contributes to their improved performance on a wide range of tasks. This tokenizer produces approximately 30% more tokens for the same text.
Anthropic, Claude Platform pricing documentation, read 18 August 2026
So a headline rate that looks 20% cheaper can lose on the invoice, because the cheaper vendor counts your document into more tokens. We couldn’t verify the cross-vendor ratio, because we found no shared corpus with token counts published by each lab. If that gap matters to your budget, the only reliable test is running your own text through each tokenizer and comparing the counts. That’s the same discipline we applied in our frontier model pricing comparison.
The cloud marketplaces sell procurement, not capability
Bedrock and Foundry exist because enterprise buying is slow. If your company has already signed with AWS or Microsoft, adding a model to that contract takes a day and adding a new vendor takes a quarter. That’s the product, and it’s a real one.
The Amazon Bedrock pricing page lists models from AI21 Labs, Amazon, Anthropic, Cohere, DeepSeek, Google, Luma AI, Meta, MiniMax AI, Mistral AI, Moonshot AI, NVIDIA, OpenAI, Qwen, Stability AI, TwelveLabs, Writer, xAI and Z AI. It sells them across tiers. Batch runs at “a 50% lower price compared to on-demand inference pricing”, Flex at “a 50% discount to Standard tier pricing”, and Priority at “a 75% premium to Standard tier pricing”. Microsoft’s Foundry Models overview claims a wider catalogue still, “over 10,000 models, with approximately 50 new models published each month”.
Breadth isn’t the same as parity, though, and the Foundry documentation is unusually clear about it. The catalogue splits in two. Models sold by Azure carry enterprise SLAs, bill through Azure meters, and “Microsoft provides support”. Models from partners and community are “billed through Azure Marketplace”, with “support and maintenance managed by the respective providers”. Anthropic’s Claude family sits in the second group, and the page directs Claude questions to Microsoft Support through a dedicated link rather than to the standard channel.
The freshness gap is the part nobody advertises, and you can measure it from two vendor pages. Bedrock’s xAI section lists one model, Grok 4.3, at $1.25 per million input tokens, $0.20 cached and $2.50 output, in three US regions. xAI’s own documentation lists grok-4.6 and grok-4.5 above it. The price matches for the model they share, so what you’re giving up is not margin, it’s a generation of capability. Bedrock’s Google section tells the same story from the other direction: it carries Gemma 3 at 4B, 12B and 27B, not Gemini.
Bedrock’s Anthropic table shows the other half of the lag, the tail that never gets cleaned up. It still lists Claude 3.5 Sonnet under “extended access” at $6.00 per million input tokens and $30.00 output. Claude Sonnet 5 is $2 and $10 first-party, so a team that standardised on the marketplace listing and stopped looking is paying three times the current rate for an older model. That’s our comparison of the two pages, not a claim either vendor makes.
Open weights move the bill from tokens to hardware and lawyers
The fourth category is the one that looks free and isn’t. Mistral’s pricing FAQ answers the self-hosting question in five words: “Yes, you can self-host our models anywhere”. The sentence after it is the one that matters. It says “Open-weight models (e.g., Mistral 7B) are Apache 2.0 licensed for research/individual use; while commercial deployments require a Mistral license with separate terms for derivatives and production use”. So the weights are open and the production use is negotiated.
Meta’s terms work differently and are stricter in one specific place. The Llama 4 Community License requires a separate licence from Meta if, on the release date, your products had “greater than 700 million monthly active users in the preceding calendar month”. It also requires that you “prominently display ‘Built with Llama'”, that derivative model names begin with “Llama”, and that distributed copies retain the copyright notice. Almost nobody hits 700 million users, which means the threshold is a competitor clause rather than a constraint on you, but the attribution rules apply to everyone.
Hugging Face sits between hosting and hub, and its billing page is worth reading before you assume it’s the cheap option. Inference Providers routes to more than 200 models with “no markup from Hugging Face”. Monthly credits run to $0.10 for free users, $2.00 for PRO, and $2.00 per seat for Team or Enterprise organisations. Those numbers are experiment money, not production money. The hosted option, hf-inference, bills compute time rather than tokens: the page’s worked example is a FLUX.1-dev request taking 10 seconds on a machine at $0.00012 per second, billed at $0.0012. As of July 2025 the same page says hf-inference “focuses mostly on CPU inference”, naming embeddings, text-ranking, classification and older models like BERT and GPT-2.
That last line is the honest limit of the open route as a drop-in. Running a frontier-class open model yourself means renting GPUs by the hour whether or not requests arrive, which is a fixed cost against a variable workload. We worked that break-even through in hosted API versus self-hosted models, and the shape hasn’t changed: utilisation decides it, and most teams overestimate theirs.
The same job, priced across six platforms
Capability comparisons drift out of date in weeks, so we priced a fixed task instead. The workload is a document summarising feature at moderate scale: 5 million input tokens and 1 million output tokens in a month, all requests under any vendor’s long-context threshold. Every rate below is the standard on-demand price from the pages linked above, read on 18 August 2026, and every total is our multiplication.
| Model and platform | Input per 1M | Output per 1M | Monthly total | Same job, batch tier |
|---|---|---|---|---|
| gpt-5.6-terra, OpenAI Platform | $2.00 | $12.00 | $22.00 | $11.00 |
| Claude Sonnet 5, Claude Platform | $2.00 | $10.00 | $20.00 | $10.00 |
| grok-4.6, xAI | $2.00 | $6.00 | $16.00 | Not published on the page we read |
| Claude Haiku 4.5, Claude Platform | $1.00 | $5.00 | $10.00 | $5.00 |
| Gemini 3.7 Flash, Gemini API | $0.75 | $3.75 | $7.50 | $3.75 |
| Mistral Large, Mistral | $0.50 | $1.50 | $4.00 | $2.00 |
| gpt-5.6-luna, OpenAI Platform | $0.20 | $1.20 | $2.20 | $1.10 |
The spread from top to bottom is ten times, and none of it is explained by which platform you’re on. It’s explained by which tier of model you picked inside it. The three mid-tier models land within $16.00 to $22.00 of each other, while OpenAI’s own range in that table runs from $2.20 to $22.00. So the tier decision moves the invoice further than the vendor decision does, which is what we found in flagship versus entry-tier models.
What the token price leaves off your invoice
Three costs sit outside the per-token rate, and they’re the ones that surprise people at the end of a quarter.
Grounded search is billed separately everywhere it exists. Google charges “$14 per 1,000 requests” for Grounding with Google Search after 5,000 free requests a month shared across the Gemini 3.x models, and charges the same for Grounding with Google Maps. Anthropic’s page prices web search at “$10 per 1,000 searches, plus standard token costs for search-generated content”. A retrieval-heavy product can spend more on searches than on generation, and neither figure appears in a per-token comparison.
Free tiers charge you in data instead of dollars. The Gemini pricing page states plainly that on the free tier, content is “Used to improve our products: Yes”, and that on the paid tier the answer is “No”. That’s a fair trade for a prototype and a bad one for customer text, and it’s the sort of clause we picked apart in what a free tier really costs you.
Then there are the modifiers that stack. Anthropic charges 1.1x on all token categories when you pin inference to US-only geography. OpenAI’s page says regional processing endpoints “are charged a 10% uplift” for eligible models released on or after 5 March 2026. Both vendors also price cache writes above base input and cache reads far below it, so a workload with a large fixed prompt behaves nothing like the table above. Prompt caching is the single biggest lever, and it isn’t visible in any headline rate.
Which category fits, and which one loses
On this evidence, the cloud marketplaces lose the comparison on the thing they’re least willing to discuss. Bedrock and Foundry are excellent at making a purchase happen inside an existing contract. They’re also a generation behind on models, split across two support regimes, and carrying stale listings at three times the current first-party rate. If your constraint is procurement, that trade is worth making. If your constraint is capability, you’re paying a real price for an easier invoice.
Hugging Face’s own hf-inference loses a narrower contest, and it says so itself. Its documentation describes a service now focused mostly on CPU inference for embeddings and smaller models. That’s a useful product and it isn’t a place to serve a chat feature, though the routing layer to other providers is a genuinely good way to compare them on one bill.
The seat subscriptions win for individual work and lose the moment the output needs to reach a customer through software. The first-party APIs win on price and on access to whatever shipped last week, and they lose if your company can’t onboard a new vendor. Open weights win on data residency and on very high steady volume. They lose on everything else, because the licence review and the idle GPU are costs the token price never shows.
Two things would change this reading. If the marketplaces close the launch-day gap, the procurement argument becomes close to free, and the first-party APIs lose most of their advantage for enterprise buyers. And if a lab publishes token counts for a shared corpus, the per-token tables become comparable for the first time, which would settle arguments that currently run on assertion. Until then, the useful move when you read generative AI news about a price cut is to check which tokenizer counts it, and which tier of model it applies to.
Get the daily rundown
One email each weekday with the AI news that matters, every claim linked to its primary source.
Free, one email each weekday, unsubscribe in one click. We never sell or share your address.
