Guides

AI cost per employee: $11.95 at the median, $7,400 at the top 1%

If you’re trying to budget for AI, the honest answer starts with one number. In July 2026 the median American business in Ramp’s spend data spent $11.95 per employee on AI. The top 10% spent $650 per employee that month, and the top 1% spent a median of $7,400. So the median firm’s AI cost is roughly a coffee per head, and the heaviest spenders are running about 620 times that, by our arithmetic on Ramp’s own figures.

That spread is the whole problem with the question. There isn’t an “AI cost” the way there’s a payroll tax rate, because two lines are being added together and they behave nothing alike. Seat licences are fixed, published and boring, so you can plan them a year out. Token usage is metered instead, and it tracks how much work you push through rather than how many people you employ. This piece prices both from the vendors’ own pages, adds up one company’s month, and then says which parts of the bill we couldn’t verify.

The median firm spends about $12 per employee, and the average is meaningless

Ramp runs a spend management product, so its index is built from what businesses actually charged, not from what they told a surveyor. That’s the strength of it, because a card charge happened whether or not anyone remembers approving it. But the weakness is stated on the page: the August 2026 sample “skews slightly more tech-y than our typical AI Index sample”, which pushes every figure up rather than down.

The token side makes the distribution obvious. Across April 2026, Ramp put the median monthly token spend at $2,246 and the average at $140,842. The average is about 63 times the median on our calculation, which is what a long tail does to a mean. So quoting the average at a budget meeting describes nobody, though it’s the figure vendors reach for because it’s the bigger one.

Percentile of businessesMonthly token spend, April 2026
Median$2,246
75th (top 25%)$14,843
90th (top 10%)$73,030
95th (top 5%)$211,409
99th (top 1%)$831,338
Average$140,842
Token spend only, excluding seat licences. Source: Ramp Token Spend Management data, April 2026.

Ramp also breaks per-employee-per-month spend by how many models a company runs, and that’s the most useful cut in the dataset. Companies using four to ten models sat at a $28 median. Eleven to twenty-five models took it to $130. Twenty-six or more reached $442. The count of models is a decent proxy for how many teams are building rather than just chatting, which is why it moves the number so hard. So the question isn’t what AI costs, but whether you’re buying assistance or shipping product.

Seat licences are the half of the bill you can predict

For most businesses the AI cost is seats, and seats are published. Microsoft’s Copilot page currently reads “Originally starting from $21.00 now starting from $18.00 user/month, paid yearly”, with the monthly-billed rate at $25.20. The discount window is narrow and the page says so: the offer runs between 1 July 2026 and 30 September 2026, annual commitment required, and the promotional price applies to the first year only. There’s a prerequisite too, because “a separate license for a qualifying Microsoft 365 plan is required to purchase Microsoft 365 Copilot”.

Anthropic’s team pricing is flatter, though the structure underneath it is more explicit. A standard Team seat is $20 per month billed annually or $25 billed monthly, for teams of 2 to 150. A Premium seat costs $100 annually or $125 monthly and carries five times the standard usage. Enterprise is quoted as “Seat price + usage at API rates” against a $20 per seat minimum, which is the clearest statement any vendor makes that the seat and the meter are separate bills.

SeatAnnual commitmentMonthly billingCondition
Microsoft 365 Copilot add-on$18.00$25.20Needs a qualifying M365 licence; promo price ends 30 Sep 2026
Claude Team, standard$20$252 to 150 seats
Claude Team, premium$100$1255x the standard usage allowance
Claude Pro, individual$17$20$200 billed up front on the annual plan
Google Workspace Business Standard£11.80Not shownUK list price, Gemini included, 300-user cap
Per user per month, as published on each vendor’s pricing page in August 2026. Google’s page served us pounds, not dollars.

Google’s row is the interesting one, because there’s no AI line on it at all. Workspace Business Standard is listed at £11.80 per user per month on the annual rate, and the description bundles in the “Gemini AI assistant in Gmail, Docs, Meet, and more”. Starter sits at £5.90 and Plus at £18.40, all capped at 300 users. The AI is inside the productivity licence you were paying for anyway, which is a real answer to “what does AI cost” and an awkward one for anyone trying to itemise it.

We couldn’t price ChatGPT seats the same way. Both openai.com and the OpenAI help centre returned HTTP 403 to our fetches, so no ChatGPT Business figure appears above. Third-party write-ups quote one, and we’ve left it out rather than cite a price we didn’t read at source. The API documentation did open, which is where the next section comes from. If the seat-versus-meter decision is the one in front of you, we’ve worked it through separately in buy a seat or run a GPU.

Token prices are low, and the model you pick moves them 200 times

Published API rates are the part of the bill people fear most and understand least. OpenAI’s pricing docs put gpt-5.6-sol at $5.00 per million input tokens and $30.00 per million output on standard, short-context requests. The mid tier, gpt-5.6-terra, runs $2.00 and $12.00. The small one, gpt-5.6-luna, costs $0.20 and $1.20. Anthropic’s pricing page lines up close: Opus 5 at $5 and $25, Sonnet 5 at $2 and $10, Haiku 4.5 at $1 and $5.

ModelInput, per 1M tokensOutput, per 1M tokens
gpt-5.6-sol$5.00$30.00
Claude Opus 5$5$25
gpt-5.6-terra$2.00$12.00
Claude Sonnet 5$2$10
Claude Haiku 4.5$1$5
gpt-5.6-luna$0.20$1.20
Llama 3.3 on Bedrock$0.15$0.15
Standard, short-context rates as published August 2026. Sources: OpenAI pricing docs, Anthropic pricing docs, AWS Bedrock pricing.

Read the top and bottom rows together and the range is the story. Output on gpt-5.6-sol costs 200 times what output costs on Llama 3.3 through Bedrock, on our division of the two published rates. Two discounts compress it further. The Batch API takes 50% off both directions at both vendors, and Anthropic’s cache reads are priced at 0.1x the base input rate, so repeated context costs a tenth of fresh context. We compared frontier models on identical work in same benchmark score, five times the bill, and the gap there was in capability, not in list price.

What that means in practice shows up in Anthropic’s own worked example. Processing 10,000 support tickets at roughly 3,700 tokens per conversation on Haiku 4.5 comes to about $37.00 total. That’s the number to hold onto, because it reframes the whole exercise. Ten thousand automated conversations cost less than two Copilot seats.

One 60-person company’s month, added up

Take a services firm with 60 staff. Twenty of them get Copilot on the annual rate. Eight people on the delivery team get Claude Team standard seats. The support inbox runs 4,000 tickets a month through Haiku 4.5, and the internal research tool fires 3,000 web searches. Every rate below came off a page we opened, and the totals are ours, so the arithmetic is checkable line by line.

Line itemWorkingMonthly cost
20 Copilot seats20 × $18.00$360.00
8 Claude Team seats8 × $20$160.00
4,000 support tickets4,000 × ($37.00 ÷ 10,000)$14.80
3,000 web searches3 × $10 per 1,000$30.00
Total60 staff$564.80, or $9.41 per employee
Our calculation, from published August 2026 rates. Annual-commitment seat pricing assumed throughout.

At $9.41 per employee, that company sits just under Ramp’s $11.95 median, and 92% of its bill is seats. The metered half comes to $44.80. Notice what happened, because it inverts the usual worry: the metered line barely registers, since the volume being metered is small. The crossover sits at roughly 140,000 tickets a month, because that’s where $0.0037 a ticket catches $520 of seats, and that’s the threshold worth knowing. It’s the one we walked through end to end in what it costs to run an AI feature.

The metered half of a normal company’s AI bill is smaller than the seat licences for two people. Until it suddenly isn’t.

Rundowns AI, from published August 2026 rates

The surcharges that don’t appear in the headline price

Both vendors publish modifiers that multiply the table above, and they stack. OpenAI charges a 10% uplift on regional processing endpoints for models released on or after 5 March 2026. Anthropic applies a 1.1x multiplier when inference_geo is pinned to “us” on Claude 4.6 and later, and Bedrock and Google Cloud add “a 10% premium over global endpoints” for regional routing. Which means data residency is a governance decision that lands on the invoice, even though nobody in procurement filed it as a cost.

Speed costs more than residency. OpenAI’s Fast mode takes gpt-5.6-sol to $10.00 input and $60.00 output, double the standard rate. Anthropic’s Fast mode prices Opus 5 at $10 input and $50 output, also double. Long context does the same job quietly: on OpenAI’s table, sol above the short-context threshold moves to $10.00 and $45.00. So feeding a model more history changes the per-token rate itself, not just the token count, which is why context bloat compounds twice.

Then there are the meters that aren’t tokens at all. Anthropic charges $10 per 1,000 web searches, $0.08 per session-hour for Managed Agents runtime, and $0.05 per hour per container for code execution beyond 1,550 free hours a month. Region matters too, because Bedrock’s Mistral rates in Europe (London) run roughly 24% above the US regions, and DeepSeek’s priority tier carries a 75% premium. Paying monthly instead of annually costs 40% more on Copilot, comparing the $25.20 and $18.00 figures Microsoft publishes.

The subtlest one isn’t a price at all. Anthropic’s pricing page notes that Claude 4.7 and later use a newer tokenizer that “produces approximately 30% more tokens for the same text”. Nothing on the rate card moved. The same document you sent last quarter simply counts higher now, so a headline price cut and a tokenizer change can cancel each other out without either showing up in a procurement comparison.

Claude 4.7 and later models and Claude Mythos Preview use a newer tokenizer that contributes to their improved performance on a wide range of tasks. This tokenizer produces approximately 30% more tokens for the same text.

Anthropic pricing documentation, August 2026

Where low cost AI is genuinely low cost

Open-weight models on a managed endpoint are the cheapest honest option, and the gap isn’t marginal. AWS Bedrock lists Llama 3.3 at $0.15 per million tokens in both directions in the US regions. Qwen3 Next 80B runs $0.15 input and $1.20 output. Mistral Large 3 costs $0.50 and $1.50. DeepSeek v3.2 sits at $0.62 and $1.85, with a flex tier at half that. Batch inference takes another 50% off.

The catch arrives when you try to reserve capacity instead of paying per token. Bedrock’s Provisioned Throughput for Llama 2 is priced at $21.18 per hour per model unit on a one-month commitment, or $13.08 on six months. Multiply the one-month rate across a 730-hour month and you’re at roughly $15,500 before a single request, on our arithmetic. Reserved capacity buys latency guarantees, but it buys them at a floor most businesses never reach on usage alone.

Free tiers sit at the other end and carry their own bill, mostly in data terms and rate limits rather than cash, which we picked apart in what each free tier really costs you. Between the two extremes, the cheap-model route works when the task is narrow and the output is short. It stops working when a weaker model needs three attempts to do what a stronger one does once, because three cheap calls that fail still cost more than one that lands.

What the spend data can’t tell you

Every per-employee figure here has a denominator problem, and the official statistics show how big it is. The Census Bureau’s Business Trends and Outlook Survey found overall AI usage hovering between 17% and 20% across 14 December 2025 to 3 May 2026. As of 3 May 2026, 37% of firms with at least 250 employees reported using AI, against 32% for firms with 100 to 249 and under 20% for the smallest. By sector, information reached 39.7% and finance and insurance 33.9%.

Ask a different survey and you get a different world. A Federal Reserve FEDS Note by Jeffrey S. Allen, published 3 April 2026, lines up three of them: BTOS at 18% of firms in December 2025, the Real-Time Population Survey at 41% of the labour force for generative AI in November 2025, and the Survey of Business Uncertainty at 78% on an employment-weighted basis. Allen attributes the spread to sampling, to unit of analysis, and to question wording, noting that BTOS respondents are “conditioned to understand the survey questions as referring to more material or extensive AI usage”.

That matters for budgeting because a per-employee average silently assumes every employee is a user. If a fifth of your headcount touches the tools, your real cost per user is five times the figure your finance team quotes. The same confusion runs through the returns side of the argument, which we set out in the benefits of AI, separated from the claims.

Two things would change the numbers above quickly. Microsoft’s Copilot discount expires on 30 September 2026, and the page is explicit that promotional pricing applies to the first year only, which means the $21.00 still printed on that page is what a second year currently lists at. And tokenizer revisions of the kind Anthropic documented move effective cost without moving list price, which no rate-card comparison catches. So anyone quoting a single “AI cost” figure has either picked a percentile or skipped the meters, and it’s worth asking which.

Get the daily rundown

One email each weekday with the AI news that matters, every claim linked to its primary source.

Free, one email each weekday, unsubscribe in one click. We never sell or share your address.

Leave a Reply

Your email address will not be published. Required fields are marked *