Comparisons

AI tools for business: $216 a seat, 30% of the tasks finished

If you’re pricing AI tools for business, the number you can get in ten seconds is the seat. Microsoft’s own page lists Microsoft 365 Copilot as an add-on at $18.00 user/month, paid yearly, shown against a struck-through $21.00, and it says plainly that “a separate license for a qualifying Microsoft 365 plan is required”. That’s $216 a seat a year by our arithmetic. The number you can’t get in ten seconds is what the seat finishes. So we went to the benchmarks that hand models actual work product and count what comes back. In a simulated company, the best system completed 30.3% of 175 professional tasks. On the open 220-task subset of a benchmark built from real expert deliverables, the best model matched or beat the human in 47.6% of head-to-head gradings.

Those two numbers do most of the work in this piece. Everything below comes from a page or a paper we opened in August 2026. That’s four vendor pricing pages, two benchmark papers, one live leaderboard, and one research group’s running account of its own randomised trial.

What a seat costs before anybody opens it

Seat pricing is the one part of this market that’s published, dated and unambiguous, so it’s the right place to start. Microsoft sells Copilot two ways: bolted onto a licence you already hold, or bundled. Anthropic’s pricing page puts a Claude Team standard seat at “$20 Per seat / month if billed annually. $25 if billed monthly”, for teams of 2 to 150, with a premium seat at $100 annually and $125 monthly. Its Enterprise answer is different in kind: “$20 per seat per month plus usage billed at API rates”, which means the invoice moves with how much your team actually runs. OpenAI’s help centre gives the same four numbers for ChatGPT Business. For a standard seat, “pricing (USD) is $25 per user per month if billed monthly and $20 per user per month if billed annually”, and a premium seat runs $125 monthly or $100 annually. Two rival vendors have landed on identical prices at both tiers, and both set the floor at two seats.

PlanRate shownBilling term shownAnnualised at that rate
Microsoft 365 Copilot, add-on$18.00 per user / monthPaid yearly (was $21.00)$216.00
Microsoft 365 Copilot, add-on$25.20 per user / monthPaid monthly$302.40
Microsoft 365 Business Standard with Copilot$23.50 per user / monthPaid yearly$282.00
Microsoft 365 Business Premium with Copilot$32.00 per user / monthPaid yearly$384.00
ChatGPT Business, standard seat$20.00 per user / monthBilled annually ($25 monthly)$240.00
ChatGPT Business, premium seat$100.00 per user / monthBilled annually ($125 monthly)$1,200.00
Claude Team, standard seat$20.00 per seat / monthBilled annually ($25 monthly)$240.00
Claude Team, premium seat$100.00 per seat / monthBilled annually ($125 monthly)$1,200.00
Google Workspace Standard£11.80 per user / monthBilled monthly, 16% off on a 1-year commitment£141.60
Google Workspace Plus£18.40 per user / monthBilled monthly, 16% off on a 1-year commitment£220.80
Notion Business£16.50 per member / monthPay monthly, “save up to 20% with yearly”£198.00
Rates as each vendor’s own page served them to us on 26 August 2026, with the annualised column our own multiplication. Google and Notion served UK pricing, so those rows are in pounds and don’t line up against the dollar rows. Google’s page also runs a 50% discount from 9 September to 9 December 2026, limited to the first 20 users for 12 months, and caps Starter, Standard and Plus at 300 users; Microsoft’s add-on page says “For up to 300 users”.

Two things in that table matter more than the headline rates. The first is the monthly-billing penalty. Paying Microsoft month to month for the Copilot add-on costs $25.20 instead of $18.00, which is 40% more for the same product by our calculation, and Anthropic’s gap is 25%. The second is that “AI included” now means very little. Google’s Workspace pricing page lists “Gemini AI assistant in Gmail” on every tier including Starter at £5.90 a user a month, so the AI isn’t the line item any more, the tier is. Notion’s page goes further and splits the bill in two: a £16.50 Business seat, then agents that are “Free to try, then $10 per 1,000 monthly Notion credits”. A seat price with a meter behind it isn’t a seat price. We worked through the three shapes this market sells in, seat, token and rented GPU, in our piece on artificial intelligence platform pricing.

The benchmark built from real deliverables stops at 47.6%

Most benchmarks ask models exam questions. GDPval, published by OpenAI in October 2025, asks them for deliverables. It covers 1,320 tasks across 44 occupations in the top nine sectors by contribution to US GDP, with each task built from work product by a professional averaging 14 years in the field, and a 220-task gold subset released openly. Grading is a blind pairwise comparison by industry experts, which is a much harder thing to game than a multiple-choice score.

The result the paper reports is that “47.6% of deliverables by Claude Opus 4.1 were graded as better than (wins) or as good as (ties) the human deliverable”. That’s the best of the models it tested, which were GPT-4o, o4-mini, o3, GPT-5, Claude Opus 4.1, Gemini 2.5 Pro and Grok 4. The same paper gives the denominator that makes the seat price legible: on the gold subset, a human expert took an average of 404 minutes and cost $361 per task. That’s 6.7 hours at an implied $53.61 an hour, both our calculation. Put the two together and a full year of the cheapest Copilot seat costs about 60% of what one of those single expert tasks costs.

Every price in the table above is published and dated. Every completion rate behind it is partial, contested, and measured on a model that has since been replaced.

Rundowns AI

The catch is buried in the appendix, and it’s the most useful finding in the paper for anyone buying. Win rates, the paper says, “are highest for shorter tasks” and “decline steadily as completion time increases”, with the top band running from zero to two hours. So the tasks these tools handle best are the ones that were cheapest to do anyway, and the expensive multi-day work is exactly where the win rate falls off.

A simulated company handed agents 175 jobs and got 30% back

GDPval measures one deliverable at a time. TheAgentCompany, from researchers at Carnegie Mellon and Duke, measures something closer to a job. It stands up a self-hosted fake software company with its own GitLab, its own chat server, its own project tracker and its own file store, then populates it with simulated colleagues. Agents get 175 tasks drawn from software engineering, project management, data science, admin, HR and finance. The agents browse, write code, run programs and message coworkers, and a programmatic evaluator checks named checkpoints rather than asking a model whether it did well.

In a setting simulating a real workplace, a good portion of simpler tasks could be solved autonomously, but more difficult long-horizon tasks are still beyond the reach of current systems.

TheAgentCompany, arXiv:2412.14161

Gemini 2.5 Pro was the best of the models tested, at 30.3% of tasks completed outright and 39.3% on the score that gives partial credit, using 27.2 steps and $4.20 of tokens per task on average. Claude 3.7 Sonnet came second at 26.3%, and GPT-4o managed 8.6%. That spread is worth holding onto, because it means the gap between the top tool and a merely recent one is larger than the gap between any two seat prices in our table.

The breakdown by platform is where a business reader should stop and look, because it splits the work by the software people spend the day inside.

Platform the task touchesTasksGemini 2.5 Pro successScore with partial credit
Plane, project tracker1741.18%51.67%
GitLab, code and issues7133.80%43.36%
RocketChat, internal chat7929.11%40.39%
ownCloud, shared files7012.86%22.44%
From Table 4 of TheAgentCompany, arXiv:2412.14161v3. Counts sum above 175 because a single task can require more than one platform.

Agents clear a third of the code work and an eighth of the shared-drive work. The authors say why, and they don’t flatter their own subject. Data science, admin and finance tasks scored lowest, “with many LLMs completing none of the tasks successfully”. Those jobs involve reading scanned documents, collecting information from several people, and grinding through tedious software. Those tasks are easier for a person than software engineering, and harder for the agent. That inversion is the single most important thing on this page if you’re buying seats for a finance or operations team rather than an engineering one. We looked at the engineering half of the same question in AI tools for developers.

Add the review time back and one tier of tool made things worse

A win rate isn’t a saving, because somebody still has to check the output and redo the misses. GDPval models exactly that. Under its “try once, then fix it yourself” scenario, GPT-5 came out 1.12x faster and 1.18x cheaper than an unaided expert, which is a 10.7% time saving and a 15.3% cost saving by our arithmetic. Those are real gains and they’re a long way from the raw comparison, where ignoring review entirely makes the same model look 90x faster.

Now name the loser. In the same table GPT-4o, at a 12.5% win rate, scored 0.87x on speed and 0.90x on cost. Below 1.0x means worse than not using it, so on our arithmetic that tier took about 15% longer and cost about 11% more than the expert working alone. A weak model didn’t deliver a smaller benefit. It delivered a negative one, because the review and rework it generated outweighed everything it produced. That’s the mechanism behind every deployment that quietly gets switched off, and it’s the reason a cheap seat is not automatically the safe choice.

The controlled trial swung from a slowdown to a maybe

Benchmarks test models. A randomised trial tests people using them, and there’s still only one well-known one in this area. METR ran it with 16 experienced open-source developers across 246 real issues from repositories they already maintained, randomising each issue to allow or forbid AI. The published result was that “when developers are allowed to use AI tools, they take 19% longer to complete issues”, with a confidence interval running from +2% to +39%. The perception gap was wider than the effect: the same developers had forecast a 24% speedup, and after living through the slowdown they still believed they’d been sped up by 20%.

That study covered early 2025, and quoting it in 2026 without the sequel would be dishonest. METR ran a second experiment from August 2025 with 10 of the original developers plus 47 new ones, and in February 2026 it published the outcome and its own doubts about it. Returning developers now showed an estimated 18% slowdown, with a confidence interval running from a 38% slowdown to a 9% speedup. Newly recruited developers showed a 4% slowdown, ranging from a 15% slowdown to a 9% speedup. Both intervals cross zero, so neither result separates from no effect at all.

METR’s explanation for why it’s changing the design is the most quotable admission in the whole subject. Between 30% and 50% of its developers said they’d stopped submitting certain tasks, because they didn’t want to do those tasks without AI. Recruitment got harder too, since people wouldn’t give the tools up for a $50 hourly rate. Both effects push the measured speedup down, which is why METR calls its own estimate a lower bound. The measurement is degrading because adoption succeeded, and that’s a genuinely awkward position for anyone who wants a clean number.

What the 2026 leaderboard changes, and what it doesn’t

Both benchmark papers tested models that are now a generation or two old, which is the standing problem with citing any of this. Artificial Analysis keeps a live re-run, GDPval-AA, over the 220-task gold subset. It swaps the expensive part, replacing human expert graders with an LLM judge that picks a winner from two anonymised outputs, and aggregates the pairings into an Elo rating “anchored to a human baseline of 1,000”. As of August 2026 the top of that board reads Claude Opus 5 at max effort on 1831, then the same model at xhigh effort on 1797. GLM-5.3-Flash sits on 1769, GLM-5.3 at max on 1763 and Grok 4.6 at xhigh on 1756.

Read that as direction, not as a score. Every model near the top sits well above the human anchor, which is a different picture from the 47.6% that human graders produced a year earlier on the same tasks. Some of that gap is genuine progress and some of it is the judge. GDPval’s own automated grader reached 66% agreement with human expert graders, which its authors note is 5% below the 71% that human graders managed with each other. A machine-graded board carries that disagreement before any model is compared. The board also lists Claude Opus 5 three times at three effort settings, on 1831, 1797 and 1719, which tells you configuration is now a variable in its own right. Our piece on why AI benchmarks keep lying to you covers the failure modes that apply here.

What we couldn’t price, and what that leaves out

Three gaps, stated rather than estimated. OpenAI’s own pricing pages returned HTTP 403 to every request we made on 26 August 2026, and its Business tab renders client-side, so the ChatGPT rows above come from its help centre instead of its shop window. That page adds “may vary by country and currency” to its dollar figures, which we can’t check from here. Google and Notion served us UK pricing, so those rows sit in pounds and can’t be lined up against the dollar rows without an exchange rate we’d have to invent. And no vendor publishes what a seat costs once support, training and the hours somebody spends checking output are counted, which is precisely the cost GDPval’s review scenario shows to be decisive.

The benchmarks have limits of their own. TheAgentCompany’s 175 tasks live inside one simulated software company, so its department mix isn’t your department mix. GDPval’s 220 graded tasks average 6.7 hours each, which excludes both the two-minute jobs and the six-month ones. Neither measures the thing most buyers actually want, which is whether a team of ten gets more done in a quarter. Our piece on where AI money shows up and where it stalls works through the survey evidence on that question, and the survey evidence is weaker than either benchmark.

What would change the answer

On the evidence we could open, the defensible position in August 2026 is narrow. Seat prices are low relative to the work, at roughly $216 to $384 a year against $361 for a single 6.7-hour expert task. Completion rates on realistic work are partial, between 12.86% and 41.18% depending on which software the job lives in. Net savings after review are real but modest, about 11% of the time and 15% of the cost, and they turn negative on a weak model. And the one randomised trial in the field says its own design is breaking down.

Three things would move that. A GDPval-style run graded by human experts on 2026 models would tell us whether the leaderboard’s jump above the human anchor survives real graders. A version of TheAgentCompany rebuilt around finance and operations software rather than a software company would tell us whether the 12.86% on shared files has moved. And a vendor publishing seat-level usage data, rather than a case study, would let somebody price the review time that all of this turns on. Until at least one of those lands, the tools are cheap, the evidence that they finish your work is thin, and those two facts aren’t in tension.

Get the daily rundown

One email each weekday with the AI news that matters, every claim linked to its primary source.

Free, one email each weekday, unsubscribe in one click. We never sell or share your address.

Leave a Reply

Your email address will not be published. Required fields are marked *