Opinion

The benefits of AI are real. Nine in ten executives don’t see them

The honest summary of the benefits of AI, as of August 2026, is that they’re real, they’re measured, and almost all of the measurement stops at the level of a single task. Randomised trials have found customer support agents resolving 14% more issues an hour, consultants completing 12.2% more work, and Swedish radiologists catching more cancers while their screen-reading workload fell 44%. Then you move up to the level of a company and the effect thins out. A survey of nearly 6,000 executives across four countries found nine in ten reporting no impact on productivity or employment at their own firm over three years.

Both of those findings are good evidence, and most coverage picks one and drops the other. So the useful question isn’t whether AI helps. It’s why a benefit that shows up reliably inside a controlled experiment keeps failing to show up in an income statement, and what that gap tells you about which claims to take seriously. Every figure below comes from a paper, filing or statistical release we opened in August 2026, and each one is dated, because the honest ones expire.

What the controlled trials actually measured

Start with the studies that randomised something, because they’re the only ones that can separate the tool from the enthusiasm of the people using it. There are three worth knowing, and they disagree in a way that turns out to be informative rather than confusing.

Erik Brynjolfsson, Danielle Li and Lindsey Raymond studied the staggered rollout of a conversational assistant across 5,179 customer support agents. Access raised productivity, measured as issues resolved per hour, by 14% on average. That average hides most of the story, and we’ll come back to it.

The second is a pre-registered field experiment run with Boston Consulting Group and written up as Harvard Business School working paper 24-013. It put 758 consultants, about 7% of the firm’s individual-contributor consultants, into three groups: no AI, GPT-4, or GPT-4 plus a prompt engineering overview. On 18 realistic consulting tasks the paper judged to sit inside the model’s capability, the AI groups completed 12.2% more tasks, worked 25.1% more quickly, and produced results rated more than 40% higher on quality than the control group’s.

The third is the one that gets quoted against the other two. METR ran a randomised controlled trial in which 16 experienced open-source developers completed 246 real tasks in mature projects, with each task randomly allowed or denied AI tools. Developers took 19% longer with AI. They’d forecast a 24% speed-up beforehand, and after living through the slowdown they still estimated they’d been 20% faster.

StudySampleTaskMeasured effect
Brynjolfsson, Li and Raymond5,179 support agentsResolving customer issues+14% issues per hour
HBS working paper 24-013, with BCG758 consultants18 tasks inside the model’s range+12.2% tasks, 25.1% faster
Same paper, one task outside that range758 consultantsA deliberately harder problem19 points less likely to be correct
METR16 developers, 246 tasksReal issues in mature repositories19% longer with AI
Figures as reported in each paper. The METR trial covered tools at the February to June 2025 frontier; the consulting experiment used GPT-4 and is dated September 2023.

Read those four rows together and the pattern is not that AI works or doesn’t. It’s that the measured benefit tracks how well-defined the task is, and how bad the person was at it beforehand. That second half is where the interesting number sits.

The gain lands hardest on whoever was worst at the job

The 14% headline from the support-agent study is an average across a very lopsided distribution. The same paper reports a 34% improvement for novice and low-skilled workers, with minimal impact on experienced and highly skilled ones. The consulting experiment found the same shape: consultants below the average performance threshold improved 43% against their own baseline, while those above it improved 17%.

Consultants below the average performance threshold improved 43% against their own baseline. Those above it improved 17%.

Dell’Acqua et al., Harvard Business School working paper 24-013, September 2023

That matters because it changes what a benefit claim even means. A vendor saying “our customers see a 40% productivity gain” may be telling the truth about a population of beginners and saying nothing at all about your senior staff. The support-agent paper’s own reading is that the tool spread the practices of the best agents to the newest ones, which is a real and valuable thing. It’s also a compression effect rather than a general uplift.

The catch for anyone budgeting on those numbers is that the compression cuts the other way on cost. If the gain concentrates in junior work, the case for buying seats for your most expensive people gets thinner, not fatter. We worked through how that logic hits a purchase decision in our piece on where AI actually pays inside a business, and the distribution is the part most pilots never check.

The clearest measured benefit is in a screening clinic, not an office

If you want the single best-evidenced benefit of AI available right now, it isn’t in knowledge work. It’s in Swedish breast cancer screening, and the reason is that somebody ran the trial properly and then waited two years for the outcome that mattered.

The MASAI trial randomised women at four screening sites in Sweden to either AI-supported screen reading or standard double reading by two radiologists. The prespecified safety analysis in The Lancet Oncology, published in August 2023 after 80,033 women were enrolled, reported 244 screen-detected cancers in the AI arm against 203 in the control arm. Cancer detection rates were 6.1 per 1,000 screened against 5.1 per 1,000, a ratio of 1.2 that did not clear conventional significance at p=0.052. The false positive rate was 1.5% in both groups. Screen-reading workload fell 44.3%, because the AI triaged most examinations to a single reader instead of two.

Detecting more cancers at screening isn’t automatically good, though, because you can hit that number by finding harmless things. The test is whether fewer cancers appear in the gap between screening rounds. The full results in The Lancet on 31 January 2026 covered 105,934 randomised women and answered it. Interval cancer rates were 1.55 per 1,000 in the AI arm against 1.76 in the control arm, a ratio of 0.88 that met the trial’s non-inferiority margin without reaching significance on its own (p=0.41). Sensitivity was 80.5% against 73.8%, which did reach significance at p=0.031. Specificity was 98.5% in both arms.

MASAI outcomeAI-supported readingStandard double reading
Cancer detection rate, per 1,000 screened6.15.1
False positive rate1.5%1.5%
Interval cancer rate, per 1,0001.551.76
Sensitivity80.5%73.8%
Specificity98.5%98.5%
Screen-reading workload44.3% lowerbaseline
Detection and workload figures from the safety analysis of 80,033 women, The Lancet Oncology, August 2023. Interval cancer, sensitivity and specificity from the full analysis of 105,934 women, The Lancet, 31 January 2026.

Notice what made that result usable. The task was narrow, the output was checkable, the comparison group did the job the old way at the same time, and the endpoint was decided before anyone looked. Almost no corporate AI deployment meets even two of those conditions, which is why almost no corporate AI deployment produces a number you can trust.

The same methods find the benefit reversing

The consulting experiment didn’t only test tasks the model could do. It included one selected to sit outside the model’s range, and on that task consultants using AI were 19 percentage points less likely to produce a correct solution than consultants without it. The paper’s framing is a jagged capability boundary: two tasks can look equally hard to a person and sit on opposite sides of it.

That’s the mechanism behind the METR result too. Those developers weren’t bad at their jobs and the tools weren’t broken. The time saved on typing got spent on prompting, waiting, reviewing and correcting, and the developers couldn’t feel the difference. They finished believing they’d been 20% faster while the clock said 19% slower, a 39-point gap by our arithmetic, and that’s the most quotable finding in the whole literature. It says self-reported productivity gains are close to worthless as evidence.

When developers are allowed to use AI tools, they take 19% longer to complete issues.

METR, Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity, July 2025

To METR’s credit, they’ve since undercut their own headline. In a February 2026 update the team said they’re changing the experiment design, because a follow-up run gave an unreliable signal. Raw results from the follow-up pointed to a speed-up of 18% for the original developers, with a confidence interval running from 38% faster to 9% slower, and 4% for new ones. The problem is selection: between 30% and 50% of developers told METR they were holding back tasks they didn’t want to attempt without AI. So the later experiment probably missed exactly the work where AI helps most, and METR says so plainly.

Nine in ten executives report no change at their own firm

Task-level gains have to survive contact with an organisation before they become a benefit anybody can bank. Mostly they haven’t yet. Ivan Yotzov, Nicholas Bloom, Steven Davis and colleagues surveyed nearly 6,000 senior executives at US, UK, German and Australian firms for an NBER working paper dated February 2026. They found 69% of firms actively using AI, and more than two thirds of executives using it themselves, for an average of 1.5 hours a week.

Then the punchline: nine in ten of those executives reported no impact on employment or productivity at their own firm over the past three years. The same people forecast a 1.4% productivity boost, 0.8% higher output and a 0.7% cut to employment over the next three. Their employees, asked separately, expected employment at their firms to rise 0.5%.

The adoption numbers are worth separating out too, because “adoption” means four different things depending on who counts. Stanford’s 2026 AI Index reports organisational adoption at 88%, up from 78% a year earlier, drawn from McKinsey’s State of AI survey. The US Census Bureau, sampling businesses rather than enterprise respondents, put current AI use at 19.8% as of 3 May 2026.

SourceWhat it countsFigureAs of
AI Index 2026, via McKinseySurveyed organisations using AI in at least one function88%2025
NBER working paper, Yotzov et al.Firms actively using AI, four countries69%February 2026
Federal Reserve FEDS noteIndividuals reporting work-related generative AI useabout 41%November 2025
Census Bureau BTOSUS businesses using AI in operations19.8%3 May 2026
Four adoption rates for the same technology in the same year. The gap is definitional: the higher numbers survey large organisations and individuals, the lowest samples businesses of every size, including the very small.

Both of those can be right at once. 37% of firms with at least 250 employees told Census they use AI, against under 20% of firms with four or fewer, so a survey weighted toward large organisations lands near 88% and a survey of all businesses lands near 20%. The Federal Reserve’s April 2026 note on tracking adoption puts firm-level use at about 18% at year-end 2025 while individual work-related use sits near 41%, which tells you a lot of the usage is people bringing their own tools. The AI Index says of its own corporate figures that they’re self-reported and “should be viewed as directional rather than comprehensive”. That’s a rare and useful sentence in an industry report.

The strongest version of the optimistic case

Here’s the best argument against everything above, and it’s a good one. Firm-level surveys measure sales per employee, which is a lagging figure that moves after reorganisation, not after tool adoption. General-purpose technologies have historically taken years to show up that way, and three years is a short window in which to expect one.

The AI Index makes that case with numbers. US productivity growth reached 2.7% in 2025, close to double the 1.4% average of the previous decade, which the report says Brynjolfsson reads as possibly the early stage of a J-curve, where organisations absorb adoption costs before the returns arrive. A study of 12,000 European firms cited in the same chapter found AI adoption lifting labour productivity 4%. OECD projections for G7 economies put annual gains at 0.2 to 1.3 percentage points over a decade.

The sharper version of the argument is that productivity is the wrong measure entirely, because most of these tools are free. The AI Index carries an estimate of US consumer surplus from generative AI, built by asking people in choice experiments what they’d accept to give up the tools for a month. Total surplus rose from $112 billion to $172 billion a year between 2025 and early 2026, and the median value per user tripled from $3.40 to $11.40. If the benefit accrues to users rather than to the firms selling access, no income statement anywhere would show it.

Even so, the same chapter carries the Penn Wharton Budget Model’s estimate that AI’s current contribution to US total factor productivity is 0.01 percentage points, described in the report’s own summary table as negligible. So the optimistic case rests on a lag, and a lag is a promise about the future rather than a measurement of the present. It might be right. It hasn’t been checked yet.

What would change this conclusion

The claim in this piece is falsifiable, which is the only thing that makes it worth writing. It’s that the benefits of AI are real, narrow, concentrated among the least skilled people doing the most structured work, and largely absent from firm-level accounts as of August 2026.

Three things would break it. The first is a repeat of the Yotzov survey in 2027 or 2028 showing executives reporting realised productivity gains anywhere near the 1.4% they forecast, since that turns a prediction into a measurement. The second is a randomised trial of AI on unstructured, senior, judgment-heavy work that finds a gain, because every clean positive result so far has come from work that could be scored task by task. The third is a redesigned METR experiment that solves its selection problem and still finds a speed-up, which would tell us the 2025 slowdown was a snapshot of immature tooling rather than a property of experienced work.

Until one of those lands, the sceptical reading holds, and it isn’t a hostile one. A 44% cut in radiologist reading workload with no loss of specificity is a genuine benefit, and so is a 34% lift for novice support agents. Neither of them is transformation, and neither of them arrived because somebody bought a licence. They arrived because somebody defined a task narrowly enough to measure. The gap between those two sentences is where most of the disappointment in enterprise AI deployment comes from, and it’s also where the honest benchmark problem starts, which we covered separately in why AI benchmarks keep lying to you.

Get the daily rundown

One email each weekday with the AI news that matters, every claim linked to its primary source.

Free, one email each weekday, unsubscribe in one click. We never sell or share your address.

Leave a Reply

Your email address will not be published. Required fields are marked *