Artificial intelligence in business: what actually got deployed
If you want to know what artificial intelligence in business actually reached production, the awkward answer is that almost nobody publishes a deployment ledger. One organisation does. The US Office of Management and Budget released the 2025 Federal Agency AI Use Case Inventory on 13 April 2026, listing 3,611 individually reported use cases across 41 agencies. We downloaded the underlying spreadsheet and counted the stage field ourselves. Only 1,040 rows are marked deployed. Another 1,479 sit in pre-deployment, 440 are pilots, and 314 are retired.
That ratio is the shape of the whole subject, and it holds outside government too. Announcements outnumber deployments, deployments outnumber measured outcomes, and the handful of places where all three exist are narrow enough to name. Everything below comes from a document we opened in August 2026: an OMB data file, an annual report filed with the SEC, two shareholder letters, two health studies, and one research group’s account of why its own measurement stopped working.
The federal inventory is the only complete deployment ledger anyone publishes
Federal agencies have to report their AI uses annually, and OMB consolidates them into one file. Each row carries a stage of development, and the wording is specific. Deployed means the use case “is being actively authorized or utilized to support the functions or mission of an agency”. Retired means it appeared in the previous year’s inventory and has since been discontinued. So the file lets you do something no vendor case study allows, which is count the failures alongside the launches.
| Stage of development | Use cases | Share of 3,611 |
|---|---|---|
| Pre-deployment (development or acquisition) | 1,479 | 41.0% |
| Deployed (actively authorised or in use) | 1,040 | 28.8% |
| Pilot (limited test) | 440 | 12.2% |
| Retired (discontinued since last year) | 314 | 8.7% |
| No stage recorded | 338 | 9.4% |
Two caveats sit inside that table, and both cut against reading it too confidently. The 338 blank rows aren’t scattered: every one of them comes from the Commerce Department (223), the Tennessee Valley Authority (59) or the Education Department (56). And of the 1,040 deployed rows, 613 carry no date for when the system became operational, so you can’t tell a 2026 launch from a decade-old model that got relabelled. The dated ones cluster recently, with 123 in 2024 and 106 in 2025.
The inventory also flags 445 use cases as high-impact, of which 227 are deployed. Reading those entries is a useful corrective, because they aren’t all software projects. The Veterans Affairs list includes ultrasound and surgical navigation units bought from vendors, which means the count mixes procurement of AI-bearing equipment with systems an agency built. That’s a wider definition than most people have in mind, and it inflates any headline figure drawn from it.
Retirement is where the file earns its keep. The 314 discontinued use cases concentrate in three agencies: Veterans Affairs (72), Health and Human Services (49) and Homeland Security (33). They include an SEC comment letter review tool, a Homeland Security disaster assistance projection, and an HHS literature review dashboard. Nothing in the file says why any of them stopped, which is a real gap, but the existence of a public retirement count is still more than the private sector offers. Our earlier piece on what enterprise AI actually reaches production found the same asymmetry from the survey side.
In annual filings, agentic AI is still mostly a seller’s word
Filings are the other place where a company has to be careful about what it claims. We queried the SEC’s EDGAR full-text search on 24 August 2026 for annual reports on form 10-K, including amendments, filed between 1 January and 1 August of each year. The phrase counts show how fast the vocabulary moved.
| Phrase in the filing text | 2024 | 2025 | 2026 |
|---|---|---|---|
| “agentic AI” | 0 | 36 | 221 |
| “generative AI” | 430 | 774 | 1,138 |
| “artificial intelligence” | 2,143 | 2,977 | 3,689 |
Read the 2026 list of 221 and the pattern is obvious. It’s dominated by software and services companies selling AI: C3.ai, Salesforce, Workday, UiPath, Zoom, Adobe, Palantir, Genpact, Concentrix. A phrase appearing in a 10-K tells you a company thinks investors care about it, which isn’t the same as a system running in a warehouse, so the count measures attention rather than deployment. That gap is the one we set out four questions for in every software company says it does AI now.
The buyer-side exceptions are worth pulling out. C. H. Robinson, a freight broker rather than a software vendor, described its deployment in its 10-K filed on 13 February 2026. The filing says the company managed roughly 37 million shipments in 2025 for 75,000 customers across more than 450,000 contract carriers, employs about 800 technologists, and that its “fleet of more than 30 AI agents is integrated with Navisphere”, the transportation management system most of its network runs on.
The same filing shows why attributing results to that fleet is harder than it looks. Average employee headcount fell 11.5% to 12,733 in 2025, and personnel expenses fell 5.9% to $1.4 billion. The company attributes both to “cost-optimization efforts and productivity improvements and the divestiture of our Europe Surface Transportation business”. That’s three causes and one number, because the filing gives no split between them. Income from operations rose 18.8% to $795.0 million over the same year, which is a real result that the filing doesn’t hand to AI alone.
The most detailed production numbers come from companies watching their own engineers
Block published the most granular internal deployment figures we found this year. Its first-quarter 2026 shareholder letter says that as of mid-April, production code changes per engineer were up over 2.5 times compared with January. Incident rates after a production code change were down over 70% against the first quarter of 2025, and down over 40% against the fourth quarter of 2025. The company began building its agent harness, goose, in early 2024.
Builderbot is executing over 200,000 operations per day and is making 15% of our production code changes nearly fully autonomously, with people only intervening to make the final decision to push to production.
Block, Q1 2026 shareholder letter
Those are the numbers a deployment ledger would want. The letter adds that in the first two weeks of April, Builderbot and other AI tools reviewed over 90% of production code change requests made, that as of early April 100% of Block employees were using AI tools for their work, and that its Managerbot product was available to over 1 million sellers. Each figure names a period, which is rarer than it should be.
The catch is that all of it is self-reported and none of it is audited. Code changes per engineer is a throughput measure, not a value measure, so a team shipping 2.5 times as many changes might be shipping smaller ones. Block itself credits the incident improvement to “a combination of our AI investments and our broader focus on reliability”, which is the same two-causes problem the freight filing has. So the disclosure is unusually good, and even so it can’t isolate the effect.
Vendor usage metrics answer a different question than deployment
Software vendors publish the most AI numbers of anyone, and they’re measuring their own funnel. Atlassian’s fourth-quarter FY26 shareholder letter reports that over 80% of Fortune 500 companies now use Rovo, that Rovo assisted actions grew over 50% quarter on quarter, and that monthly active users of its MCP server and Teamwork Graph CLI more than doubled during the quarter to pass 1 million.
Then come the outcome claims, and they change register. The letter says customers adopting Rovo complete 20% more Jira work items and create or edit 25% more Confluence pages than non-adopters, and that adopters grow their annual recurring revenue more than twice as fast. That’s a comparison between two groups who chose differently, so the companies buying AI features first may simply be the companies already growing fastest. The letter doesn’t claim otherwise, and a reader shouldn’t either.
There’s a denominator problem underneath it too. “Over 80% of the Fortune 500 use Rovo” counts organisations where somebody has it switched on, which could be nine people in one department. That makes it a licence statistic rather than a deployment statistic. We worked through what per-seat AI actually costs an organisation in our piece on AI cost per employee, and the same distinction between accounts and use runs through it.
Hospital documentation is the one deployment with both a count and an outcome
Clinical documentation is the clearest case we found of an industry-wide deployment that somebody counted independently. A study published in the American Journal of Managed Care on 28 January 2026 measured adoption of ambient AI scribes, which listen to a patient visit and draft the clinical note for a doctor to check. According to the summary from Emory’s Rollins School of Public Health, where the lead author works, nearly two-thirds of the 2,784 US hospitals using the Epic record system had adopted ambient AI tools by 2025, and three products accounted for more than 80% of them: DAX Copilot, Abridge and ThinkAndor.
Adoption wasn’t even, which is the finding the authors lead with. It ran higher at larger hospitals, nonprofit ones, metropolitan ones, and those with stronger operating margins. Ilana Graetz, the professor of health policy and management who led the work, is quoted saying the results “reveal uneven adoption patterns that could have important equity implications”. A deployment count that varies with a hospital’s margin is telling you about balance sheets as much as about technology.
The outcome side has its own study. Emory Healthcare and Mass General Brigham surveyed 1,430 clinicians, 557 at Emory and 873 at Mass General Brigham, and published the results in JAMA Network Open on 22 August 2025. Emory’s write-up reports a 30.7% absolute increase in documentation-related well-being at 60 days at Emory, and a 21.2% absolute reduction in burnout prevalence at 84 days at Mass General Brigham. Even so, those are survey measures over roughly two to three months rather than clinical outcomes or savings, and the authors don’t present them as either.
The measurement itself got harder in 2026
The cleanest study of whether AI tools speed people up has now been withdrawn from service by the group that ran it. METR published a randomised trial in 2025 finding that AI use caused experienced open-source developers to take 19% longer on tasks, with a confidence interval from +2% to +39%. It started a bigger follow-up in August 2025, and in a note published on 24 February 2026 it explained why the new data doesn’t work.
We have observed a significant increase in developers choosing not to participate in the study because they do not wish to work without AI, which likely biases downwards our estimate of AI-assisted speedup.
METR, We are Changing our Developer Productivity Experiment Design, 24 February 2026
The follow-up ran with 10 developers from the original study plus 47 new recruits, at $50 an hour instead of $150. METR reports an estimated speedup of -18% for the returning group, with a confidence interval from -38% to +9%, and -4% for the new recruits, interval -15% to +9%. Both intervals cross zero, which means neither one settles anything. Between 30% and 50% of developers told the researchers they’d stopped submitting some tasks because they didn’t want to attempt them without AI, which biases the sample toward the work AI helps least.
That’s a quietly important result for anyone budgeting on productivity gains. The randomised design that could have priced the benefit is being redesigned because AI adoption broke its control group. METR says it believes developers are more sped up now than in early 2025, and says in the same breath that its data is “only very weak evidence” for how much. Nobody else is running a comparable trial in public.
What a deployment claim has to carry before it means much
Put these sources side by side and a pattern falls out, because the same four qualifiers keep deciding which claims hold. The claims that survive scrutiny name a stage, a denominator, a period and a measurer. The federal inventory names the stage, which is why it can report 314 retirements. Block names the period on every figure, so mid-April 2026 doesn’t get quietly compared with a different month. The hospital study names an independent measurer, because Rollins researchers counted Epic installations rather than asking vendors for numbers.
The claims that don’t survive tend to fail on the denominator. “Over 80% of the Fortune 500” counts organisations, not workflows. A headcount reduction with three stated causes isn’t an AI saving. An adopter-versus-non-adopter growth gap is a comparison of self-selected groups. None of these is dishonest, and each becomes misleading the moment it’s quoted without the qualifier its own source attached. Our piece on where AI money shows up and where it stalls works through the survey side of the same problem.
So what would change this picture? A second large organisation publishing a stage-coded inventory with retirements would do it, because one ledger from one government isn’t a base rate for anything. So would a large deployer disclosing a retirement count of its own, or a randomised trial that survives contact with a workforce that already uses the tools. Until one of those arrives, the honest summary of artificial intelligence in business in 2026 is that we can count what got announced, we can partly count what got switched on, and we can almost never price what it did.
Get the daily rundown
One email each weekday with the AI news that matters, every claim linked to its primary source.
Free, one email each weekday, unsubscribe in one click. We never sell or share your address.
