Enterprise artificial intelligence: what actually reaches production
Search for enterprise artificial intelligence and you get two kinds of page: vendors reporting record adoption, and surveys that agree with them. Both are accurate, but both dodge the question a buyer actually has, which is how much of this reaches production. The measured answer is a small fraction. MIT’s Project NANDA tracked task-specific enterprise AI tools through the first half of 2025 and found 5% reached production. Deloitte’s third quarterly survey, fielded in mid-2024, found 68% of executives said 30% or fewer of their generative AI experiments had gone fully live. S&P Global Market Intelligence put the share of companies abandoning most of their AI initiatives at 42%. The US Census Bureau, which asks firms directly instead of asking executives about strategy, found 18% of firms using AI in any business function.
Those studies used different samples, different definitions and different years, so the agreement between them is worth more than any one number, because four wrong methods rarely land in the same place. What follows is the production evidence from documents we opened in August 2026: what the funnel looks like, why projects stall, what separates the teams that ship, what it costs, and where the headline figure is weaker than it sounds. Every figure links to the page it came from.
Four measurements, four methods, one answer
The most detailed funnel comes from MIT’s Project NANDA, whose July 2025 report The GenAI Divide: State of AI in Business 2025 reviewed more than 300 publicly disclosed AI initiatives, interviewed representatives from 52 organizations and surveyed 153 senior leaders. It puts enterprise generative AI investment at $30 to $40 billion and finds 95% of organizations getting zero return on it. For embedded, task-specific tools the drop-off is steep: “Sixty percent of organizations evaluated such tools, but only 20 percent reached pilot stage and just 5 percent reached production.”
General-purpose chat tools behave completely differently in the same report, which is the detail most summaries drop. Over 80% of organizations had explored or piloted ChatGPT and Copilot, and nearly 40% reported deployment. So the 5% figure isn’t a statement about whether staff use AI. It’s a statement about whether a system got wired into a workflow that the business depends on, which is a much higher bar.
| Study | What it measured | Result | Sample |
|---|---|---|---|
| MIT Project NANDA, July 2025 | Task-specific GenAI tools reaching production | 5% | 52 orgs, 153 leaders, 300+ public initiatives |
| Deloitte, Q3 2024 | Executives whose org moved 30% or fewer experiments fully into production | 68% | 2,770 respondents |
| S&P Global Market Intelligence, reported March 2025 | Companies abandoning most AI initiatives | 42%, up from 17% | 1,000+ respondents, North America and Europe |
| US Census Bureau, April 2026 working paper | Firms using AI in a business function | 18%, or 32% employment-weighted | US firms, Census survey data |
The abandonment figure comes from S&P Global Market Intelligence, as CIO Dive reported in March 2025: the share of companies dropping most of their AI initiatives “jumped to 42%, up from 17% last year”, across a survey of more than 1,000 respondents in North America and Europe. Deloitte’s number is older and phrased differently. Its third quarterly State of Generative AI in the Enterprise report, fielded across 2,770 respondents in May and June 2024, found “a large majority of respondents (68%) saying their organization has moved 30% or fewer of their GenAI experiments fully into production”.
Adoption is not deployment, and the government data shows the gap
Survey work asks executives what their organization is doing, so the answer comes back flattered. The Census Bureau asks firms, at scale, on a fixed instrument. Its April 2026 working paper The Microstructure of AI Diffusion, by Kathryn Bonney, Cory Breaux, Emin Dinlersoz, Lucia Foster, John Haltiwanger and Aditya Pande, reports that 18% of firms used AI in a business function, rising to 32% on an employment-weighted basis. That gap between firm counts and employment weights is the whole story of enterprise AI: large employers are in, but most companies still aren’t.
The same paper finds adoption is shallow where it exists. It reports that 57% of AI users run it in three or fewer business functions, led by sales and marketing at 52%, strategy and business development at 45%, and IT at 41%. Employment effects barely register either, though not for want of trying: AI-related employment decreases occur in only 2% of firms, and 66% of users rely on AI purely to augment existing work.
The Bureau’s Business Trends and Outlook Survey, collected between 14 December 2025 and 3 May 2026 and written up by Adam Grundy, Cory Breaux and Dhanapati Khatiwoda, shows the same split by size and sector. Overall use “hovered between 17% and 20%”. Firms with at least 250 employees reported 37%, firms with 100 to 249 reported 32%, and fewer than 20% of firms with four or fewer employees said they used it. By sector, Information sat at 39.7% and Finance and Insurance at 33.9%, while Retail Trade was around 14%.
Set that against the adoption headline and the gap is obvious. Stanford HAI’s 2026 AI Index reports that “organizational AI adoption continued to rise in 2025, up to 88% of surveyed organizations”, with generative AI used in at least one business function at 70% of them. The same chapter then notes that “AI agent deployment was in the single digits across nearly all business functions”. Near-universal adoption sits beside single-digit agent deployment in the same report, and the two are measuring different things.
Adoption counts logins. Production counts the processes a company would have to pause if the system went down. Almost nothing has crossed from the first to the second.
Rundowns AI
Why the pilots stall
NANDA surveyed executive sponsors and frontline users across its 52 organizations and asked them to rate barriers on a 1 to 10 frequency scale. Unwillingness to adopt new tools ranked highest, which surprises nobody. Model output quality ranked second, which should, because the same people rating enterprise tools as unreliable were heavy ChatGPT users themselves. The report’s reading is that the problem isn’t raw capability but memory: tools that don’t learn, don’t retain context, and break on edge cases.
One interview in the report makes the point better than the survey does. A corporate lawyer whose firm spent $50,000 on a specialized contract analysis tool kept defaulting to ChatGPT anyway.
It’s excellent for brainstorming and first drafts, but it doesn’t retain knowledge of client preferences or learn from previous edits. It repeats the same mistakes and requires extensive context input for each session. For high-stakes work, I need a system that accumulates knowledge and improves over time.
Corporate lawyer at a mid-sized firm, quoted in The GenAI Divide, MIT Project NANDA
That’s a deployment complaint, not a capability complaint, and it’s the argument we made from the funding side in our piece on the boring layer nobody wants to sell. The work that decides whether a pilot ships is access review, process documentation, data mapping and accountability. None of that shows up in a demo, so none of it gets budgeted early, and the bill still arrives later.
What the benchmark evidence says about reliability
Survey data tells you what buyers believe. Benchmarks on realistic tasks tell you what the systems can actually do, and the February 2026 paper Agent-Diff, by Hubert M. Pysklo, Artem Zhuravel and Patrick D. Watson, is one of the few that tests enterprise software rather than puzzles. It runs nine models over 224 tasks against containerized replicas of four real APIs: Box, Calendar, Linear and Slack. Success is scored as whether the expected change in environment state was achieved, which is a harder test than judging the output text.
Under the paper’s no-documentation baseline, where the agent has to discover endpoints by exploring, DeepSeek-v3.2 led with an overall score of 88.1 and a 76% pass rate. Llama-4-Scout came last at 38.0 and 29%. Injecting the relevant API documentation into the prompt raised pass rates by 7.0 percentage points on average, which helps but doesn’t close the gap. Two caveats matter: the nine models tested sit in the fast and small tier rather than at the current frontier, and the tasks run inside four applications rather than across a multi-week business process.
Even so, a 76% pass rate is the ceiling in that study, not the floor. A quarter of attempts failing is fine for a drafting assistant but unusable for an unattended process, which is roughly where the 2026 research on long-horizon agents lands too. The number that decides deployment isn’t the average score, it’s the cost of the tail.
What separates the teams that ship
NANDA’s most useful finding isn’t the failure rate, it’s the split inside it. In its sample, external partnerships with customized tools reached deployment around 67% of the time, against roughly 33% for tools built internally. The report states this plainly as a correlation and warns that it “may reflect organizational capabilities rather than implementation approach alone”, since the firms that buy may simply be better at procurement. Speed splits the same way, and the direction is counterintuitive. Top-performing mid-market companies reported 90 days from pilot to full implementation, while enterprises, defined there as firms above $100 million in revenue, took nine months or longer.
| Pattern | What NANDA’s sample reported |
|---|---|
| Bought and co-developed with a vendor | ~67% reached deployment |
| Built internally | ~33% reached deployment |
| Mid-market top performers | 90 days from pilot to implementation |
| Enterprises above $100M revenue | Nine months or longer |
| Companies with a paid LLM subscription | 40% |
| Companies whose staff use personal AI tools for work | Over 90% |
That last pair is the shadow AI economy, and it’s the strongest evidence in the whole report. Only 40% of companies had bought an official LLM subscription, while workers at over 90% of the surveyed companies were regularly using personal AI tools for work anyway. So the demand is real and the procurement is what’s missing, which is a very different diagnosis from the one the 95% headline implies.
Where the return showed up is the other surprise. NANDA documented back-office savings of $2 million to $10 million annually from eliminating business process outsourcing in customer service and document processing, a 30% cut in external creative and content costs, and $1 million saved annually on outsourced risk management in financial services. Those gains came from reduced external spend rather than headcount cuts, which matches the Census finding that employment reductions show up in only 2% of firms.
What it costs, and where the budget actually goes
Per-seat pricing is the part a buyer can check without a survey. As of August 2026, Microsoft lists Microsoft 365 Copilot for business at $21.00 per user per month on an annual subscription, with a promotional $18.00 first year running to 30 September 2026, and it requires a qualifying Microsoft 365 Business plan underneath. For a thousand seats that’s $252,000 a year by our arithmetic, spent before anyone measures whether it changed an outcome, and it’s the reason the deployment question is a finance question rather than a technology one.
Budgets are still growing regardless. Andreessen Horowitz’s survey of 100 CIOs across 15 industries, published in June 2025, found enterprise leaders expecting around 75% growth in AI spend over the following year, with 37% already running five or more models against 29% the year before. One CIO in that survey put it as “what I spent in 2023 I now spend in a week”. Growing spend and flat production rates can coexist for a while, but that’s exactly the condition the abandonment figures describe, so the two datasets aren’t in conflict.
Where the money lands is the last piece. NANDA’s executives, asked to allocate a hypothetical $100, sent roughly 70% to sales and marketing, though the report’s own summary text puts that figure at 50% in two other places. Either way the bias runs toward functions with visible metrics, because attribution is easier there, while the documented savings sat in the back office. If you want the arithmetic underneath a single feature rather than a whole programme, we worked one through properly here.
The strongest case against the 95% number
A sceptical reader should push hard on the headline, and the report invites it. NANDA states its own limits: “These figures are directionally accurate based on individual interviews rather than official company reporting. Sample sizes vary by category, and success definitions may differ across organizations.” Fifty-two interviews and 153 conference-recruited survey responses make a qualitative base, though the number got quoted as though it were census data.
The internal inconsistencies matter too. The executive summary says only 2 of 8 major sectors show meaningful structural change, while a later section says 7 of 9 sectors show none. The budget allocation appears as 70% in the body and 50% in two summaries. Neither error changes the direction, but both tell you the precision isn’t there, which is why quoting 95% to the percentage point is a misuse of the document.
Timing is the other weakness. Deloitte’s survey closed in June 2024, the S&P finding was reported in March 2025, and NANDA’s research period ended in June 2025. Tooling moved a lot in the year after that, and the Census data collected to May 2026 has firms expecting more use rather than less, so the picture isn’t frozen. The honest reading is that the pilot-to-production gap was large and well documented through mid-2025, and that nothing published since has closed it. That’s a weaker claim than the headline and it’s the one the evidence supports.
What would change this conclusion
Three things would move it, and all three are checkable without waiting for a vendor to tell you. The first is the Census series: if firm-level use climbs past the 22% that respondents projected for the following six months, and if the share running AI in more than three business functions rises with it, then depth is following breadth. The second is the AI Index agent number. Single-digit agent deployment across business functions is the cleanest measure of production use available, which is why a jump there would be harder to explain away than any vendor survey.
The third is a replication of the NANDA funnel with a bigger, non-conference sample and a published definition of production. We haven’t found one. Until someone runs it, the 5% figure stays what its authors called it, which is directional, and the safer version of the finding is the one four independent studies agree on: adoption is broad, deployment is narrow, and the gap between them is organizational rather than technical.
Get the daily rundown
One email each weekday with the AI news that matters, every claim linked to its primary source.
Free, one email each weekday, unsubscribe in one click. We never sell or share your address.
