Artificial intelligence and business: 18% of firms, mostly in sales
If you want to know how businesses actually use artificial intelligence, the most honest answer available is narrow, uneven, and thinly measured. The US Census Bureau asked a nationally representative sample of firms between November 2025 and January 2026. Some 18% said they’d used AI in at least one business function, rising to 32% when you weight by employment. Among those adopters, 57% used it in three functions or fewer. The single most common function was sales and marketing, reported by 52% of firms that use AI anywhere.
That’s the adoption picture, and it’s the easy half, because counting who bought something is simpler than counting what it did. The results picture is much thinner, because most functions have no published outcome attached to them at all. What follows goes function by function: what got measured, who measured it, and what the number was actually compared against. That last part matters more than it sounds, and it’s where most of the headline figures quietly fall apart.
The function ranking comes from one nationally representative survey
Almost every “AI by function” chart you’ll see traces back to vendor surveys of self-selected customers. The exception is the AI supplement to the Census Bureau’s Business Trends and Outlook Survey, analysed in an April 2026 working paper by Kathryn Bonney, Cory Breaux, Emin Dinlersoz, Lucia Foster, John Haltiwanger and Aditya Pande, and circulated through NBER. It asks about 15 named functions, so the ranking is a measurement rather than a marketing claim, which is why it anchors everything below.
Conditional on using AI somewhere, sales and marketing leads at 52%, followed by strategy and business development at 45%, information technology at 41% and R&D at 40%. Production, supply chain, quality control and distribution all sit well below that. Weight the same data by employment and the order shifts: IT comes first, then finance and accounting, then sales and marketing. That matters because big employers and small ones are evidently doing different things with the same technology.
Underneath the function layer sits a worker layer, though it’s a good deal smaller. Workers used AI on work tasks in 23% of firms, or 41% employment weighted. Writing and editing dominates at 85% of those firms, with information search a distant second at 50%, and 65% of firms confine it to three tasks or fewer. Anthropic’s own usage data points the same way: its June 2026 Economic Index report, covering 10 April to 10 June 2026, found that work conversations most often produce documents and reports (20%), then explanations (9%), then email drafts (7%).
Two completely independent datasets, a federal survey and a model provider’s own traffic, agree that the main thing businesses do with AI is write things.
Rundowns AI analysis of Census BTOS and Anthropic Economic Index data
Sales and marketing is the most used function and the worst measured
Because it’s the most common function, sales and marketing is where you’d most want a hard number, and the hard number is not flattering. The best one comes from a series of randomised field experiments run across seven workflows at a large cross border online retail platform during 2023 and 2024, written up by Lu Fang, Zhe Yuan, Kaifu Zhang, Dante Donati and Miklos Sarvary. Their June 2026 revision reports effects on sales that range from nothing detectable to 16.3%, and the spread across workflows is the finding.
| Workflow | Effect on sales | Observations |
|---|---|---|
| Pre-sale service chatbot | +16.3% (p<0.01) | 44,614 consumers |
| Search query refinement | +2.9% (p<0.05) | 1,849,382 consumers |
| Product description generation | +2.1% (p<0.05) | 4,772,937 consumers |
| Marketing push message | +1.6%, not significant | 13,715,528 consumers |
| Google advertising title | -4.5%, not significant | 1,244,016 products |
| Chargeback defence | +15% defence success rate | platform’s internal estimate |
| Live chat translation | +5.2% consumer satisfaction | platform’s internal estimate |
But the two pure advertising workflows are the ones that failed. Rewriting Google advertising titles produced a negative point estimate that wasn’t statistically distinguishable from zero, and the push message result didn’t clear significance either. Across the four applications with positive sales effects, the authors put the implied annual value at roughly $4.6 to $5.2 per consumer. The gains ran through conversion rates rather than bigger baskets, and return rates didn’t worsen. Even so, the effects were largest for less experienced buyers, which is a pattern worth holding on to.
Customer service has the cleanest number, and it’s about novices
So notice which workflow won in that table. It wasn’t advertising copy, it was customer service, and that fits the one large worker level study anybody has. Erik Brynjolfsson, Danielle Li and Lindsey Raymond studied the staggered rollout of a conversational assistant across 5,179 customer support agents. Access raised issues resolved per hour by 14% on average. But the average conceals the mechanism, because novice and low skilled workers improved 34% while experienced and highly skilled ones barely moved.
The retail chatbot result carries the same shape, and reading the fine print changes what 16.3% means. Consumers in the control group weren’t handed to a human agent. They got an automated message saying live support wasn’t available, because the platform prioritised human agents for post sale queries. So the comparison is assistance against nothing, which is not the comparison a buyer is making. The authors ran further arms to check it, though.
Comparing consumers who receive GenAI-assisted service with human escalation to those served exclusively by human agents shows that the former spend 11.5% more. We interpret this comparison as a conservative lower bound on GenAI’s sales productivity impact in pre-sale customer support.
Fang, Yuan, Zhang, Donati and Sarvary, arXiv 2510.12049
Set that against what vendors publish for the same function. Salesforce’s Agentforce metrics page, as of September 2026, leads with 3.2 billion Agentic Work Units completed in its second quarter, up 97% on the first. The page defines a work unit as one discrete task, with examples including “a prompt processed, a reasoning chain completed, or a tool invoked”. A prompt processed is an activity count, not a result.
Below that come per customer figures: 58% chat containment at CVS Health, 95% case deflection at Live Nation, 89% of routine messaging inquiries at Canada Goose, 62% of support cases resolved autonomously at Xero. The catch is that containment tells you a human didn’t take the ticket. It doesn’t tell you the customer’s problem went away, and we went through the same gap when we priced business AI tools against benchmarked completion rates.
Software and IT is the function where the measurement broke
IT ranks third by firm count and first by employment weight, so it ought to be the best evidenced function of all. It’s the opposite, and the reason is that somebody actually tried to measure it. METR ran a randomised controlled trial in which 16 experienced open source maintainers worked 246 real issues in their own repositories, each issue randomly assigned to allow or forbid AI tools. The result, published in July 2025, was that tasks took 19% longer with AI available. The developers had forecast a 24% speedup, and even so they still believed afterwards that they’d been sped up by 20%.
That study got quoted everywhere, though usually without the follow up. METR ran it again from August 2025 with 57 developers across 143 repositories and more than 800 tasks, then published in February 2026 that the second round couldn’t be trusted. The point estimates had moved toward speedup, but both confidence intervals crossed zero, running from -38% to +9% for returning participants and from -15% to +9% for new recruits. The selection problem had grown larger than the effect they were trying to measure.
Between 30% and 50% of developers told METR they’d withheld tasks from the experiment because they didn’t want to do those tasks without AI. Others declined to take part at all, so the pool itself was skewed. One participant put it plainly: “I avoid issues like AI can finish things in just 2 hours, but I have to spend 20 hours.” The researchers now say they believe developers are more sped up in early 2026 than their 2025 estimate suggested, while stating that their own data is only very weak evidence for the size of it. That’s an unusually honest position, and it leaves the largest employment weighted function without a defensible productivity number.
Finance and accounting moved the work before it moved the headcount
Finance ranks second on the employment weighted list, and it has the study the software function lacks. Jung Ho Choi and Chloe Xie combined a survey of 277 professional accountants with proprietary field data from an AI enabled accounting platform serving 79 small and medium sized businesses, covering over 200,000 transaction level records. It ran in the Journal of Accounting Research in 2026.
What they found is reallocation rather than reduction, which is the distinction most coverage collapses. Effort shifted systematically away from routine data entry toward business communication and quality assurance, general ledgers got more granular, and month end closing got faster. Accountants intervened selectively when the AI’s confidence scores were low, which is the complementarity you’d hope for. The caution is in the same paper: a framed field experiment showed that while AI assistance raised classification accuracy on average, relying on non consensus AI recommendations can raise the risk of error. So accuracy went up in aggregate while the tail got worse.
Legal and compliance keeps the only public count of its own failures
Legal sits in the middle of the Census function ranking, and it’s the one function where the failures are catalogued in public. Damien Charlotin maintains a database of court decisions in which a tribunal explicitly found that a party relied on hallucinated material. As of its 19 September 2026 update it listed 2,044 cases, 1,397 of them in the United States. Allegations don’t qualify, because a court has to have found it.
The quarterly series in that database is the part worth sitting with: 404 decisions in the fourth quarter of 2025, 450 in the first quarter of 2026, and 411 in the second. Of the 2,024 cases where an outcome is recorded, 370 drew a monetary sanction and 158 a disciplinary referral. But the party mix cuts against the obvious reading, and that is the part most summaries drop. In the third quarter of 2026, 57.9% of the tracked cases involved self represented litigants and 37.4% involved lawyers. A good deal of this is people without counsel, not firms with a broken workflow.
What correlates with performance is breadth, not any single function
Pick any one function and the evidence is patchy. But look across functions and a pattern shows up. The Census authors ran a latent class analysis over the 15 functions and got five recognisable types of firm, which is a more useful map than a ranked list because it tells you what real adoption profiles look like.
| Firm class | Share of all firms | Average employees | Functions using AI | Worker tasks using AI |
|---|---|---|---|---|
| Comprehensive adopters | 1.2% | 34 | 12.0 | 5.0 |
| Institutional and administrative integrators | 4.2% | 41 | 7.3 | 3.7 |
| Technical strategists | 3.4% | 22 | 5.6 | 3.2 |
| Marketing specialists | 8.6% | 18 | 2.3 | 1.4 |
| Minimalist adopters | 10.2% | 30 | 2.1 | 1.3 |
| Non-users | 72.3% | 18 | n/a | n/a |
Comprehensive adopters, the firms running AI across 12 of 15 functions, are 1.2% of all businesses. Marketing specialists and minimalist adopters together are nearly 19%, and both average about two functions. So the median adopter isn’t transforming anything, though the press releases suggest otherwise. It’s using a model in one or two places, most likely to write copy.
The regressions point the same way. Firms using AI anywhere were 3.2 percentage points more likely to report above average current performance and 8.3 points more likely to expect it in six months, with a 6.4 point higher chance of a recent sales increase, against baselines of roughly 31% and 13%. Associations with employment decreases were smaller, at 1.7 and 1.8 points against a 9% to 10% baseline. Breadth of functional deployment tracked performance most strongly, while operational investment, not worker task use, tracked headcount reduction.
The authors name reverse causality themselves: profitable firms can afford to adopt. This is correlation in a cross section, and it’s a different exercise from counting what actually reached production.
What would change these numbers
Three things, and none of them are vendor announcements, which is the point. The first is METR’s redesign, because a credible software productivity estimate would fill the largest hole in the evidence. The second is the next BTOS supplement: firms told the Census Bureau they expected adoption to reach 22% within six months, and the largest expected gains sat in sales and marketing, strategy, customer service, and finance and accounting. If that shows up, the function ranking holds and deepens. If it doesn’t, the 18% starts to look like a ceiling.
The third is more workflow level experiments of the kind the retail platform ran, because they’re the only design that separates a function that works from a function that got budget. Right now exactly one function, customer support, has both a large worker level study and a large field experiment pointing the same direction, and even there the biggest measured effect came from serving customers who were previously getting nothing. Everything else is either an adoption rate or a vendor metric. We said much the same when we looked at where business AI pays and where it stalls, and the newer evidence hasn’t moved that conclusion as far as you’d expect.
Get the daily rundown
One email each weekday with the AI news that matters, every claim linked to its primary source.
Free, one email each weekday, unsubscribe in one click. We never sell or share your address.
