Opinion

What the AI industry gets wrong about its own users

The AI industry has a picture of who uses its products, and it’s wrong in a specific way. It imagines someone technical, curious, and willing to iterate. Most users are none of those things and never will be.

That gap explains a lot of otherwise puzzling product decisions, and a fair amount of the disappointment when features ship.

The AI industry’s imagined user

AssumedActually
Rephrases when the answer is poorConcludes the tool doesn’t work and leaves
Verifies output against a sourceTakes fluent text at face value
Enjoys exploring capabilitiesWants one task finished
Understands what the model can’t doHas no mental model at all
Will learn to prompt wellTypes the question they’d ask a person
Who products are designed for, against who opens them.

Every row is a design decision made on the left and experienced on the right.

Nobody iterates

The single largest mistake is assuming users will try again with a better prompt.

Prompting genuinely works. Google Research showed in Chain-of-Thought Prompting that eight worked examples could reach state of the art on a maths benchmark, beating a fine-tuned model. That’s a large effect from how you ask.

But it’s a technique, and techniques have to be taught. Someone who asks a question, gets a mediocre answer and closes the tab has learned that the product is mediocre, which is a reasonable conclusion from their evidence.

Shipping a product whose quality depends on user skill, without teaching the skill, is shipping a product that fails for most people.

Trust is calibrated wrongly, in both directions

Users either believe everything or nothing, and almost nobody lands in the middle where the tool is actually useful.

Over-trust is the industry’s own doing. Models are tuned on human preference, and preference rewards confident, well-formed answers, as OpenAI’s InstructGPT work demonstrated when a 1.3B tuned model was preferred over 175B GPT-3.

So the output sounds equally certain whether it’s right or invented, and users reasonably read confidence as reliability. Our piece on the vocabulary problem covers how the language makes this worse.

Under-trust is the mirror image. One bad experience produces permanent dismissal, because a person who was told the tool is intelligent and then caught it being wrong updates hard.

Neither group is behaving irrationally. Both are responding sensibly to a product that gives them no way to tell a reliable answer from an unreliable one.

The chat box was a default, not a decision

Almost every AI product presents as a conversation, and it’s worth asking why, since the answer isn’t that research showed it works best.

Chat is what the first successful product used, so it became the pattern. But a blank box is the hardest possible interface for someone with no mental model, because it offers no clue what the system can do.

The features that survive tend not to be conversational. Coding assistants live inside the editor. Retrieval search returns results. Both replace a task rather than opening a dialogue about it.

Retrieval is instructive here. The original RAG paper combined “pre-trained parametric and non-parametric memory”, and what users experience is a search box that understands meaning, which needs no new mental model at all.

That’s the pattern worth copying. The best AI features look like an existing thing working better, not like a new thing requiring instruction.

The industry is discovering this the expensive way

You can watch the correction happening in public.

Microsoft merged its Copilot apps and killed AI-generated podcasts, Group Chats, Copilot Labs and its animated assistant. TechCrunch read the merger as an admission the prior strategy was too complicated to compete.

Look at what died. Every one was a bet on AI as something exploratory and social, which is exactly the imagined user rather than the real one.

TechTimes read the consolidation as evidence of a paid-adoption problem, which is what you’d expect when a product is built for a user who mostly doesn’t exist.

Who is underserved

Three groups get almost nothing designed for them, and they’re large.

People doing repetitive work who’d benefit most from automation are rarely given tools shaped to their workflow. People without technical vocabulary can’t describe what they want to a blank box. And people who need to be certain, in regulated or high-stakes work, get no reliable way to tell a sound answer from a confident one.

That last group matters most, and it’s where verifiability keeps proving decisive. DeepSeek’s R1 work got its strongest results on tasks where answers can be checked mechanically, and products inherit that property or lack it.

The counter-argument

A fair objection says early products should target early adopters, and that’s textbook rather than a mistake.

Technology usually reaches enthusiasts first and gets simplified later, and designing for the mainstream too early produces something nobody wants. Users do adapt, and prompting norms are spreading without anyone teaching them formally.

The reply is that the money has already moved past that stage. Enormous sums are committed to infrastructure, including up to $500 billion arranged through Nvidia, and spending at that scale assumes mainstream adoption rather than an enthusiast base.

What better products would look like

They’d show capability rather than presenting a blank box. They’d surface uncertainty visibly rather than answering everything in the same confident register. They’d fit an existing workflow instead of asking people to bring their work to a chat window.

And they’d assume one attempt. If quality depends on a second prompt, most users never see the good version.

Cost pressure may force some of this. Prices now span $6 to $30 per million output tokens, and products that burn tokens on users who abandon after one attempt are paying for failure at scale.

Which is an odd sort of good news. The economics of serving people badly are unattractive enough that fixing the onboarding problem has a financial argument behind it, not just a design one.

What would change this analysis is published retention data showing sustained use by non-technical users. Companies report adoption when the numbers flatter them and go quiet otherwise, and the silence on retention is itself close to an answer.

Get the daily rundown

One email each weekday with the AI news that matters, every claim linked to its primary source.

Free, one email each weekday, unsubscribe in one click. We never sell or share your address.

Rundowns AI Desk

The Rundowns AI desk covers artificial intelligence research, tools, business and policy. Every factual claim we publish links to the primary source it came from, so readers can check it themselves.

Leave a Reply

Your email address will not be published. Required fields are marked *