Funding & Business

Who keeps the money in the AI stack: chips, models or apps

Who actually makes money from AI? The answer depends on which layer of the stack you’re looking at, and the margin is distributed far less evenly than the attention is.

Four layers, each with different economics, different competitive dynamics, and different odds of still existing in five years.

The four layers of AI stack economics

LayerWhat it sellsCompetitive position
SiliconAccelerators, memory, packagingExtremely concentrated
InfrastructureData centres, power, cloud capacityCapital-intensive, few players
Foundation modelsAPI access to trained modelsSeveral credible rivals, prices falling
ApplicationsProducts built on topCrowded, low barriers
The AI value chain, top to bottom.

Margin concentrates where substitution is hardest. That’s the whole analysis in one sentence, and it explains why the layer with the least public attention captures the most value.

Read the table from the bottom up and the pattern is clear. Applications have thousands of competitors and almost no barrier to entry. Silicon has a handful, and building a new entrant takes years and tens of billions.

Silicon holds the strongest position

Everyone above needs accelerators, and there is no comparable substitute available at volume. That’s about as good as a market position gets.

The clearest evidence of that leverage came this month. Six private capital firms signed non-binding agreements toward up to $500 billion of AI data centre financing, with Nvidia guaranteeing the collateral value of its own hardware.

Jensen Huang told CNBC his chips are an “investable asset”. A supplier that can make its product into collateral is not competing on price.

The layer nobody writes headlines about is the one arranging half a trillion dollars of financing for its own customers.

Infrastructure is capital-hungry and defensible

Below the models sit data centres, and the barrier here isn’t technology. It’s land, power and permitting, none of which respond to being wanted urgently.

Grid connections are granted on utility timescales, and cooling for modern racks means construction rather than configuration. Our piece on where the compute shortage actually sits works through why power became the binding constraint.

The returns are real but slower, and the risk is different. You’re betting that hardware with a four to six year usable life will stay economically relevant across a much longer financing term.

Demand at least is not in doubt. Epoch AI tracks frontier training compute growing 5x per year since 2020, and inference demand grows on top of that as products ship.

The catch is who carries the risk when it turns. The Motley Fool pointed out that a residual-value guarantee is only as good as the guarantor’s willingness to honour it in a downturn, which is exactly when it would be called.

Foundation models are where the fight is

This is the layer everyone writes about, and its economics are getting worse rather than better.

Several labs now ship broadly comparable capability, so buyers can switch. Grok 4.6 matched GPT-5.6 Sol Max at 61 on the Artificial Analysis index while charging $6 per million output tokens against $30.

That’s the signature of commoditisation. When two suppliers offer the same measured capability and one is five times cheaper, the expensive one is selling brand and switching costs rather than a technical advantage.

The pricing is externally checkable too, which matters. Published Artificial Analysis figures put the gap at roughly 60% on input and five times on output, and a price list doesn’t move with interpretation the way a benchmark does.

Open weights push harder in the same direction. Meta released a 30B model under Apache 2.0 and says its most advanced model is next, which sets a price ceiling at zero for anyone able to self-host.

CNBC read that release as a swipe at OpenAI and Anthropic, and commercially it functions as one. Meta doesn’t sell model access, so giving weights away costs it nothing and costs its rivals their pricing power.

That is the classic move of a company attacking a layer it doesn’t need to profit from, and it works because Meta monetises attention elsewhere.

Applications have the weakest moat and the best access

The top layer is where most companies live, and it’s the easiest to enter and the hardest to defend.

If your product is a thin interface over someone else’s API, a competitor can rebuild it quickly and the model provider can absorb it into their own offering. That risk is not hypothetical.

But it’s also the layer with the customer relationship, the proprietary data and the workflow integration, and those are genuine moats when a company actually has them.

Distribution counts for a lot here. A product already embedded in a company’s daily workflow can add AI features and keep the account, while a better standalone tool struggles to get installed at all.

The test worth applying is simple. If the model underneath your product were replaced with a competitor’s tomorrow, would customers notice or leave? If neither, you have a real business. If both, you were reselling the model.

The layer that gets overlooked entirely

There’s a fifth position that doesn’t fit the stack diagram, and it may be the most defensible of all: deployment services.

IBM is betting heavily on it, folding GPT-5.6, Codex and ChatGPT Work into its consulting platform and building a practice of thousands of consultants, per its own announcement.

The insight is that enterprises can’t deploy this themselves. Integration, compliance sign-off and change management are the bottleneck, and none of them are model problems.

Selling implementation capacity for whichever model wins is a hedge against the entire model layer commoditising, which is a reasonable thing to hedge against.

Where AI stack economics land the money today

Because gross margin behaves so differently by layer, it’s worth being concrete about the shape of each business rather than the headline revenue.

Silicon earns high margins on scarce products, which is why the layer can afford to underwrite its own customers’ financing. Infrastructure earns lower margins on enormous capital, so it depends on utilisation staying high for years.

Model providers sit in the worst position of the four. They carry frontier training costs, face falling prices, and compete against a free option that improves annually. That combination has historically been fatal to pricing power.

Applications keep whatever they can defend, which varies enormously between a company with proprietary workflow data and one with a chat box over an API.

Everyone is trying to escape their layer

Once you see the map, the strategic moves become legible as attempts to stop being pinned to one row.

Chip vendors financing data centres are reaching down into infrastructure. Model labs building consumer apps are reaching up into applications. Application companies fine-tuning open weights are reaching down into models.

Microsoft’s consolidation reads this way too. Merging its Copilot apps and moving Deep Research behind a paywall is an application-layer company trying to convert distribution into subscription revenue.

The counter-argument

A serious objection says this analysis is backwards, and it deserves stating.

Silicon looks unassailable now, but hardware markets have collapsed before, and a chip vendor guaranteeing residual value on its own product is exposed precisely when demand turns. Databricks raising $5 billion at a $190 billion valuation shows the application layer can attract capital on a scale that funds vertical integration.

And in every previous computing wave, value migrated upward over time. Hardware commoditised, then operating systems, then applications captured the customer. CNBC noted the Databricks round was its second this year, which is what a well-capitalised application layer looks like.

There’s no strong reason AI must break that pattern. The counter to the counter is that previous waves didn’t have a supplier whose product depreciated to near-worthlessness in four to six years, which changes the shape of the bet at every layer above it.

What would change the picture

A credible second source of accelerators at volume would break silicon’s position faster than anything else, and it’s the development the whole chain is watching for.

Model prices continuing to fall while capability converges would confirm that the model layer is becoming infrastructure rather than product, which is already the direction of travel.

Consolidation would tell you something too. If model labs start being acquired by infrastructure or application companies rather than raising independently, that’s the market deciding the middle layer isn’t a standalone business.

Valuations are the noisiest signal in that picture, and our guide to what a headline valuation actually means covers what the number leaves out.

And application companies reporting durable gross margins would show the top layer has found real moats. Worth noting that almost none of this is verifiable from filings, since most of the interesting companies are private and the public ones bury AI inside larger segments.

Get the daily rundown

One email each weekday with the AI news that matters, every claim linked to its primary source.

Free, one email each weekday, unsubscribe in one click. We never sell or share your address.

Rundowns AI Desk

The Rundowns AI desk covers artificial intelligence research, tools, business and policy. Every factual claim we publish links to the primary source it came from, so readers can check it themselves.

Leave a Reply

Your email address will not be published. Required fields are marked *