The AI compute shortage moved from chips to power
Why is AI compute still scarce years into the buildout? Because the shortage moved. It started as a chip shortage, and it is now mostly a power and financing problem wearing a chip shortage costume.
Following where the constraint actually sits explains a lot of otherwise strange corporate behaviour.
Demand grows faster than any supply chain
Start with the number that drives everything. Epoch AI reports training compute for frontier language models growing 5x per year since 2020, a doubling roughly every 5.2 months, with the top models up about ten-thousandfold over that span.
No physical industry expands at 5x annually. Fabs take years to build, and the equipment inside them has its own multi-year lead times.
So a gap opens by arithmetic, not by mismanagement. Even a perfectly run supply chain would fall behind a demand curve shaped like that.
The AI compute shortage keeps moving
This is the part most coverage gets wrong, because it treats “GPU shortage” as one fixed thing.
| Constraint | Why it binds | How fast it can ease |
|---|---|---|
| Leading-edge wafers | Very few fabs can make them | Years, new fabs are multi-year projects |
| Advanced packaging | Stacking memory onto logic is its own capacity | Quarters to years |
| High-bandwidth memory | Few suppliers, tight qualification | Quarters |
| Data centre power | Grid connections and generation | Years, often the slowest of all |
| Financing | Somebody must fund the build | Weeks, if capital markets cooperate |
Relieve one and the next becomes binding. That is why “the shortage is easing” and “we cannot get capacity” keep being said in the same quarter by people who are both telling the truth about different links.
Power is the constraint nobody can engineer around quickly
A data centre needs a grid connection, and grid connections are granted on timescales set by utilities and regulators rather than by chip vendors.
You can buy accelerators with money. You can’t buy a substation on the same timescale, and in several markets the interconnection queue is now the gating item for new capacity.
Cooling compounds it. Modern accelerator racks draw enough power that air cooling stops working, and retrofitting liquid cooling into an existing facility isn’t a software update. It’s construction.
That’s why so much new capacity is greenfield. It’s often easier to build beside a power source than to upgrade a site that already has tenants.
Chips are a purchasing problem. Power is a permitting problem, and permitting does not respond to urgency.
Which is why financing became the story
If hardware is available and sites are the limit, the question becomes who pays for the build. That reframing explains Nvidia’s most consequential move this month.
Six private capital firms signed non-binding agreements toward up to $500 billion of AI infrastructure financing, and Nvidia made it work by guaranteeing the collateral value of its own chips. Jensen Huang told CNBC his chips are an “investable asset”.
Read that as a diagnosis. A chip vendor underwriting residual value is a vendor whose customers are constrained by balance sheets rather than by allocation.
The customers agree, judging by how they’re raising. Databricks closed $5 billion at a $190 billion valuation with proceeds pointed at AI agent products, and CNBC noted it was the company’s second round this year.
Two rounds in twelve months isn’t what a company does when capital is comfortable. It’s what a company does when the thing it needs costs more than it expected.
The depreciation problem underneath it
Lenders need to know what the hardware is worth if they have to seize it, and the honest answer is uncomfortable.
Goldman Sachs puts usable accelerator lifespans at four to six years, with economic obsolescence from newer generations eroding value on top of physical wear. Older cards get pushed to lower-margin inference work, which drags resale prices further.
So the asset securing a decade-scale infrastructure loan has a half-life measured in a few years, and the company setting that obsolescence schedule is the same one guaranteeing the value. The Motley Fool called that the catch, and it is.
Efficiency quietly reduces the need
The supply side is not the only thing moving. Epoch AI also reports pre-training compute efficiency improving roughly 3x per year, meaning the same capability costs progressively less compute to reach.
That’s a compounding rate applied against the compute curve, so the effective progress rate is the two multiplied together rather than either alone.
Architecture contributes directly. Mixture-of-experts models hold enormous parameter counts while activating a fraction per token, so serving cost decouples from model size.
Training recipes matter too. DeepMind’s compute-optimal work showed that a smaller model trained on more tokens beat a larger one trained the old way, which is a straight efficiency gain from allocation alone.
But inference is eating the savings
Here is the twist that keeps demand high even as efficiency improves.
Reasoning models spend compute at request time rather than only during training. DeepSeek’s R1 demonstrated how much capability that unlocks, and every such request costs more than a direct answer would.
So compute demand shifted from a large one-off training bill to a permanent per-request one, which is a worse shape for capacity planning. Our piece on when that spend is justified works through the arithmetic.
What the AI compute shortage means if you are buying
Three practical consequences follow for anyone smaller than a frontier lab.
Rental pricing moves with the binding constraint, not with chip availability, so quotes can rise in a period when hardware is reportedly plentiful. Ask what is actually scarce in the region you are buying in.
Older hardware is often the better deal for inference, because the market prices it against frontier training demand it will never serve. If your workload fits, the discount is real.
It’s also worth checking what you’re actually renting. An instance advertised on accelerator count can differ several-fold in memory bandwidth, and bandwidth is usually what limits inference throughput rather than raw compute.
And long commitments carry a specific risk here. Locking in multi-year capacity on hardware with a four to six year usable life means betting that efficiency gains will not make your contracted compute uneconomic before the term ends.
Who actually feels the squeeze
Scarcity isn’t distributed evenly, and that’s worth being explicit about because the headlines describe the frontier while most people live somewhere else.
Frontier labs mostly aren’t short of chips. They have multi-year supply agreements and the balance sheets to prepay, so their constraint is sites and power rather than allocation.
Mid-sized companies feel it most. Too big to run on a handful of rented instances, too small for a supply agreement, they’re buying from the spot market at whatever the current binding constraint has done to prices.
Small teams and individuals are, oddly, fine. Consumer hardware keeps improving and open-weight models keep shrinking, so the local option gets more viable each year even while industrial capacity stays tight. That’s part of why self-hosting maths has changed.
The reading that says this eases sooner than expected
The bear case for scarcity deserves airing, because plenty of serious people hold it.
Capacity ordered during a shortage arrives after it. Fab expansions, packaging lines and data centres commissioned at peak demand come online years later, and semiconductor history is largely a history of gluts following shortages.
Add efficiency compounding at 3x per year and demand growth that eventually has to slow, and you get a plausible path where the industry ends up with more capacity than it needs around the time the last projects complete.
Nvidia’s residual-value guarantee looks rather different under that scenario. It’s a commitment made at peak that would be called on precisely when the collateral is worth least.
What would signal the shortage is genuinely over
Watch rental spot prices for sustained falls rather than vendor commentary, because vendors describe allocation and the spot market describes scarcity.
Watch whether the Nvidia financing MOUs convert into binding commitments, since non-binding agreements are cheap to sign and tell you about intent rather than capital.
And watch grid interconnection queues in the major data centre markets. If those shorten, capacity follows within a year or two. If they don’t, no amount of chip supply fixes it.
One caution on all of this. Almost nothing about frontier compute is disclosed. Contract terms, allocation, effective utilisation and real cost per training run are private, so the public picture is assembled from vendor commentary and analyst estimates rather than filings.
That’s a genuine limit on how confidently anyone can call this, including us. The figures in this piece come from Epoch AI’s tracking and from reporting on the Nvidia financing, both of which are external to the companies whose capacity is in question.
Get the daily rundown
One email each weekday with the AI news that matters, every claim linked to its primary source.
Free, one email each weekday, unsubscribe in one click. We never sell or share your address.
