
Hyperscalers are preparing to spend roughly US$700 billion in 2026 to make artificial intelligence more capable and cheaper to run. That sounds like a straightforward growth story. More models, more compute, more demand, more data centres.
The risk is that the spending is buying the very forces that weaken part of the demand case.
AI build-outs are usually debated through the depreciation question: will the chips become obsolete too fast to earn back their cost? That is a real issue, but it is not the only one. A more important pressure may come from below. Once a model is good enough to complete a defined task, the buyer no longer needs the newest frontier model for that job. And once the cost of running that level of performance keeps falling, more of the work can move onto hardware the buyer already owns.
The problem is not that AI demand disappears. It probably does not. Cheaper intelligence creates more use. The problem is price. The share of work that can be done by a good-enough model on local hardware sets a ceiling on what a data centre can charge for that work.
That ceiling is falling.
The four largest US technology platforms have guided toward extraordinary capital spending. Amazon has pointed to about US$200 billion, Microsoft to about US$190 billion for the calendar year, Alphabet to US$175 billion to US$185 billion, and Meta to US$115 billion to US$135 billion. McKinsey’s widely cited estimate puts worldwide data-centre capital needs near US$6.7 trillion by 2030, with about US$5.2 trillion AI-specific. Morgan Stanley’s estimate is lower, near US$3 trillion through 2029, but still leaves a financing gap it puts around US$1.5 trillion.
These sums create fixed obligations: debt, power contracts, long-dated capacity deals, and depreciation schedules. Those obligations must be serviced regardless of what price inference commands.
The spending buys two things at once. It buys capability, so models can do more. It also buys efficiency, so a fixed level of AI performance becomes cheaper. Both are desirable. Both also undercut scarcity pricing.
Also Read: How AI and blockchain could make commerce decisions more accountable
Capability has a bound for each task. A tax return, a legal draft, a customer-support exchange, or a code review requires a model good enough to finish the job to a competent standard. Once that threshold is crossed, a better model does not make the completed task more complete. The buyer then moves from “best available model” to “cheapest model that clears the bar.”
Efficiency then moves the same work toward commodity pricing. Epoch AI has found the price of reaching a fixed performance milestone falling between ninefold and several-hundredfold a year in some cases. Andreessen Horowitz has tracked inference at constant capability falling at roughly tenfold a year. Gartner expects inference on a trillion-parameter model to cost more than 90 per cent less in 2030 than in 2025, with on-device and edge inference among the drivers.
That matters because the edge is no longer theoretical. Gartner puts AI PCs near 31 per cent of the market in 2025. Counterpoint puts penetration closer to two-fifths. Canalys expects more than 200 million AI PCs shipped annually by 2028, and IDC expects neural processing units to be near-universal in new PCs by then.
Capability is moving with the hardware. A 120-billion-parameter open model now runs on a single desktop appliance at roughly 32 tokens per second. A 70-billion-parameter model runs on a mini-PC costing about US$1,500 at 12 to 15 tokens per second. Apple’s unified-memory Macs and Nvidia and AMD desktop systems put comparable capability on millions of desks. The best open-weight models now trail the leading closed systems by low single-digit points on neutral capability indexes, with several available under permissive licenses.
This does not mean the edge replaces the frontier. It does not. Frontier training remains capital-bound. The largest facilities have minimum efficient scale that owned devices cannot touch. Workloads needing very long context, low latency at high concurrency, proprietary cloud data, always-on orchestration, or the newest reasoning models still belong in centralised infrastructure.
But that is exactly the distinction. The defensible data-centre market is the work that structurally needs the data centre. The fragile slice is ordinary inference that once lived in the cloud only because a capable model could not run anywhere else.
Also Read: The barrier to AI adoption was never budget, it was knowing where to start
For that slice, the buyer has a standing choice: rent compute from a data centre, or run the workload on owned hardware. The amortised cost of the owned option becomes the maximum sustainable cloud price before work defects. A chip can be fully utilised and still earn a thinning margin if the price it commands is set by what the same job costs on a machine the customer already bought.
This is why the depreciation debate misses part of the issue. Obsolescence asks whether a chip’s useful life is shorter than the accounting schedule. Edge substitution asks whether the work the chip serves is still scarce enough to command the assumed price. Those are different risks.
The price signals are already mixed, which is what should make the issue interesting rather than dismissible. Hourly rates for prior-generation accelerators fell sharply from around US$8 in early 2024 into a US$1.50 to US$3.50 band by late 2025, as newer chips arrived and new providers entered the market. That looked like commodity pressure. Then rates reversed: one-year rental contracts rose about 40 per cent from an October 2025 low of US$1.70 to US$2.35 by March 2026, with on-demand capacity sold out. The rebound shows that demand is still strong enough to absorb capacity. It does not prove the ceiling has vanished.
McKinsey’s own estimate captures the uncertainty. It expects roughly 60 per cent to 65 per cent of US and European AI workloads to sit on hyperscaler infrastructure by 2030. That still leaves a third or more elsewhere, and the cloud-versus-edge split is a live swing factor.
The strongest objection is simple: total demand may outrun the whole problem. Agentic workflows can use five to thirty times more tokens than a chatbot exchange. Older chips can flow down into inference rather than strand. One Bernstein estimate says a five-year-old accelerator can still earn about US$0.93 an hour against US$0.28 of cash cost, implying a contribution margin above 70 per cent. If AI-generated productivity gains keep arriving, demand expansion may swamp pricing pressure.
That objection is serious. The likely outcome is not collapse. It is segmentation. Frontier work remains centralised and expensive. Commodity work gets cheaper. Some of that commodity work stays in data centres because management, integration, security, and convenience matter. Some moves to owned hardware because the economics become too obvious to ignore.
The watch points are clear. If open-weight models keep closing the gap, the local option strengthens. If memory shortages keep edge hardware expensive, the ceiling falls more slowly. If frontier capability re-widens, centralised infrastructure keeps more pricing power. If commodity accelerator rental rates soften while AI PC penetration rises, the pressure is binding.
The AI build-out may still work. But the risk is not just that racks sit idle. The sharper risk is that the racks are busy carrying work whose price is being set somewhere else.
—
Editor’s note: e27 aims to foster thought leadership by publishing views from the community. You can also share your perspective by submitting an article, video, podcast, or infographic.
The views expressed in this article are those of the author and do not necessarily reflect the official policy or position of e27.
Join us on WhatsApp, Instagram, Facebook, X, and LinkedIn to stay connected.
The post The US$5 trillion AI data-centre buildout unleashes the paradox that limits its returns appeared first on e27.
