The AI Cloud Math is Broken, and It's Creating a Power Shift Within Big Tech

Dow Jones
昨天

Buying AI hardware outright is becoming more economical than renting AI power through the cloud

Renting computing power has stopped getting cheaper fast enough to matter.

In July, I spent EUR64,780, or about $75,500, on four artificial-intelligence graphics cards and the server that holds them, rather than renting the same processing power from Amazon.com (AMZN) by the hour. And if enough companies run that calculation at a much greater scale, it changes who profits from the AI boom.

Let's look at the math: When it went on sale in April 2025, the Nvidia (NVDA) RTX PRO 6000 Blackwell was initially listed at $8,565. By June 2026, the price jumped up to $13,250. In August 2026, those who wanted the graphics card needed to shell out a staggering $16,000, or roughly 87% above the initial price.

Enterprise hardware is supposed to get cheaper after it ships. When it does the opposite, the reasoning behind a great deal of corporate AI spending is called into question.

The assumption has been that renting always wins, and it was a reasonable one. Renting means paying a cloud provider like Amazon's AWS by the hour for time on somebody else's chips. Hardware depreciates and demand comes in bursts, so a bought machine sits idle between them, as the price of processing falls every year. Owning meant tying up capital in equipment that lost value while you learned how to use it.

Opinion: AMD is betting on dirt-cheap AI chips, but financing them is a major question mark

Most of that no longer holds. Hardware is not depreciating on schedule, and renting has stopped getting cheaper fast enough to matter. The last argument for renting is that you pay only for the hours you use, and those idle hours are disappearing as more of the work runs through AI. So, is now the time to own your own hardware rather than rent it?

To answer that question, you first need to calculate the cost of buying versus renting the same equipment. The RTX PRO 6000 Blackwell is a useful measuring stick, since it's sold both to the cloud providers and the companies that would otherwise rent from them. Nvidia ships the server version of the card through Dell Technologies (DELL), Hewlett Packard Enterprise (HPE) and Lenovo (HK:992) $(LNVGY)$, and the large clouds rent it by the hour. Amazon uses it in its G7e instances, describing the configurations as cost-effective for generative AI inference, which is the everyday workflow in which a trained AI model answers questions or processes documents.

Amazon's four-GPU instance costs $16.57 an hour, or about $12,100 a month. A server built around the same four processors costs about $75,500 to put on your own premises. Run that Amazon instance around the clock, and you have spent the same amount of money in a little over six months. Run your own machine seven hours a day, and it pays for itself in under two years against renting from Amazon. Against the middle of the rental market, where four of these processors go for about $10 an hour, it takes just under three years.

Tech leaders at more than 60 companies told Deloitte that they start weighing on-premises alternatives once cloud bills reach 60% to 70% of the cost of buying the equivalent hardware.

Many companies are past that line without realizing it. The FinOps Foundation's State of FinOps 2026 surveyed 1,192 users overseeing more than $83 billion in cloud spending and found that 98% of them now manage AI costs - up from 31% two years ago - and believe that AI pricing is less transparent and more variable than it is for the conventional cloud.

Falling prices used to cover the difference. YipitData found that effective API prices fell 6% in May 2026, after falling 39% in the second half of 2025. Tokens, which are the units of data that AI models process and the unit providers bill for, are still getting cheaper - just no longer fast enough to absorb the amount a company can spend.

Opinion: Optical stocks have a China problem that most investors are missing

Uber (UBER) reportedly exhausted its entire 2026 AI budget by April, after Claude Code spread across roughly 5,000 engineers faster than its finance models assumed. In May, Microsoft's (MSFT) Experiences and Devices division reportedly wound down most of its Claude Code licenses, also on cost grounds.

The bills are only going to get bigger. Gartner expects the cost of a single agentic workflow to rise more than fivefold through 2028. The firm calls this the inference paradox: Cheaper AI tokens invite more elaborate work, rather than smaller bills. I watch it happen on my own bills for the AI coding tools my team uses; a task that took one request a year ago now takes dozens. Two years ago, developers would ask AI a question and get an answer. Now, they hand the work to a program that breaks the task into steps, tests each one and retries what failed - routing through several models that are each billed separately. Run a few of those workflows at once and the machine never goes quiet, especially before a deadline.

There is only one way off a bill charged by the token, which is to stop buying tokens. Open-weight models are now good enough to run on on-premises machines, so you only need to pay for electricity instead.

Which brings the price of equipment back in focus. Nvidia has not yet explained why the price of a certain graphics card has jumped up 87%, but memory is the likeliest driver. Every RTX PRO 6000 Blackwell carries 96 gigabytes of high-speed memory, leaving it exposed to the industrywide memory shortage. SK Hynix's (SKHY) (KR:000660) CEO Kwak Noh-jung told Reuters that 2027 will be "the worst year in the industry's history" for supply, with demand exceeding what his company can produce beyond 2030. Nvidia is feeling it as well; CFO Colette Kress told analysts on Aug. 26 that the company is "experiencing extreme pricing conditions in memory," and that margins will bottom to about 71% by the end of 2026.

The shortage has already changed how the hardware is sold. When I ordered my server in late July, the vendor held the quoted price only on the condition that I paid before a set date, and told me plainly that missing it meant repricing.

The shortage does not spare the renter, either. Amazon buys the same graphics cards from the same supply, so the shortage reaches its customers at a higher rate.

Don't miss: Apple has gotten too predictable. Can its new CEO bring back the element of surprise?

The order books show that companies are responding. Nvidia has started reporting its data-center sales in two parts, separating the large cloud providers from everyone else. In its latest quarter, reported on Aug. 26, providers bought $49 billion worth of hardware while everyone else - from specialist AI hosts to governments and ordinary companies - bought $40 billion. Dell booked $24.4 billion in AI orders in the quarter ended May 1, leaving a $51.3 billion backlog. At HPE, orders more than doubled, and CEO Antonio Neri attributed the increase to "accelerated customer investments in agentic AI and AI inferencing," alongside private cloud adoption for AI.

The buyers are visible too. At Dell's customer conference in May, CEO Michael Dell said chief information officers are aggressively pivoting to hybrid AI, and showcased three big customers that run Dell servers on their premises: Eli Lilly (LLY) switched on LillyPod, a supercomputer built from 1,016 Nvidia Blackwell processors that it owns rather than rents; Samsung (KR:005930) uses Dell hardware in chip design; and Honeywell (HON) partnered with Dell and Nvidia last year.

A second reason to own hardware has nothing to do with cost. The large cloud providers claim they cannot satisfy the existing demand. Alphabet (GOOG) (GOOGL) CEO Sundar Pichai said Google continues to be supply-constrained, while Amazon's Andy Jassy said his company will fall short of demand in 2026 and expects the same in 2027. Nvidia CFO Kress said the same, calling the company's forecast for next year a supply-constrained outlook. When capacity is rationed, the question stops being what an hour costs and becomes whether you can get one.

The obvious objection is that cloud revenue is accelerating, rather than shrinking. AWS grew revenue 37% in the second quarter to $42.2 billion, clocking its fastest growth rate in 18 quarters, and Google Cloud grew 82% to $24.8 billion.

Both things hold, because they involve largely different customers. Amazon disclosed that Anthropic and OpenAI have each committed to multiyear, multigigawatt purchases of its Trainium chips, with Anthropic alone securing up to 5 gigawatts. A handful of labs buying at that scale can lift cloud revenue while ordinary companies move routine work onto machines they own. No provider separates the companies buying for themselves from the hosts renting the same chips out again, so the proportion is unknown from outside, and the market prices it as though the answer were zero.

The limits are worth stating. Electricity, cooling and maintenance add roughly 10% to the cost of owning, whichever way you run the machine. Staff expenses look like the bigger problem, until you remember that most midsize to large companies already employ the people who run their own data centers. The racks and the power are already there; what changes is what goes into them at the next refresh. A bank putting accelerated servers in a room it has run for years is not building anything - it is buying different hardware.

My own purchase is not the case to generalize from. I run a two-person company; the server sits in a room rather than a rack, and solar panels cover much of what it draws. Amazon also wraps its four processors in far more memory and storage. The arithmetic above is written for a company that already has the racks, the power and the people, not for a shop the size of mine. What we have in common is the calculation, not the scale.

应版权方要求,你需要登录查看该内容

免责声明:投资有风险,本文并非投资建议,以上内容不应被视为任何金融产品的购买或出售要约、建议或邀请,作者或其他用户的任何相关讨论、评论或帖子也不应被视为此类内容。本文仅供一般参考,不考虑您的个人投资目标、财务状况或需求。TTM对信息的准确性和完整性不承担任何责任或保证,投资者应自行研究并在投资前寻求专业建议。

热议股票

  1. 1
     
     
     
     
  2. 2
     
     
     
     
  3. 3
     
     
     
     
  4. 4
     
     
     
     
  5. 5
     
     
     
     
  6. 6
     
     
     
     
  7. 7
     
     
     
     
  8. 8
     
     
     
     
  9. 9
     
     
     
     
  10. 10