AI Token Price Collapse: Goldman Sachs Warns the Unit Economics of the AI Boom Are Breaking

Deep News
2 hours ago

The core investment narrative driving AI infrastructure is facing a severe test at the unit economics level. The speed at which Token prices are collapsing has outpaced market expectations, fundamentally shaking the revenue logic that has propped up AI sector valuations for the past two years.

According to the latest report from Rich Privorotsky, head of Goldman Sachs' One-Delta trading desk, the Silicon Data LLM Token Expenditure Index (ticker: SDLLMTK), which measures the usage-weighted average price per million tokens, fell 29% in August to roughly $0.97 at month-end. This marks an all-time low and represents a cumulative decline of over 50% from the May peak of approximately $2.05. It is the first time the index has broken below the $1-per-million-token threshold.

Meanwhile, a JPMorgan data center report shows that OpenRouter platform Token usage grew about 47% month-over-month in August, yet dollar spending increased only about 7%. This scissors effect between volume growth and price decline directly exposes the central contradiction in current AI investment logic: stock market returns are calculated in dollars, not token counts, and hyperscaler capital expenditure is based on dollar revenue expectations.

Privorotsky explicitly noted in the report that "more demand plus lower prices, if price declines outpace consumption growth, does not automatically constitute a positive."

Token Price Collapse: Not Disappearing Demand, But a Unit Price Breakdown

The Silicon Data LLM Token Expenditure Index is not a measure of Token demand volume but a usage-weighted price index. The company itself has cautioned the market that the index's decline could stem from listed price cuts, user migration to lower-cost open-source models, or both, and does not necessarily mean AI usage is contracting.

However, this is precisely the problem. OpenRouter routing volumes have surged, while H100 GPU rental prices are falling — a trend that contradicts the narrative of "exploding Token demand." Privorotsky points out that August data shows a clear pattern: volume up significantly, prices down even more, and dollar spending virtually flat.

Goldman Sachs' conclusion is that equity pricing over the past two years has been built on the assumption that "more tokens equal more revenue," and August data is the first to clearly shatter that assumption.

Competition Intensification and Local Inference: Structural Erosion of the Per-Token Pricing Model

In the report, Privorotsky openly expressed skepticism about "per-token billing" as a durable business model. The logic of pricing cloud inference per million tokens depends on three premises holding simultaneously: models are too large to run locally, users cannot replace them with open-source alternatives, and workloads are spiky enough that building in-house compute is uneconomical. These three premises are currently unraveling in tandem.

On the hardware front, RTX Spark-class laptops, DGX Spark workstations, Mac Studio systems capable of running 70-billion-parameter models, and NPUs with compute ranging from 40 to 75+ TOPS have compressed the marginal cost of a large volume of tokens to near electricity-price levels. Once a company's monthly API bill exceeds the amortized cost of a $5,000 to $15,000 device, that company transitions from being a Token consumer to a one-time hardware buyer — rather than a recurring software subscription user contributing 40% profit margins.

On the model competition front, Meta began rolling out Muse Spark 1.3 on September 2, which in independent benchmarks is now on par with GPT-5.6 Sol and Claude Opus 5, showing particular strength in agent tasks and code generation. Privorotsky noted that "frontier advantage is measured in weeks, not a moat, just a product cycle." Whenever a suboptimal model gets close enough and cheaper, enterprises bypass the premium tier, and the Token price index is the aggregate reflection of these routing decisions.

Capital Expenditure Logic Under Pressure: Credit Markets Send Early Warnings

The Goldman Sachs report points out that the current AI capital expenditure cycle is not flexible operating spending but locked-in heavy-asset commitments. According to the report, five rated hyperscalers' combined 2026 capital expenditure is projected at approximately $737 billion, accounting for roughly 38% of their revenue. Moody's has already issued warnings about free cash flow compression, balance sheet transitions from asset-light to asset-heavy structures, and lease commitments that are not reflected as bonds but substantively constrain issuers.

Credit markets have reacted ahead of equity markets. According to a tally of relevant bonds, 78 of the 91 hyperscaler bonds issued in 2026 had already fallen below their issue price by the end of August. Privorotsky describes this as "credit repricing" and notes that equity valuation multiples are a lagging indicator.

The logic chain is clear: if Token prices fall 30% while usage grows 20%, inference revenue declines; with inference revenue declining while depreciation and interest from the 2025–2027 construction cycle continue to ramp up, return on investment collapses; once ROI collapses, the market does not need a dramatic "bubble burst" narrative — it simply needs to adjust valuation multiples to reflect a utility-type enterprise holding massive assets but with constrained pricing power.

GPT-6 Astra: The Only Remaining Narrative Reversal Variable

In the report, Privorotsky identifies OpenAI's next-generation model Astra as the only near-term catalyst capable of potentially turning the situation around. On September 1, OpenAI stated that Astra had met the critical cybersecurity threshold under its Preparedness Framework, becoming the first model to be included in that category.

However, the Goldman Sachs trading desk remains cautious. A frontier model with access restrictions, required monitoring, and higher friction of use does not automatically replenish the Token revenue pool. If high-value workloads remain in controlled testing without entering public billing systems, Astra could even reduce the number of billable tokens.

The report concludes that Astra is the last near-term variable that could shift the revenue mix back toward high-value tiers. Until then, August's Token price data is the most honest signal of the current market state — demand can grow infinitely, but if the price at which that infinite demand arrives cannot cover the cost of the debt issued to build data centers, stock prices can still fall.

Disclaimer: Investing carries risk. This is not financial advice. The above content should not be regarded as an offer, recommendation, or solicitation on acquiring or disposing of any financial products, any associated discussions, comments, or posts by author or other users should not be considered as such either. It is solely for general information purpose only, which does not consider your own investment objectives, financial situations or needs. TTM assumes no responsibility or warranty for the accuracy and completeness of the information, investors should do their own research and may seek professional advice before investing.

Most Discussed

  1. 1
     
     
     
     
  2. 2
     
     
     
     
  3. 3
     
     
     
     
  4. 4
     
     
     
     
  5. 5
     
     
     
     
  6. 6
     
     
     
     
  7. 7
     
     
     
     
  8. 8
     
     
     
     
  9. 9
     
     
     
     
  10. 10