The core narrative around AI infrastructure investment is facing severe challenges at the unit economics level. Token prices have collapsed faster than expected: the Goldman Sachs Silicon Data LLM Token Expenditure Index fell 29% in August alone to approximately $0.97, dropping below the $1 per million tokens threshold for the first time. Meanwhile, usage on the OpenRouter platform increased by 47%, but dollar spending rose by only 7%. This widening gap between volume growth and price decline reveals a fundamental contradiction: returns in the stock market are measured in dollars, not token quantities. This marks a fundamental erosion of the revenue logic that supported AI sector valuations over the past two years—a warning that credit markets have already sounded.Article author and source: Wall Street Journal
The core narrative around AI infrastructure investment is facing severe challenges at the unit economics level. The pace of token price declines has surpassed market expectations, fundamentally undermining the revenue logic that supported AI sector valuations over the past two years.
According to the latest report from Rich Privorotsky, head of the Goldman Sachs One-Delta trading desk, the Silicon Data LLM Token Expenditure Index (ticker: SDLLMTK), which measures the weighted average price per million tokens used in the market, fell 29% in August alone, closing the month at approximately $0.97—a record low—and is now more than 50% below its May high of approximately $2.05. This marks the first time the index has dropped below the $1 per million tokens threshold.

Meanwhile, according to JPMorgan’s data center report, OpenRouter’s token usage increased by approximately 47% month-over-month in August, while dollar spending rose by only about 7%. This divergence between volume growth and price decline directly reveals the core contradiction in today’s AI investment logic: returns in the stock market are measured in dollars, not in token quantity, and capital expenditures for ultra-large data centers are calculated based on expected dollar revenues.
Privorotsky clearly stated in the report, "Increased demand combined with lower prices does not automatically constitute a positive signal if prices fall faster than consumption grows."
Token price collapse: It’s not that demand has vanished, but that the unit price is falling apart.
Silicon Data's LLM Token Expenditure Index is not a measure of token demand, but rather a usage-weighted price index. The company itself has alerted the market that the index's decline may result from price reductions, users migrating to lower-cost open-source models, or both—and does not indicate a contraction in AI usage.
Yet, this is precisely the issue: OpenRouter’s routing volume has surged significantly, while H100 GPU rental prices have fallen—contradicting the narrative of “surging token demand.” Privorotsky noted that August’s data reveals a clear pattern: volume rose sharply, prices dropped even more, and dollar spending remained nearly flat.
Goldman Sachs concluded that equity valuations over the past two years were based on the assumption that "more tokens equal more revenue," and August's data first clearly broke this assumption.
Intensifying competition and local inference: Structural erosion of the per-token pricing model
Privorotsky stated in his report that he is skeptical about "billing by tokens" as a sustainable business model. The logic of pricing cloud inference per million tokens relies on three assumptions holding true simultaneously: the model is too large to run locally, users cannot replace it with open-source alternatives, and the workload is sufficiently bursty that building in-house compute is uneconomical. All three assumptions are currently unraveling at once.
On the hardware side, RTX Spark-class laptops, DGX Spark workstations, Mac Studio systems running models with 70 billion parameters, and NPUs with compute power ranging from 40 to 75+ TOPS have reduced the marginal cost per token to near the level of electricity expenses. Once a company’s monthly API bill exceeds the amortized cost of a $5,000 to $15,000 device, that company shifts from being a token consumer to a one-time hardware buyer, rather than a recurring software subscriber contributing a 40% profit margin.
On the model competition front, Meta’s Muse Spark 1.3, launched on September 2, has already entered the same tier as GPT-5.6 Sol and Claude Opus 5 in independent benchmark tests, excelling particularly in agent tasks and code generation. Privorotsky noted, “Frontier advantages are measured in weeks and do not constitute a moat—only product cycles.” Whenever a slightly inferior model becomes close enough in performance and is priced lower, enterprises will bypass the premium tier, and the Token price index is precisely the aggregated reflection of this routing decision.
Capital expenditure logic is under pressure: credit markets have led the way in issuing warnings.
Goldman Sachs' report indicates that this AI capital expenditure cycle is not flexible operating spending, but rather locked-in commitments to heavy assets. According to the report, the combined capital expenditures of five rated hyperscale cloud providers are projected to reach approximately $737 billion in 2026, accounting for about 38% of their revenue. Moody’s has issued warnings regarding compressed free cash flow, a balance sheet transition from light to heavy assets, and lease commitments that, while not reflected as debt, effectively constrain issuers.
The credit market has reacted ahead of the stock market. According to a survey of relevant bonds, 78 out of 91 ultra-large-scale cloud provider bonds issued in 2026 had already fallen below their issue prices by the end of August. Privorotsky referred to this as "credit repricing" and noted that stock valuation multiples are lagging indicators.
The logic chain is clear: a 30% drop in token price and a 20% increase in usage lead to declining revenue; as depreciation and interest expenses rise during the 2025–2027 construction phase, while revenue falls, the return on investment collapses; once the return on investment collapses, the market no longer needs a dramatic “bubble burst” narrative—simply reprice the valuation multiples to reflect a utility-like enterprise with substantial assets but limited pricing power.
GPT-6 Astra: The only remaining narrative reversal variable
In the report, Privorotsky identified OpenAI’s next-generation model, Astra, as the only near-term catalyst capable of turning the tide. On September 1, OpenAI stated that Astra had met key cybersecurity thresholds under its Preparedness Framework, becoming the first model to be included in this category.
However, Goldman Sachs’ trading desk remains cautious. An aggressive model with restricted access, requiring monitoring and high usage friction does not automatically supplement the token revenue pool. If high-value workloads remain confined to controlled testing and never enter the public billing system, Astra may even reduce the number of billed tokens.
The report concludes that Astra is the last near-term variable capable of rebalancing the revenue mix toward higher-value tiers. Prior to this, August's token price data remains the most accurate indicator of current market conditions—demand can grow indefinitely, but if the price at which infinite demand materializes fails to cover the cost of debt issued to build data centers, the stock price can still decline.
