JPMorgan's July "Data Center Watch" showed that AI inference request prices and short-term rental prices for high-end GPUs cooled in tandem at the time.
The direct implication of this trend is that the most strained phase of AI infrastructure is beginning to ease. For ordinary users and enterprise customers, lower input token prices make invoking large models more affordable. For AI application companies and cloud providers, the decline in H100 rental rates indicates that compute resources are no longer in such short supply that everyone is waiting in line.
However, this is not a sign of declining demand. The tracked sample shows that LLM token processing volume has increased several-fold since the beginning of the year, with top models still operating at 70%-80% utilization, and high-performance GPU utilization exceeding 80%. The price cooling appears to be part of a broader catch-up by supply, open-source models, and inference efficiency to meet demand, rather than a sudden weakening of AI demand.
In July 2026, the input token prices for mainstream models continued to decline. According to the sample in this report, the input token price for some mainstream models ranged approximately between $0.2 and $5 per million tokens, with an average monthly decrease of about 3% in July.

This range requires qualification. The output token prices for high-performance proprietary models such as GPT-4o and Claude 3.5 Sonnet are typically significantly higher than their input token prices, and the pricing of lightweight models like GPT-4o mini cannot be directly compared with that of flagship models. In other words, the lower prices in the report primarily reflect input tokens, lightweight models, or the sampling methodology of hosting providers, and should not be interpreted as indicating that the full invocation costs of all mainstream models have all decreased to the same range.
Even so, the direction remains clear. For AI application companies, unit inference costs for tasks such as Q&A, code generation, search enhancement, or customer service calls are declining. Over the past few years, one of the most difficult financial calculations in large model commercialization was “the more you use it, the more you lose.” The falling price per input token has at least brought some high-frequency applications closer to an affordable cost range.
More importantly, the price reduction occurred against the backdrop of increased usage volume. From the first half of 2026 through July, the number of LLM tokens processed in the sample reports grew several-fold compared to the beginning of the year, with some niche segments seeing growth of over four times. Utilization rates for leading models from OpenAI, Anthropic, and Meta remained consistently high at 70%-80%, and the increased adoption of the open-weight Llama series models was also a major driver of rising usage volume.
The market is not seeing a price drop due to lack of usage, but rather a reduction in unit pricing driven by increased demand, combined with contributions from model service providers, open-weight ecosystems, and infrastructure supply.
The rapid growth of open-weight models was another key theme in the July token market.
The report shows that the open-weight model's token usage increased by 26% month-over-month in July and reached 63 times the level of the same period last year; its volume-weighted average price rose by 36% month-over-month, and token expenditures increased by 71% month-over-month.
This means that open-weight models are not simply squeezing proprietary models through low pricing. As their performance gradually approaches state-of-the-art levels, the invocation costs, usage volumes, and spending for some open-weight models are rising in tandem.
In July, the share of spending on open-weight models rose to 20%, up from 12% in June and 10% in May. Although closed-source models accounted for only 29% of total tokens, they still contributed 80% of spending, indicating that high-performance closed-source models continue to capture the majority of commercial value—yet open-weight models are rapidly catching up.
Based on token usage, the top five models in July were MiMo v2.5, DeepSeek v4 Flash, GLM 5.2, DeepSeek v4 Pro, and MiniMax M3, collectively accounting for 50% of total tokens.
By token expenditure, the top five are Claude Opus 4.8, Claude Opus 4.7, Fable 5, Kimi K3, and GPT-5.6 Sol, accounting for a combined 54% of total spending. Kimi K3’s entry into the top five spending models indicates that open-weight models are entering the high-value segment previously dominated by closed-source models.

For enterprise clients, the increased availability of models helps reduce dependence on a single proprietary vendor. However, according to this report, the impact of open-weight models extends beyond price reduction to include competition for higher-value use cases and budgets.
A similar structural differentiation has emerged in the GPU rental market.
In July 2026, the average rental price for H100 GPUs in the non-hyperscale cloud market was $2.68 per GPU hour, a 1.1% month-over-month decline. This marked the first monthly decrease after seven consecutive months of increases in H100 rental rates.
This decline is far below the 15% to 25% range and cannot be described as "H100 rental rates dropping by 20%." The report tracks the average monthly rental price per GPU hour, not the weekly rental price per card.
This indicates that the supply side is beginning to respond. More GPUs are entering the cloud rental market, intensifying competition among cloud service providers and computing power platforms, and short-term rental prices are no longer solely determined by extreme shortages. Compared to the phase of severe scarcity of high-end GPUs, the pressure of being unable to buy or rent, or having to wait in line for cards, has eased.
However, high-end resources remain expensive. The H100 is still significantly more expensive than the A100, and overall GPU utilization remains at 75%-90%, with high-performance GPU utilization exceeding 80%. As long as utilization stays within this range, the decline in rental prices appears more like the elimination of a portion of the scarcity premium rather than evidence of a broad oversupply of computing power.

For AI companies, this change has two layers of impact.
First, the cost of inference workloads is more likely to decrease. Inference requires continuous, stable, and low-latency computing power; a decline in GPU rental rates will improve the cost structure of API providers, AI search services, code assistants, and enterprise agent products.
Second, the cost pressure of training large models is unlikely to disappear simultaneously. High-end training depends not only on the number of GPUs but also on cluster interconnectivity, memory bandwidth, scheduling efficiency, and stable power supply. A decrease in H100 rental costs does not equate to a proportional reduction in large-scale training expenses, nor does it mean all AI startups can access equivalent cluster resources at low prices.
The least thoroughly relaxed aspect of price movement is in memory.
In May-June 2026, sample reports showed a rapid increase in DRAM spot prices, with categories such as DDR5 16Gb accumulating gains of over 90%-100%, stabilizing only by July. In contrast, HBM prices remained elevated due to demand for AI accelerators, with contract prices still lagging spot prices by one to two quarters.
For high-end AI accelerators, the GPU chip itself is not the only bottleneck—HBM, advanced packaging, and data center delivery all impact actual supply. Even if some GPU prices in the rental market have declined, upstream memory and high-end cluster delivery will still constrain the overall pace of cost reduction.

This is why "the tightest period of computing power has begun to ease" cannot be directly written as "the computing power shortage has ended." The decline in rental market prices indicates that some short-term supply pressures have been alleviated. However, high-end training and large-scale inference clusters remain constrained by memory bandwidth, advanced packaging, and the shipment pace of next-generation GPUs.
Please note the time frame. This report reflects price and utilization changes in the July 2026 sample, making it more suitable as a snapshot of easing AI compute constraints in mid-2026, rather than a comprehensive overview of prices across all cloud providers, regions, and long-term contracts.
The decline in token and GPU rental prices is beneficial for AI application companies, but not necessarily all good news for model service providers.
Lower invocation costs help stimulate demand, increase API usage, and enable more enterprises to integrate large models into real business workflows. However, if prices fall faster than inference efficiency improves, profit margins for model providers and cloud platforms will be squeezed. Especially as open-weight models close the gap and customers gain stronger bargaining power, the high-price advantage of proprietary APIs will continue to face pressure.
For hardware and cloud service providers, the decline in short-term H100 rental rates signals a normalization of the exceptionally high returns driven by extreme shortages. Actual prices received by different customers can vary significantly due to long-term contracts, regional differences, bulk volume, and variations in spot price reporting methodologies. A mere look at the weekly average rental rate cannot directly imply that all cloud providers’ revenues will decline in tandem.
The clearest signal from this price tracking is that by mid-2026, AI infrastructure supply will have caught up with part of the demand, causing per-call and short-term rental costs to begin cooling. However, the boundaries are clear: token usage is still growing rapidly, high-end GPU utilization remains high, output token prices are still expensive, and HBM remains tight. Cost reductions are first occurring in the most easily marketable segments, while the hardest to loosen remain high-end clusters and upstream memory supply.
律动 BlockBeats


