AI inference drives new memory demand in the semiconductor industry

iconMetaEra
Share
AI summary iconSummary
AI and crypto news highlight rising memory demand in the semiconductor industry as AI inference grows. MetaEra reports that output tokens per query are increasing more than fivefold annually, driving KV cache and agent AI as key focus areas. Nvidia CEO Jensen Huang, at GTC Taipei 2026, stated that AI memory will reshape storage, with KV offloading and increased CPU demand as primary drivers. Nvidia introduced Dynamo and CMX, while Arm, Intel, and AMD are launching new CPUs for agent AI in 2026. Industry trends point to memory innovation as a core focus.
The arrival of the AI inference era is fundamentally reshaping the demand landscape of the semiconductor memory industry. As the average number of output tokens per query surges by more than fivefold annually, memory demands arising from KV cache management and agent AI deployment have become one of the most challenging and最具市场潜力的新兴领域 in AI infrastructure.

Article author and source: Semiconductor Industry Observer

At the GTC Taipei conference in June 2026, NVIDIA founder and CEO Jensen Huang explicitly stated that AI memory systems will fundamentally transform storage systems and identified memory systems as one of the most challenging components of AI infrastructure. This assessment directly points to two structural demand drivers: the need to offload KV caches driven by inference workloads, and the expanding demand for CPU memory fueled by the rise of agent-based AI.

The impact of these trends on the storage industry chain has already become evident. NVIDIA has successively launched the Dynamo software platform and the CMX context memory storage platform, while major chipmakers such as Arm, Intel, and AMD are set to release a new generation of CPU products targeted at agent AI in 2026. The industry is accelerating its transition from throughput-oriented architectures to low-latency-oriented architectures.

Inference-side scaling: Explosive growth in tokens is reshaping hardware requirements. The hardware demands during the AI inference phase differ fundamentally from those during the training phase.

According to NVIDIA’s public data, since the second half of 2024, the average number of output tokens per query has increased by more than fivefold annually, reaching approximately 30,000 to 40,000 tokens. This trend indicates that the industry has entered NVIDIA’s “Three Scaling Laws” phase of test-time scaling for inference.

According to TrendForce analysis, AI inference imposes three core hardware requirements: higher queries per second (QPS), longer context windows, and more inference steps and agent cycles. These three demands drive structural changes in memory requirements across three dimensions: model weights, KV cache, and agent AI.

Model weights belong to static memory allocation, and their memory footprint is directly tied to the number of model parameters, calculated as: Total model weight size = Number of parameters × Bytes per parameter. As model sizes continue to grow, this static memory requirement forms the foundational base of the inference system’s memory demand.

KV cache: Dynamic expansion drives offloading technologies and a new market for SSD PODs. The KV cache is the primary source of memory pressure during the inference phase.

KV cache stores key-value vectors generated during the prefill phase of inference to avoid redundant computations during decoding, and it involves dynamic memory allocation. Its total size is determined by the number of layers, the number of KV heads, the dimension per head, sequence length, batch size, and precision, and it grows nonlinearly with increasing dialogue length and batch size.

In long-context, high-batch inference scenarios, when the GPU's HBM capacity is insufficient, the system is forced to discard the KV cache and re-execute the prefill computation, resulting in increased latency and higher total cost of ownership (TCO).

To address this bottleneck, NVIDIA released the KV cache offloading software Dynamo in March 2025, offloading less frequently accessed KV caches to higher-capacity, lower-cost storage tiers such as CPU memory and SSDs, ensuring data remains reusable during the decoding phase.

In conjunction with Dynamo, NVIDIA launched the CMX Context Memory Storage Platform in January 2026, managed by BlueField-4 DPUs and built on BlueField-4 STX racks, with 64 BlueField-4 DPUs managing approximately 9,600 TB of capacity per rack, introducing a new Pod-level context storage layer (G3.5) between local SSDs (G3) and shared storage (G4).

Notably, the BlueField-4 DPU architecture model showcased at COMPUTEX 2026 is equipped with SK Hynix’s PEB210 E1.S and PE9010 M.2 SSD samples. As NVIDIA, Google, and other companies陆续 launch SSD POD platforms, demand in this niche market is expected to continue rising.

Agent AI: The CPU-to-GPU ratio is being restructured to 1:1, driving increased demand for LPDRAM as agent AI scales deployment, triggering another profound transformation in AI server architecture.

In AI agent workflows, the model must actively perform planning, tool invocation, decision-making, and agent operations, with all orchestration, data routing, and sub-agent evaluation tasks handled by the CPU. Huang Renxun noted that agents live in a nanosecond-scale world, where ultra-low latency is the top priority, significantly increasing the importance of CPU architecture.

TrendForce expects that as the deployment of agent AI scales up, the workload ratio between CPUs and GPUs will shift from the traditional 1:4 or 1:8 toward approximately 1:1, creating significant incremental demand for the CPU market and simultaneously driving structural growth in CPU memory demand.

NVIDIA will launch the Vera CPU in 2026, designed specifically for agent AI workloads. According to the original specifications, Vera supports up to 1.5 TB of LPDDR5X memory capacity, triple that of the previous-generation Grace CPU.

However, according to TrendForce’s latest survey, NVIDIA has decided to halve the SOCAMM memory capacity of its next-generation Vera Rubin superchip module due to insufficient LPDRAM production capacity allocated to NVIDIA in suppliers’ initial 2027 production plans; this adjustment does not reflect a decline in NVIDIA’s overall memory demand.

In the broader CPU market, 2026 is set to become the year of a comprehensive product generational shift toward agent AI. Intel has launched the Xeon 6+ (Clearwater Forest), AMD has unveiled the EPYC Venice, Arm has introduced the Arm AGI CPU, and Ampere’s AmpereOne MX is also expected to enter mass production this year. The emergence of a multi-party competitive landscape will further accelerate the release of CPU memory demand.

Two key drivers are aligning, creating a structural opportunity for the storage industry chain. Overall, AI inference is reshaping the memory demand landscape from two independent yet synergistic dimensions.

First, inference workloads are rapidly increasing the consumption of KV caches, and KV cache offloading technologies are directing large volumes of data to CPU memory and SSD PODs; as relevant platforms accelerate deployment, demand visibility in this niche market continues to rise.

Second, agent AI is pushing the CPU-to-GPU workload ratio toward 1:1, creating new incremental market opportunities for CPUs and their accompanying LPDRAM that did not exist before.

For investors in the storage supply chain, these trends indicate that enterprise SSDs, LPDRAM, and related DPU-compatible storage products are emerging as new focal points for AI infrastructure investment, beyond HBM.

Disclaimer: The information on this page may have been obtained from third parties and does not necessarily reflect the views or opinions of KuCoin. This content is provided for general informational purposes only, without any representation or warranty of any kind, nor shall it be construed as financial or investment advice. KuCoin shall not be liable for any errors or omissions, or for any outcomes resulting from the use of this information. Investments in digital assets can be risky. Please carefully evaluate the risks of a product and your risk tolerance based on your own financial circumstances. For more information, please refer to our Terms of Use and Risk Disclosure.