NVIDIA has restarted the AI inference prefill acceleration GPU project "Rubin CPX," with production scheduled to begin in the first quarter of 2027. Designed for long-context prefill scenarios, the product features 168GB of HBM4 memory and delivers compute performance approaching that of the standard Rubin GPU, with a single-card power consumption of 2,300 watts. It employs an independent MGX ETL rack design, supporting flexible deployment configurations ranging from 64 to 256 GPUs. The interconnect architecture uses a hierarchical design, with NVLink for intra-tray connectivity and Ethernet for inter-tray communication, balancing performance and cost. Rubin CPX will work alongside standard Rubin GPUs in a 1:1 ratio, with CPX handling prefill and KV cache generation, while Rubin GPUs handle decoding—optimized specifically for long-context inference workloads.Article author and source: Wall Street Journal
NVIDIA has quietly revived a chip project once believed by the market to have been abandoned, targeting the high-cost prefill stage of AI inference.
NVIDIA has restarted the AI inference prefill acceleration GPU project "Rubin CPX" and made significant adjustments to the product design. According to current plans, the new version of Rubin CPX is expected to begin production in the first quarter of 2027.
The restart of this project sends a clear signal: NVIDIA is intensifying efforts to address the performance and cost bottlenecks in pre-filling during AI inference. As the context length of large models continues to grow, the pre-filling stage must process vast amounts of input data and generate KV caches, whose efficiency directly impacts the overall cost and economic viability of AI inference deployment.
Guo Mingchi stated that currently more than 50% of AI inference workloads come from input context processing and KV cache construction.
The chip specifications have been significantly adjusted, bringing its computing power close to that of the standard Rubin.
In terms of compute power and power consumption, the new CPX performs close to the standard Rubin GPU, with a maximum single-card power draw of 2,300 watts. For memory, the new CPX has been upgraded to 168GB of HBM4, a significant improvement over the previous 128GB GDDR7 solution, but still lower than the standard Rubin GPU’s 288GB of HBM4.
According to Guo Mingqi, the total HBM4 capacity of the 8-GPU CPX compute tray is approximately 1.34 TB, sufficient to meet the demands of most long-context prefill workloads and their corresponding KV cache requirements. This indicates that the Rubin CPX is not merely focused on maximizing general-purpose computing power, but has been redesigned to optimize memory capacity and prefill efficiency for long-context scenarios.
Switch from shared racks to dedicated deployment for greater configuration flexibility.
The rack architecture is also a key focus of this adjustment.
In the previous plan, CPX was scheduled to share a rack with Rubin GPU; the new version instead uses a dedicated MGX ETL rack, allowing customers to independently scale CPX computing power according to their actual needs.
Specifically, customers can choose deployment scales of 64, 128, 192, or 256 CPX GPUs. Each rack module consists of 64 CPX GPUs, including eight compute trays, each equipped with eight CPX GPUs, along with one switch tray.
The modular design allows customers to flexibly increase CPX computing power based on actual pre-loaded load requirements, without having to purchase fixed, large-scale Rubin racks.
Layered interconnection design reduces pre-fill deployment costs
On the interconnected architecture, Rubin CPX employs a layered design to balance performance and cost.
Within a single compute tray, eight CPX GPUs are scaled up via NVLink. Each CPX GPU has an NVLink bandwidth of 1 to 1.5 TB/s, which is lower than the 3.6 TB/s of standard Rubin GPUs.
Within and between trays and rack modules, CPX uses Spectrum-6 Ethernet copper L1 links; inter-rack module connections are established via Spectrum-6 switches in each module using OSFP fiber links.
Compared to solutions that rely entirely on high-bandwidth GPU interconnects, this hierarchical architecture places greater emphasis on cost optimization for the pre-filling scenario.
Collaborate with Rubin GPU to target long-context reasoning.
Rubin CPX is not a standalone general-purpose GPU, but rather operates as a prefill accelerator alongside the Vera Rubin NVL72.
According to Guo Mingqi, NVIDIA recommends deploying CPX and Rubin GPUs in a 1:1 ratio: CPX handles the pre-filling computation and generates the KV Cache, which is then transmitted to the Rubin GPU via Ethernet RDMA for subsequent decoding.
From a product positioning perspective, Rubin CPX is not designed to replace the standard Rubin GPU, but rather to further divide the AI inference process, assigning prefill and decoding tasks to different GPUs.
Guo Mingqi positions it as "the best cost-performance solution for long-context prefilling." As AI model context windows continue to expand, computational demands and KV Cache sizes during the prefilling phase are steadily increasing; NVIDIA’s revival of the CPX project may aim to further reduce the cost of long-context inference through dedicated hardware and heterogeneous deployment.
