NVIDIA Vera Rubin CPU Enters Mass Production, Targets AI Infrastructure

iconMetaEra
Share
AI summary iconSummary
NVIDIA’s Vera Rubin NVL72 GPU has entered mass production, with over 300 partners and more than 350 factories across 30 countries involved. The chip features 88 Olympus cores and has been shipped to clients including OpenAI and SpaceX. As the Fear & Greed Index reflects growing market confidence, NVIDIA is rolling out full-rack systems powered by its custom CPU to compete with Intel and AMD. Altcoins to watch may experience increased movement as demand for AI infrastructure continues to rise.
NVIDIA's Vera Rubin NVL72 has entered mass production ramp-up, with a global supply chain spanning over 350 factories in 30 countries and involving more than 300 partners. The Vera chips have already been delivered to customers including OpenAI, Anthropic, and SpaceX, with an expected shipment of 1.3 million units this year and an average price of approximately $5,000 per chip. NVIDIA is extending its focus from the GPU market to the CPU market, betting that AI agents will create a new "CPU moment," with the long-term server CPU market potentially reaching $200 billion. The Vera chip features 88 Olympus cores, prioritizing single-threaded performance over core count, achieving up to 1.8 times the speed of x86 CPUs on specific tasks. NVIDIA is no longer selling chips alone but instead offering complete server rack systems incorporating its proprietary chips, directly challenging AMD and Intel’s core interests in the server CPU market.

Article author and source: Tencent News

NVIDIA is extending its rivalry from the GPU market into the CPU market.

Ahead of AMD's AI-related event, NVIDIA disclosed on Tuesday, Eastern Time, the latest progress on its next-generation Vera Rubin platform: the Vera Rubin NVL72 has entered mass production ramp-up, with NVIDIA stating that the platform’s supply chain spans more than 350 factories across 30 countries and involves over 300 partners.

Meanwhile, NVIDIA further disclosed performance data for the Vera CPU and Vera Rubin platform under AI agent workloads, aiming to demonstrate that as AI shifts from “answering questions” to autonomously planning, invoking tools, and executing tasks, CPUs are becoming the new frontier of AI infrastructure.

The complete specifications, benchmark results, and architectural details of NVIDIA’s data center CPU product, Vera, are essential for potential customers to fully evaluate the chip. The company stated that the Vera chip has already been delivered to customers such as OpenAI, Anthropic, and SpaceX in June.

Vera's launch is the latest step in NVIDIA's vertical integration strategy. The company aims to sell complete server rack systems featuring its own chips to customers, rather than selling individual chips alone. According to Wolfe Research, the average price per Vera chip is approximately $5,000, with an estimated shipment volume of around 1.3 million units this year. For AMD and Intel, NVIDIA's move directly threatens their core interests in the server CPU market.

This is not the first time NVIDIA has announced that the Vera CPU has entered mass production. At the end of May this year, NVIDIA announced that Vera had entered full-scale production and stated that the chip can perform specific tasks 1.8 times faster than traditional x86 CPUs. In mid-May, NVIDIA also delivered the first Vera CPU systems to customers such as Anthropic, OpenAI, and SpaceX AI. This announcement further emphasizes the scaling of mass production for the entire Vera Rubin platform, customer deployments, and real-world performance in production environments.

The rise of AI agents has made the CPU a critical bottleneck once again.

NVIDIA's core rationale for betting on CPUs is that AI applications are evolving.

The primary task of traditional generative AI is to generate responses based on user input, with GPUs handling large-scale model computations and CPUs primarily managing peripheral scheduling tasks. However, as AI agents rapidly evolve, AI systems are now autonomously breaking down tasks, invoking external tools, executing code, accessing data, and iteratively evaluating results.

During this process, the CPU must frequently handle a large volume of low-latency, real-time tasks. NVIDIA believes that AI agents are not solely GPU-dependent workloads: each agent’s execution environment, tool invocation, task orchestration, and long-context data retrieval all require CPU involvement.

When NVIDIA previously introduced Vera, it noted that agent-based AI is creating a new "CPU moment." The company reasoned that as AI systems shift from "answering questions" to "taking actions," the role of the CPU within the AI factory will become significantly more prominent.

NVIDIA even expects the long-term size of the server CPU market to reach $200 billion. This figure appears extremely aggressive compared to the traditional server CPU market, but NVIDIA’s reasoning is that the future market boundaries for CPUs may extend beyond conventional enterprise servers to include AI inference, AI agents, reinforcement learning, data processing, and various control and orchestration tasks within AI infrastructure.

This is also why NVIDIA is trying to redefine the rules of CPU competition.

NVIDIA bets on "single-core performance," not core count.

Vera's design roadmap differs significantly from that of traditional server CPUs.

NVIDIA stated that Vera is its first CPU designed specifically for agent-based AI, featuring 88 custom-designed Olympus cores and 1.2 TB/s of memory bandwidth. NVIDIA previously indicated that Vera delivers a 50% improvement in single-core performance compared to traditional CPUs; in tests released at the end of May, Vera completed specific tasks 1.8 times faster than x86 CPUs.

In the latest disclosure, NVIDIA further emphasized Vera's core design priorities: rather than simply adding more cores, Vera places greater emphasis on single-threaded performance, inter-core communication bandwidth, and memory access latency.

NVIDIA stated that Vera’s custom Olympus core delivers 2x single-threaded performance, 3x higher inter-core bandwidth, and 40% lower memory latency compared to competitive chiplet designs. Its goal is to enable AI agents to complete tasks faster and return computational resources to the GPU sooner, thereby improving overall utilization of the AI factory.

The underlying business logic is straightforward: GPUs are among the most expensive and scarce computing resources in AI data centers. If CPUs process tasks too slowly, GPUs may sit idle. NVIDIA aims to reduce this "idle time" by specifically optimizing CPUs.

In production environment tests conducted by DeepInfra on the Vera CPU, NVIDIA stated that Vera can support up to 1.6 times more concurrent AI agents and achieve up to 2.2 times faster task orchestration. Please note that these figures are derived from specific customer test environments and do not imply that Vera outperforms all general-purpose CPU workloads.

The competition has escalated from a single chip to an entire "AI factory."

Vera CPU is just one part of NVIDIA's broader system strategy.

On the Vera Rubin platform, NVIDIA integrates the Vera CPU with the Rubin GPU, Groq 3 LPX, Spectrum-6 networking chips, ConnectX-9 SuperNIC, and BlueField-4 into a co-designed system.

NVIDIA stated that the Vera Rubin NVL72 is currently ramping up to mass production globally, with partners such as CoreWeave, Google Cloud, Microsoft Azure, and Oracle Cloud Infrastructure already beginning to deploy related racks. NVIDIA also noted that its supply chain spans more than 350 factories across 30 countries, involving over 300 partner companies.

In customer testing, CoreWeave reported that after running DeepSeek-R1 on the Vera Rubin NVL72, it achieved a tenfold increase in tokens generated per megawatt per second compared to the previous-generation Grace Blackwell NVL72. Google Cloud has already launched A5X instances based on the Vera Rubin NVL72, with NVIDIA stating that, in specific scenarios, it enables lower inference costs and higher token throughput per megawatt.

This means NVIDIA's competitors are no longer just single accelerator chips like AMD's Instinct GPUs or Intel's Gaudi.

NVIDIA aims to control the entire infrastructure: CPUs handle scheduling, GPUs handle computation, networking chips enable high-speed interconnects, and software manages scheduling and optimization, delivering systems at the rack or even data center level to customers.

For AMD and Intel, the threat of this competitive approach is that NVIDIA may not need to compete head-on with them in the traditional CPU market. Instead, NVIDIA aims to enter through the most high-growth segment of AI servers—targeting workloads such as agent-based AI, reinforcement learning, and high-performance AI inference—before gradually expanding the scope of its CPU applications.

AMD and Intel's traditional advantages face new challenges.

For a long time, Intel and AMD have dominated the server CPU market.

However, NVIDIA now believes that AI infrastructure is reshaping the value hierarchy of CPUs. In the past, competition among server CPUs focused primarily on core count, general-purpose computing power, and overall throughput; in the era of agent-based AI, the importance of low latency, single-threaded performance, memory bandwidth, and task orchestration efficiency is rising.

This is precisely the market gap NVIDIA believes it can fill.

From a product perspective, Vera can be used as a standalone CPU or combined with NVIDIA GPUs to form the Vera Rubin system. NVIDIA also positions it as a key component connecting AI agents with GPU computing resources.

NVIDIA’s previously disclosed customer list already includes AI companies such as Anthropic, OpenAI, and SpaceX AI, as well as cloud service providers like ByteDance, CoreWeave, and Oracle Cloud Infrastructure; server manufacturers including Dell, Hewlett Packard Enterprise, Lenovo, and AMD are also advancing systems based on Vera.

This gives NVIDIA an advantage over traditional CPU manufacturers: it can provide customers with complete solutions through a combination of GPUs, CPUs, networking, and software.

But this does not mean NVIDIA can easily challenge AMD and Intel.

Competition in the CPU market extends beyond chip performance to include software ecosystems, customer certifications, server compatibility, and long-term supply relationships. Large cloud service providers and enterprise customers typically require several years to transition their hardware platforms.

Therefore, whether Vera can truly expand from AI-specific use cases to the broader server market depends on whether customers can achieve a clearly noticeable advantage in performance and cost.

NVIDIA is opening a new revenue curve beyond GPUs.

NVIDIA's coordinated release of production and performance details for Vera Rubin just ahead of AMD's AI event also carries clear competitive market implications.

Over the past few years, NVIDIA has almost monopolized the most profitable segment of AI infrastructure through its GPUs. However, as AMD continues to expand its AI accelerator offerings and Intel seeks to break into the AI chip market, attention is increasingly turning to whether NVIDIA can sustain its high growth rate.

Meanwhile, AMD and Intel's CPU businesses have regained market attention due to new demand expectations driven by agent-based AI.

NVIDIA's response is: If the value of AI data centers is expanding from single-GPU computing to complete AI factories, NVIDIA does not need to cede the CPU market to its competitors.

From this perspective, Vera is not merely a server CPU launched by NVIDIA, but rather a further extension of its vertical integration strategy.

NVIDIA wants customers to purchase not individual GPUs, CPUs, or network chips, but an entire infrastructure capable of running AI models and AI agents directly.

If agent-based AI truly becomes the core driver of the next wave of computing demand, NVIDIA may be competing not just for market share in the GPU sector, but for the entire AI server computing value chain.

This also means that AMD and Intel will soon face not just each other, but NVIDIA—a company that already possesses capabilities in GPUs, CPUs, networking, software, and complete system solutions.

Disclaimer: The information on this page may have been obtained from third parties and does not necessarily reflect the views or opinions of KuCoin. This content is provided for general informational purposes only, without any representation or warranty of any kind, nor shall it be construed as financial or investment advice. KuCoin shall not be liable for any errors or omissions, or for any outcomes resulting from the use of this information. Investments in digital assets can be risky. Please carefully evaluate the risks of a product and your risk tolerance based on your own financial circumstances. For more information, please refer to our Terms of Use and Risk Disclosure.