NVIDIA has aggressively disclosed real-world performance data for its new Vera Rubin platform ahead of AMD’s annual product launch, and officially revealed the full specifications of its in-house designed Vera CPU. The Vera CPU delivers nearly double the performance of x86 chips on agent-based AI tasks, with latency improved by up to six times. When running the DeepSeek R1 model, the Vera Rubin NVL72 platform achieves ten times higher token throughput per megawatt compared to the previous-generation GB200 NVL72. OpenAI, Anthropic, and SpaceX have already received their first shipments. As NVIDIA’s first self-designed server CPU, Vera consumes 250–450 watts of power, supports up to 1.5TB of memory, and is expected to ship 1.3 million units this year. This marks NVIDIA’s formal entry into the server CPU market, directly challenging AMD and Intel. The agent-based AI segment is projected to drive the server CPU market to a potential $200 billion.Article author and source: Wall Street Journal
Just before its major competitor AMD’s annual product launch, NVIDIA struck a powerful blow by aggressively releasing real-world performance data for its new Vera Rubin platform and officially disclosing the full specifications of its in-house CPU chip, Vera.
On Tuesday, NVIDIA disclosed that the Vera CPU delivers nearly double the performance of x86 chips on agent-based AI tasks, with latency reduced by up to six times. Meanwhile, early production tests by cloud partner CoreWeave showed that the Vera Rubin NVL72 platform achieves ten times higher token throughput per megawatt when running the DeepSeek R1 model compared to the previous-generation Blackwell-based GB200 NVL72 system.
The timing of the release of this data is highly significant. AMD will hold its annual product launch event, "Advancing AI," in San Francisco this Thursday. NVIDIA’s coordinated release of performance data at this moment is widely interpreted as a deliberate effort to generate market momentum. For cloud providers and enterprise customers evaluating investments in next-generation AI infrastructure, this data will directly influence their purchasing decisions.
Meanwhile, NVIDIA has officially announced that the Vera CPU has completed its first deliveries in June, with customers including OpenAI, Anthropic, and SpaceX. This marks NVIDIA’s formal entry into the CPU market, directly challenging AMD and Intel’s traditional domains.
Vera CPU unveiled: NVIDIA officially enters the server CPU market. On Tuesday, NVIDIA revealed the full specifications, benchmark results, and architecture details of its data center CPU product, Vera—key data required for potential customers to comprehensively evaluate the chip. The company stated that Vera chips have already been delivered to customers including OpenAI, Anthropic, and SpaceX in June.
Vera is NVIDIA's first server CPU designed from the ground up with a custom microarchitecture codenamed "Olympus Core," rather than using off-the-shelf designs provided by Arm.
Hannah Coutand, Head of Product Marketing for NVIDIA’s Vera, explained that the chip’s design focuses on single-core speed, high memory bandwidth, and low latency, aiming to "enable agents to return to the GPU as quickly as possible and keep the GPU consistently highly utilized." On the hardware specification level, Vera’s power consumption ranges from 250 to 450 watts, with each chip supporting up to 1.5 TB of low-power memory.
In the latest benchmark results, NVIDIA demonstrated that Vera achieves 1.9x higher performance and 6x lower latency than x86 chips on agent-based AI tasks. NVIDIA also stated that Vera outperforms AMD’s flagship EPYC Turin CPU by nearly 100% on certain industry-standard benchmarks. Previously, NVIDIA noted that Vera delivers 50% higher overall performance on AI agent tasks compared to x86 chips.
The launch of Vera is the latest step in NVIDIA’s vertical integration strategy. Wolfe Research estimates that the average selling price of each Vera chip is approximately $5,000, with an expected shipment volume of around 1.3 million units this year. Ian Buck, NVIDIA’s Vice President of Hyperscale Computing, stated that agent-based AI makes CPUs “more indispensable than ever” and predicted that the overall market size for server CPUs could eventually reach $200 billion.
Vera Rubin real-world test: CoreWeave successfully deployed and validated the industry’s first Vera Rubin NVL72 in early June, conducting a full-stack validation covering power, cooling, networking, and compute. The data disclosed represents the first publicly available real-world performance measurements of the Vera Rubin NVL72 silicon.
Testing against the DeepSeek R1 model as a baseline, under identical interaction response targets, the Vera Rubin NVL72 generates ten times the number of tokens per megawatt-second compared to the GB200 NVL72. CoreWeave also noted that optimizations made for Vera Rubin can be retroactively applied to the GB200 NVL72 system, increasing its throughput per megawatt by more than fourfold within three months. NVIDIA stated that these results were validated through over 250,000 independent configurations and more than 1.4 million GPU hours of testing.
From an architectural perspective, the Vera Rubin NVL72 rack integrates 72 Rubin GPUs and 36 Vera CPUs, connected via a 260 TB/s fully interconnected NVLink 6 architecture, with native support for NVFP4 precision. CoreWeave states that this efficiency gain enables customers to process more inference traffic within the same power budget, or run the same workload with less power, thereby reducing the cost per token.
CoreWeave also disclosed specific use cases: a global cybersecurity company expects to run threat detection inference at ten times the performance per watt; an autonomous programming agent company expects to scale agent tasks with significantly reduced token costs; and an AI-native search engine expects to expand its real-time search service to more users without exceeding response time limits.
Infrastructure Co-Optimization: Integrated Hardware and Software Unlock Greater Efficiency NVIDIA emphasizes that the aforementioned performance improvements are not solely due to the chip itself, but rather the combined result of a hardware-software co-design strategy. Through dynamic optimization of the entire infrastructure and energy stack, NVIDIA states that it can deploy up to 40% more GPUs within the same power budget.
In terms of cooling, NVIDIA employs a 45°C closed-loop liquid cooling system that saves approximately 4 million gallons of water per megawatt per year compared to standard cooling methods. At the network level, the sixth-generation NVLink 6 interconnect architecture delivers 2.3 times higher throughput for large language model inference decoding than Ethernet-based architectures; the Spectrum-X platform achieves 1.6 times faster remote direct memory access bandwidth, reduces switch count by 1.7 times, improves optical power efficiency by five times, and increases reliability by ten times.
NVIDIA’s latest Spectrum-6 platform has begun shipping to AI factory customers including CoreWeave, Microsoft, Nebius B.V., SpaceX AI Corp., and Tesla. NVIDIA is currently accelerating shipments to customers and partners such as Google Cloud, Microsoft Azure, Meta, Oracle Cloud Infrastructure, Dell Technologies, OpenAI, and CoreWeave.
NVIDIA remains a challenger in the CPU market. Despite Vera’s impressive performance metrics, NVIDIA still faces significant market share challenges in the CPU market. Gartner analyst Kevin Knox stated that AMD is currently the primary competitor in the enterprise AI server CPU segment. Reports indicate that Intel holds approximately 66.8% of the server CPU market, while AMD holds around 33%, with AMD continuously gaining market share and establishing strong partnerships with hyperscale cloud providers.
Driven by increased demand expectations for CPUs fueled by the rise of agent-based AI, AMD and Intel have seen their stock prices rise by 128% and 149% respectively this year, far outpacing NVIDIA’s approximate 8% gain over the same period, making them the top-performing chip stocks of 2026.
At this stage, NVIDIA’s list of major cloud service provider partners includes only Oracle and does not yet encompass other leading cloud vendors. Hannah Coutand acknowledged that Vera is currently in the "early adoption phase," but OpenAI plans to begin large-scale deployment of the Vera chip as early as this quarter.
Karl Freund, founder of Cambrian AI Research, summarized NVIDIA's strategic intent as: "Their goal is to free customers from dependence on Intel or AMD CPUs, and they are determined to capture this revenue. They have chosen to focus on developing a unique CPU that no one else on the market currently offers."
