NVIDIA Unveils Vera Rubin NVL72, Increases DeepSeek Throughput 30x

icon MarsBit
Share
AI summary iconSummary
NVIDIA released on-chain news showcasing Vera Rubin NVL72 test results, with DeepSeek-V4-Pro throughput increasing 30x. Token costs decreased 35x compared to GB300 NVL72. The platform includes Vera CPU and Groq 3 LPX, now being utilized by SpaceXAI. New token listings may follow, as the performance gains suggest broader AI deployment.

That’s insane.

Just now, NVIDIA unveiled, for the first time in its history, on-chip test data for its next-generation flagship cabinet, Vera Rubin NVL72.

Moreover, this real-world test directly attracted DeepSeek -V4-Pro, running the most authentic "agent coding" tasks!

The results are astonishing: compared to the current flagship GB300 NVL72, Vera Rubin’s throughput per megawatt has surged by up to 30 times, and the token cost has dropped by as much as 35 times!

Agent

Someone exclaimed: "I wouldn't even dare write a pitch deck this bold to deceive investors!"

Another netizen commented wryly: “Previously, I thought $30,000 for an H200 was expensive, but now looking back, the H200 doesn’t even qualify as the entry-level model.”

Agent

Let’s take a close look.

From H200 to GB300, throughput for real-world agent workloads increased by up to 30x, with costs approximately doubling.

From GB300 to Vera Rubin, throughput has surged by up to 30 times, and the price per rack has roughly doubled again.

According to Huang’s Law, NVIDIA’s biggest technological breakthrough over the past two years hasn’t been the GPU itself, but rather achieving a 900-fold speed increase while raising prices fourfold!

Agent

This time, NVIDIA announced to everyone a counterintuitive truth: the LLM era is over, the Agent era has arrived, and all previous AI benchmark scores are obsolete!

Agent

Meanwhile, the Vera CPU, designed specifically for agents, has been installed overnight by SpaceX AI, owned by Musk, and is even being prepared for launch into space.

Also today, NVIDIA's Groq 3 LPX has entered full-scale production.

Gemma 4 31B runs on top, delivering an incredible speed of up to 3,400 tokens per second. The throughput of a trillion-parameter model surges 35 times.

Agent

Agent

These new records will rewrite the commercial landscape of the Agent ecosystem starting tonight!

Why must NVIDIA release Vera Rubin?

The latest data from OpenRouter provides the answer: in the real world, an Agent AI task consumes 15 times more tokens than a standard chat conversation!

For example, if an agent is tasked with researching a company for investment decisions, the agent and sub-agents continuously reason, and the accumulated tokens become input for the next step—making long-context processing the most critical limiting factor for AI agents.

Agent

The more agent interactions, the higher the throughput—when the number of agents increased tenfold, tool calls doubled, creating new demands driven by the tension between computing power and algorithms.

Agent

Don't confuse chat with agents—the old benchmarks are obsolete!

At this point, traditional AI benchmarks (such as fixed-length 8K/1K sequence tests) become completely ineffective.

NVIDIA officially states: Performance measurement must evolve! We can no longer measure only single inference requests—we must capture entire agent workflows.

To this end, they used the AgentX benchmark from SemiAnalysis.

This test is no longer a rigid Q&A format, but rather a replay of real-world "coding sessions" featuring contextual growth, tool usage, and sub-agent generation.

Agent

This is why the H200 struggles in the new agent battlefield, and Vera Rubin is destined to reign supreme.

Vera Rubin NVL72 × DeepSeek-V4-Pro:

The hash power beast has arrived

To demonstrate the strength of the Vera Rubin NVL72, NVIDIA directly used the open-source champion— DeepSeek Conducted on-chip testing for V4-Pro (1.6T).

As soon as the data was released, the results were astonishing—

Under AgentX workloads, Vera Rubin NVL72 achieves up to 30 times higher throughput per megawatt compared to GB300 NVL72!

Agent

Please note that the comparison here is not with the outdated H200, but with the highly sought-after flagship GB300.

GB300 at the same DeepSeek During the V4-Pro test, the throughput per megawatt has improved by 15 times compared to the H200.

Yet Vera Rubin raised the bar another 30 times on the shoulders of GB300, further elevating the entire Pareto curve!

What does this mean?

For AI factories constrained by power supply, this directly expands the boundaries of physical laws.

Under the same power budget, Vera Rubin can do 30 times the work of conventional agents.

For megawatt- or even gigawatt-scale data centers, this is equivalent to magically creating 30 times the computing power assets.

Moreover, NVIDIA DSX MaxLPS technology enables power management at the GPU, rack, and workload levels, allowing up to 40% more GPUs to be deployed within the same megawatt budget, further enhancing the AI factory’s throughput per megawatt!

Agent

Token costs drop 35-fold! The "ledger" of the Agent industry is completely rewritten.

Per megawatt of throughput, the direct impact is the cost of generating each token. The surge in performance leads to the most immediate consequence: a nuclear explosion in the business model.

NVIDIA announced that the Vera Rubin NVL72 reduces the cost per million tokens by up to 35 times compared to the GB300 NVL72!

Agent

When inference costs plummet by 35-fold, 24/7 digital employees will become a reality, and consumer-facing super apps will experience a massive surge.

Old Huang’s move has directly opened the door to the era of agent applications.

Agent

Ultimate collaborative design: How did Huang achieve this?

You might ask: How can it be 30 times faster? Is it due to the process technology of a single chip?

Clearly not.

The remarkable performance improvement of Vera Rubin NVL72 is achieved through "extreme co-design."

Agent

For example, segregated services, distributed KV caching, KV-aware routing, MegaMoE, and more.

Agent

Additionally, NVFP4 quantization directly compresses model weights to 4-bit precision, significantly reducing memory usage without sacrificing output quality, thereby dramatically boosting throughput.

And sixth-generation NVLink, combined with large MoE models, provides an interconnect network that is 10 times faster and offers three times lower latency than standard Ethernet, enabling DeepSeek This large model based on the MoE architecture can seamlessly invoke different expert subnetworks across 72 GPUs.

This is no longer about selling graphics cards; Huang is selling an entire AI power plant.

Agent

NVIDIA Groq 3 LPX enters full-scale production:

3,400 tokens per second—Code Agent, times are changing!

At the Hot Chips 2026 conference, NVIDIA also announced a stunning development: Groq 3 LPX has entered full-scale production!

Agent

Groq 3 LPX is an exclusive expansion for the Vera Rubin NVL72 data center platform.

It is specifically customized for this system, with a focus on low-latency inference acceleration.

This black technology was acquired by NVIDIA in December last year for $20 billion from the startup Groq.

Agent

AI agents may encounter decoding delays; to eliminate this pain point, Groq 3 LPX seamlessly separates large-context processing from token generation.

NVIDIA's solution is to have the Rubin GPU handle large-context tasks and delegate the ultra-fast token generation to the Groq LPX.

One handles reading, the other handles writing—each focuses on what they do best.

Agent

In a full rack-scale deployment, up to 256 LP30 accelerators can work in tandem with GPUs via ultra-high-bandwidth interconnects to form an enterprise-grade inference engine.

Agent

As a result, Groq 3 LPX directly achieved a record-breaking output speed.

In Artificial Analysis's benchmark, Groq 3 LPX ran Gemma 4 31B with a 100,000-token long context window, achieving an astonishing output speed of 3,400 tokens per second!

This makes it four times faster in response time than competitors for workloads that are extremely sensitive to latency.

Agent

Multi-step agent tasks that used to take hours to complete can now be finished in just minutes!

Agent

Groq 3 LPX reduces the time to generate 5,000 tokens from 50 seconds to 1.5 seconds, a 34-fold improvement.

Agent

In coding tasks, the Groq 3 LPX achieves a maximum output speed of 6,981 tokens per second.

Agent

Jensen Huang stated that advancing ultra-fast token generation through LPX has achieved "another massive leap" in AI throughput and responsiveness.

Agent

Nebius has deployed this chip in the Nebius Token Factory to provide an unparalleled experience of instant responses at every step of the agent loop.

Agent

Following closely is Groq itself.

Agent

The first CPU designed specifically for Agents is released,

Musk rushed into space overnight!

The most ambitious move in this official announcement is the Vera CPU.

For the Agent, NVIDIA specially designed a CPU.

The Vera CPU is specifically designed to power agents.

It is equipped with 88 proprietary Olympus cores and high-bandwidth LPDDR5X memory, delivering a direct bandwidth of up to 1.2 TB/s.

Agent

Why is a completely new CPU needed in the era of large models?

At Hot Chips, NVIDIA's Vera CPU executives stated directly, "Agentic AI is the most complex computing task ever."

Agent

A single task requires the Agent to perform hundreds of steps behind the scenes: invoking tools, executing Python code, navigating context, and processing massive amounts of data...

All these courier scheduling tasks are forced onto the CPU to handle alone. Therefore, NVIDIA must create a dedicated CPU for the Agent.

Precisely because of this, Musk's SpaceXAI announced overnight the official large-scale deployment of NVIDIA's Vera CPU!

As early as May this year, NVIDIA quietly shipped the first Vera samples to Company A, OpenAI, and SpaceX AI for initial testing.

Just three months later, SpaceXAI eagerly moved it into full deployment.

Even more astonishing, SpaceXAI is building a gigawatt-scale computing facility based on the Vera Rubin platform to power Grok.

Moreover, this computing power will be sent directly into space.

Musk stated, "Two industry leaders have partnered to design an optimized version of the Vera Rubin NVL72, which will be deployed at scale by 2028."

Agent

The first-generation "Starmind" AI satellite from SpaceXAI is planned to use the Vera Rubin NVL72 rack-level system.

Agent

From a single GPU to an entire Agent factory

On the same day, two major chips, Vera CPU and Groq 3 LPX, entered full-scale production.

This is significant for NVIDIA.

Agent

In the era of large models, NVIDIA has secured dominance in training and inference through its GPUs.

In the agent era, its reach has expanded to CPUs, inference accelerators, NVLink, and more.

What Lao Huang sells is no longer just a GPU.

He is selling an AI factory capable of continuously producing tokens.

This time, Lao Huang is rebuilding the foundation for the entire "Agent Era."

Reference materials:

https://x.com/MinLiBuilds/status/2091915873661686204

https://developer.nvidia.com/blog/nvidia-vera-rubin-and-blackwell-set-a-new-standard-for-agentic-ai-performance-per-watt/

https://developer.nvidia.com/blog/inside-nvidia-groq-3-lpx-the-low-latency-inference-accelerator-for-the-nvidia-vera-rubin-platform/

https://developer.nvidia.com/blog/maximizing-ai-factory-performance-per-watt-with-nvidia-dsx-maxlps/

This article is from the WeChat public account "New Intelligence Yuan," authored by ASI Revelation.

Disclaimer: The information on this page may have been obtained from third parties and does not necessarily reflect the views or opinions of KuCoin. This content is provided for general informational purposes only, without any representation or warranty of any kind, nor shall it be construed as financial or investment advice. KuCoin shall not be liable for any errors or omissions, or for any outcomes resulting from the use of this information. Investments in digital assets can be risky. Please carefully evaluate the risks of a product and your risk tolerance based on your own financial circumstances. For more information, please refer to our Terms of Use and Risk Disclosure.