NVIDIA Completes Its Five-Layer AI Ecosystem to Redefine Token Economics

iconMetaEra
Share
AI summary iconSummary
NVIDIA unveils a five-layer AI ecosystem to transform token economics, integrating energy, chips, infrastructure, models, and applications. At GTC Taipei, CEO Jensen Huang promoted the concept that "Compute equals revenue," positioning NVIDIA as the architect of the AI factory. The Vera Rubin platform unifies CPU, GPU, networking, and storage, while DSX AI Factory and Dynamo enhance efficiency. New models such as Nemotron and Cosmos are designed for AI agents and physical AI, with expansion into robotics. The end-to-end solution supports ecosystem growth and aligns with emerging AI + crypto trends.
Jensen Huang introduced the concept that “computing power equals revenue” at GTC Taipei, positioning NVIDIA as the chief architect of an AI factory, evolving from a GPU giant. NVIDIA proposed the Five-Layer Cake Theory to redefine the token production system, comprising five layers from bottom to top: energy, chips, infrastructure, models, and applications. At the chip layer, the Vera Rubin platform integrates CPUs, GPUs, networking, and storage; at the infrastructure layer, NVIDIA unveiled the Vera Rubin DSX AI Factory reference design and the Dynamo open-source inference framework; at the model layer, it is advancing the Nemotron and Cosmos foundational models; and at the application layer, it has launched agent-based and physical AI solutions. Through full-stack co-design, NVIDIA creates a closed-loop ecosystem spanning from energy to applications, making it difficult for competitors to replicate—and is redefining the metrics of AI business value through token economics.

Article author and source: Semi-Industry纵横

At GTC Taipei, Jensen Huang presented an equation: “Compute equals revenue.” The underlying message is that AI companies, cloud providers, and enterprise customers are not buying hardware—they are investing in intelligent capacity that generates sustainable revenue over time.

NVIDIA is no longer just a chip company—it’s more like the general contractor of a large-scale token factory. Jensen Huang aims to be the chief architect ensuring everything runs smoothly. In this factory, competition isn’t about the capability of individual machines, but about output per unit of energy across the entire production line. Huang has provided NVIDIA’s answer through his “five-layer cake” theory: each layer, from energy to applications, has been redefined as a stage in the token production system. These five layers, from bottom to top, are: Energy, Chips, Infrastructure, Models, and Applications.

While competitors are still optimizing parameters at a single layer, NVIDIA is already optimizing and positioning every layer of the stack, leveraging multiplicative effects to leave competitors far behind.

01 Energy: AI factories begin with electricity as the foundation of computing power

From AI-driven wind and solar power forecasting to high-voltage direct current distribution systems, and on to intelligent energy storage and stable power supply, the energy layer ensures the efficient operation of AI factories and the reliable production of tokens.

The future competition among AI factories will first be a race for “how much intelligence per kilowatt-hour” can be produced. Jensen Huang introduced the “economics of the token factory,” arguing that under fixed power constraints, the core metric of competitiveness is no longer peak computing power, but “tokens per watt.” NVIDIA’s Vera Rubin high-performance computing platform has improved performance per watt by 10 times, reducing the cost per token to one-tenth. Second is the competition in energy supply capacity. In 2025, NVIDIA’s venture capital arm, NVentures, entered the energy sector for the first time, participating in a $650 million investment round in TerraPower and investing in Commonwealth Fusion Systems (CFS). Beyond exploring cutting-edge clean energy, NVIDIA has also invested in startups and established companies optimizing data center power solutions, including chips, computing, and grid management—such as Emerald AI and Utilidata. In August 2025, NVIDIA updated its official website to include its partners for 800V DC power architecture, with Chinese companies Innoscience and Megmeet being selected.

When electricity becomes a hard constraint for AI factories, keeping pace with innovation on the energy side is equivalent to securing bargaining power over upstream raw materials. NVIDIA’s investments in energy, supply chain integration, and technological empowerment ensure the stability of supply across the entire five-layer cake.

02 Chip: Vera Rubin Platform Compute Power Package

As a prime example of NVIDIA's extreme co-design, the NVIDIA Vera Rubin platform integrates the Vera CPU, Rubin GPU, NVLink 6 switch, ConnectX-9 SuperNIC, BlueField-4 DPU, Spectrum-6 Ethernet switch, and Groq 3 LPU.

NVIDIA combines these chips into a single unit because AI models are extremely sensitive to communication bandwidth and latency. NVLink 6 and CPO (co-packaged optics) switches address software-level communication bottlenecks through tight physical integration. More importantly, in the era of co-design, the cost of mixing and matching components has risen sharply.

First, let’s look at the Vera CPU, NVIDIA’s first CPU designed specifically for agent AI. The Vera CPU features 88 custom-designed Olympus cores from NVIDIA, spatial multithreading technology, and a memory subsystem with LPDDR5X bandwidth of up to 1.2 TB/s. For agent workloads, Vera completes tasks 1.8 times faster than x86 CPUs. More importantly, as part of the Vera Rubin NVL72 platform, the Vera CPU is paired with NVIDIA GPUs via NVLink-C2C interconnect technology, delivering up to 1.8 TB/s of coherent bandwidth—about seven times the bandwidth of PCIe Gen 6—enabling high-speed data sharing between CPU and GPU.

Now consider the network: NVLink 6 delivers the fast, seamless GPU-to-GPU communication required by today’s large-scale MoE models. Each GPU supports 3.6 TB/s of bandwidth, and each Vera Rubin NVL72 rack provides 260 TB/s of bandwidth. Spectrum-X is the world’s first mass-produced Ethernet silicon photonics switching platform; the next-generation Spectrum-X switches are built on CPO (Co-Packaged Optics) technology, further reducing power consumption, improving reliability, and enhancing AI productivity by integrating silicon photonic devices with the switch ASIC. The ConnectX-9 SuperNIC offers up to 1.6 Tb/s throughput, breakthrough acceleration features, and optimized network performance, delivering ultra-low-latency 800 Gb/s networking that accelerates data transfer, optimizes RoCE performance, and ensures consistent, predictable network performance for demanding AI workloads.

Finally, storage. The BlueField-4 DPU is a data center processing unit introduced by NVIDIA, part of the comprehensive BlueField platform and also a component of the Rubin platform. NVIDIA has built the BlueField-4 STX storage architecture based on the BlueField-4 DPU. The first rack-scale deployment integrates the new CMX context memory storage platform. The CMX platform treats KV Cache as a new AI-native data type, specifically designed to store and retrieve KV Cache data generated during LLM inference, making context a high-bandwidth resource shared across AI cluster systems. At GTC Taipei 2026, NVIDIA’s upgraded Vera BlueField-4 STX platform will introduce NVIDIA DOCA security libraries and microservices into the AI storage layer, helping enterprises protect data, agents, and context memory as they deploy agent-based AI into production environments.

In comparison, AMD’s Instinct MI350+EPYC disaggregated architecture follows a modular approach, but the CPU-GPU communication bandwidth bottleneck severely limits data transfer efficiency between CPU and GPU, especially in Agent scenarios requiring frequent data exchange. Meanwhile, proprietary chip teams like Google TPU, AWS Trainium, and Microsoft Maia have built their own vertically integrated closed loops, yet their scale cannot match NVIDIA’s full-stack ecosystem. When the Vera CPU is tightly coupled with the Rubin GPU via NVLink-C2C at a bandwidth of 1.8 TB/s, customers find it extremely difficult to replace Vera with an x86 CPU, as such a replacement would result in a significant drop in token generation efficiency.

03 Infrastructure: From "Buying Servers" to "Building a Factory"

Chips need to operate within AI infrastructure. Jensen Huang has designated infrastructure as a third layer, encompassing land, power, cooling, networking, and systems that orchestrate thousands of processors into a single machine. Essentially, this is what NVIDIA repeatedly refers to as the “AI factory.” NVIDIA aims to shift its customers from “buying servers” to “building AI factories.”

So how do you build an AI factory? IBM defined the mainframe standard with System/360; AWS defined the cloud architecture standard with the Well-Architected Framework. NVIDIA has defined the AI factory standard through its AI Infrastructure Hardware Guidelines and NVIDIA Dynamo. Leveraging its strong influence in intelligent computing centers, NVIDIA has released the Vera Rubin DSX AI Factory Reference Design, a set of guidelines for co-designed AI infrastructure. This reference design outlines how to design, build, and operate the entire AI factory infrastructure stack—including compute, Spectrum-X Ethernet networking, and storage—to achieve repeatable, scalable, and exceptional cluster performance. The accompanying documentation also provides industry partners with best practices for designing, building, and operating power, cooling, and control systems, enabling seamless hardware-software integration and scalable deployment. NVIDIA clearly aims for all future AI factories to adopt its standards, thereby locking in long-term adoption of its chips. Replacing NVIDIA chips isn’t merely swapping out a board—it requires dismantling the entire factory’s design logic. And the “operating system” of this factory is Dynamo. Dynamo is an open-source, low-latency, modular inference framework designed to serve generative AI models in distributed environments. It enables seamless scaling of inference workloads across large GPU clusters, delivering intelligent resource scheduling and request routing, optimized memory management, and efficient data transfer.

Model 04: Accelerating Agent Deployment to Enhance AI Factory Efficiency

NVIDIA not only provides hardware solutions but also actively invests in the model layer. Nemotron is NVIDIA’s most widely deployed family of proprietary models. Its recently launched Nemotron 3 Ultra is a mixture-of-experts model with a total of 550 billion parameters and 55 billion activated parameters per inference, capable of handling orchestration and complex reasoning tasks within autonomous workflows: making architectural decisions during extended coding sessions, synthesizing information across hundreds of research sources, and validating thousands of interdependent constraints. Cosmos is NVIDIA’s foundational model for physical AI (robotics, autonomous driving), designed primarily to generate physically accurate synthetic data and action policies, reducing industry costs for data acquisition. Cosmos 3 addresses a core challenge in physical AI: enabling robots, intelligent vehicles, or visual agents to generalize in the real world despite limited training data and fragmented simulation stacks.

For vertical industries, NVIDIA also offers specialized models, such as BioNeMo, its vertical model platform for life sciences. NVIDIA is not aiming to become the “next OpenAI”; rather, it aims to demonstrate that its hardware can run these models at lower token costs and prove that NVIDIA’s full-stack solution delivers the optimal tokens per watt. This enables customers, for whom “computing power equals revenue,” to choose NVIDIA’s hardware to deploy these proven, highly efficient models.

05 Application: The Real Market

At the top layer of the five-tier cake is the user-facing domain that generates real value, including applications of agents, enterprise automation, and physical AI in fields such as drug discovery, industrial robotics, and autonomous driving. This is the ultimate monetization layer of NVIDIA’s full-stack strategy and the output endpoint of the entire “AI factory.” Through defining agent standards, providing physical AI development platforms, and certifying industry-specific skill modules, NVIDIA enables an ecosystem of AI-powered applications to flourish within its ecosystem.

Agentic AI is NVIDIA’s current focus. Following the surge in popularity of OpenClaw, NVIDIA quickly responded by launching NemoClaw, an enterprise-grade security deployment software stack designed to complement OpenClaw with sandbox isolation, access control, and scalable operations capabilities, aiming to capitalize on OpenClaw’s momentum. To ensure security, NVIDIA also introduced Verified Agent Skills, which integrate transparency, provenance, security validation, and authenticity checks directly into the agent’s capability layer, empowering developers to expand autonomous agents with greater confidence. Physical AI is NVIDIA’s strategic lever for the future. Recently, NVIDIA partnered with Unitree Robotics and Sharpa to unveil an open humanoid robot reference design built on the NVIDIA Isaac GR00T platform. According to the official roadmap, the Isaac GR00T humanoid robot reference platform is scheduled to be officially launched by Unitree Robotics by the end of 2026.

From energy and chips to infrastructure, models, and applications, NVIDIA’s full-stack technology is driving advancement. Today, all five layers of this “cake” are in place. At the Chain Expo held in Beijing on June 22, we saw localized implementation cases in China and NVIDIA’s ecosystem of Chinese partners on its exhibition booth. In the energy layer, companies such as Xingneng Xuanguang and Energy Quantum leverage accelerated computing and digital twin technologies to advance future energy R&D, while Jinpan Technology and Chint use NVIDIA AI for end-to-end design and operations. In the chip layer, NVIDIA provides its Spectrum-X Ethernet platform to Chinese partners. In the infrastructure layer, CloudEdge Information has developed the iMDC liquid-cooled containerized AI data center solution based on NVIDIA products. In the model layer, SGLang achieves high-performance inference of DeepSeek-V4 on NVIDIA GPU servers. In the application layer, mid-range intelligent humanoid robots and Lingrui P1 industrial heavy-duty quadruped robots, built on NVIDIA solutions, were unveiled simultaneously.

06 NVIDIA plans to "sell all five layers of the cake together"

NVIDIA truly aims to build a complete closed-loop system spanning five layers. The Vera Rubin platform integrates CPU, GPU, networking, and storage into a single package; Dynamo serves as the AI factory operating system orchestrating resources across the entire ecosystem; DSX digital twins validate the flow of every watt of power before construction begins; and the Enterprise Agent Toolkit fully connects business logic into NVIDIA’s ecosystem. NVIDIA doesn’t just provide bricks—it also offers the design blueprint (DSX Blueprint), the construction crew (Dynamo orchestration), and even the laborers (NemoClaw).

Jensen Huang’s ambitions extend far beyond being a token factory in the digital world. Today, agents like OpenClaw are beginning to serve as the operating systems for agent-based computing; in the future, everything—from cloud agents to personal computers, autonomous vehicles, humanoid robots, satellites, base stations, and factory equipment—will run on agents. And the ultimate raw material for all these deployments is tokens. NVIDIA is redefining the metrics of AI business economics using tokenomics, encapsulating vertical-domain expertise into pricing-capable agent capabilities through CUDA-X skillization, and enabling full-stack co-design that makes any alternative solution extremely difficult to match in token output under the same power budget. Customers who choose NVIDIA’s full-stack solution gain proven, quantifiable intelligent output; those who opt for non-full-stack alternatives face endless challenges in system integration and optimization.

A five-layer cake, not one layer can be missing. Jensen Huang’s strategy has long surpassed chips and data centers—he is building an integrated technological and business framework.

Disclaimer: The information on this page may have been obtained from third parties and does not necessarily reflect the views or opinions of KuCoin. This content is provided for general informational purposes only, without any representation or warranty of any kind, nor shall it be construed as financial or investment advice. KuCoin shall not be liable for any errors or omissions, or for any outcomes resulting from the use of this information. Investments in digital assets can be risky. Please carefully evaluate the risks of a product and your risk tolerance based on your own financial circumstances. For more information, please refer to our Terms of Use and Risk Disclosure.