Organized & Compiled by Shenchao TechFlow

Guest: Jensen Huang, Founder and CEO of NVIDIA
Host: Dwarkesh Patel
Podcast source: Dwarkesh Patel
Jensen Huang – Will Nvidia’s moat persist?
Broadcast date: April 16, 2026
Key Points Summary
This article explores issues such as the potential impact of TPUs (Tensor Processing Units) on Nvidia’s dominance in AI computing, the company’s control over advanced chip supply chains, whether AI chips should be sold to China, why Nvidia does not directly become a hyperscaler, and its choices in chip architecture and investment strategy, through a conversation with Nvidia CEO Jensen Huang.

Key Insights Summary
Regarding NVIDIA's core essence and moat
- Ultimately, something must always convert electrons into tokens, and this conversion process—along with making tokens grow more valuable over time—is not easily commoditized. Our job is to do as many necessary things as possible and as few unnecessary things as possible to achieve this conversion.
- Why are they (the supply chain) willing to invest in me and not others? Because they know NVIDIA has the capacity to absorb their supply and sell it through our downstream channels. NVIDIA’s downstream supply chain and demand scale are so massive that they’re willing to bet on the upstream.
- All bottlenecks last no more than two or three years, without exception. What I’m concerned about are downstream issues—such as policies that hinder energy development. Without energy, you cannot build an industry; without energy, you cannot advance reindustrialization.
The Competition Between CUDA and ASICs
- NVIDIA does accelerated computing, not tensor processing units. Accelerated computing encompasses a much broader range... We are the only company that accelerates a wide variety of applications and has a vast ecosystem.
- NVIDIA's GPUs are accelerators, more like F1 race cars. Anyone can drive them up to 100 km/h without issue, but pushing them to their limits requires considerable expertise. We extensively use AI to generate our own kernels.
- NVIDIA's computing stack is the best value in the world—without exception. No other platform has demonstrated to me a better performance-to-total-cost-of-ownership ratio today.
About Company Philosophy and Investment
- We should do as many necessary things as possible, and as few unnecessary things as possible. This means: the work we do in building our computing platform, if we didn’t do it, I truly believe no one else would. But the world doesn’t lack cloud services—if we didn’t do it, someone else would step in.
- My mistake was not fully realizing they had no other options—venture capital will never invest $50 to $100 billion in an AI lab. That was my blind spot. If I could rewind time, and if NVIDIA had been as large as it is today back then, I would have been happy to do so.
Regarding the Chinese market and geopolitics (most watched section)
- What’s the best way to create a secure world? Labeling them as victims or turning them into enemies is probably not the best answer. They are competitors, and we want the United States to win, but I believe maintaining dialogue and continuing research exchanges may be the safest approach.
- This loser mentality, this loser premise, means nothing to me. I cannot accept the idea of abandoning a market based on the premises you’ve described. It makes no sense because I don’t believe America is a loser, and neither is our industry.
- When you have abundant energy, it can compensate for insufficient chips. If your electricity is completely plentiful and nearly free, why care about performance per watt? Just stack older chips—you don’t need anything more. A 7nm chip is essentially equivalent to the Hopper generation and is more than sufficient.
- The day DeepSeek was launched on Huawei chips was a bad outcome for our country... If all AI models perform best on someone else's tech stack, you need to explain—can you really say that’s good for the United States?
On Software, Agents, and the Future
- Today, we are limited by the number of engineers; tomorrow, these engineers will be supported by a large number of agents, allowing us to explore the design space in ways never seen before... tool usage will cause software companies to scale rapidly—the reason this hasn’t happened yet is that agents are not yet skilled enough at using tools.
- One thing you can count on from NVIDIA: This year, Vera Rubin will be outstanding; next year, Vera Rubin Ultra will arrive; the year after, Feynman will come. Every year, you can rely on us as reliably as you rely on a clock.
About Architecture and Workloads
- We could do it this way (develop multiple architectures), but we don’t have a better idea. We’ve run all of them in the simulator, and the results were worse, so we won’t go that route.
- In the past, higher throughput was always better. But we believe there may be a scenario where tokens with very high individual value can compensate for lower throughput, making the trade-off worthwhile. This is why we decided to expand the Pareto frontier.
Is Nvidia's supply chain advantage its greatest moat?
Host Dwarkesh: We’ve seen a sharp decline in valuations across a wide range of software companies, as markets anticipate that AI will commoditize software. There’s a seemingly straightforward logic here: NVIDIA sends GDS2 files to TSMC, which manufactures logic chips and switches, then combines them with HBM from SK Hynix, Micron, and Samsung, and ships them to Taiwanese ODMs for final system assembly. From this perspective, NVIDIA is essentially a software company—with others handling its manufacturing. If software becomes commoditized, could NVIDIA become commoditized too?
Jensen Huang:
Ultimately, something must always convert electrons into tokens, and this conversion process—along with making tokens grow more valuable over time—is not easily commoditized. Making one token more valuable than another is like making one molecule more valuable than another—it involves the art, engineering, science, and invention that we are witnessing unfold in real time. The transformation, processing, and underlying science required to create a token are far from fully understood, and this journey is far from over. I don’t believe it will be commoditized. The framework you described is essentially my own mental model of the company: input is electrons, output is tokens, and in between is NVIDIA. Our job is to do as many necessary things as possible—and as few unnecessary things as possible—to achieve this conversion. “As few as possible” means: if I don’t need to do it myself, I bring in a partner and make it part of my ecosystem. If you look at NVIDIA today, we likely have the world’s largest partner ecosystem—encompassing the entire upstream supply chain, all downstream computer manufacturers, application developers, and model creators.
AI is a five-layer cake (application, model, infrastructure, chip, energy), and we have ecosystems at every layer. We try to do as little as possible, but the part we must do happens to be extremely difficult. I don’t think this will be commoditized; in fact, I don’t believe enterprise software companies or tool vendors will be commoditized either. Today, most software companies are essentially tool manufacturers, but what I see is the opposite of what most people see: the number of AI agents will grow exponentially, and so will the number of tool users. The number of instances using Synopsys Design Compiler will likely surge dramatically, as will the number of agents using layout tools and design rule checkers. Today, we’re limited by the number of engineers; tomorrow, those engineers will be supported by a vast army of agents, enabling us to explore design space in ways never before possible—all using today’s existing tools. Tool usage will cause software companies to scale at an unprecedented rate; the reason this hasn’t happened yet is that agents aren’t yet skilled enough to use these tools effectively. Either these companies will build their own agents, or agents will become powerful enough to use these tools—I believe both will happen.
Host Dwarkesh: In your latest earnings report, you have nearly $100 billion in purchase commitments for foundries, memory, and packaging. SemiAnalysis reported that this figure could be as high as $250 billion. One interpretation is: NVIDIA’s real moat is securing scarce manufacturing capacity for years to come—others may have accelerators too, but can they get memory? Can they get logic chips? Is this NVIDIA’s biggest moat over the next several years?
Jensen Huang:
This is one of the things we can do that others find extremely difficult. We’ve made tremendous commitments upstream—some explicit, as you mentioned, and others implicit. For instance, many upstream investments are initiated by our supply chain partners because I tell their CEOs: “Let me show you how massive this industry will become, let me explain why, let me walk you through the projections, and let me show you what I see.” Through this ongoing process of informing, inspiring, and collaborating with CEOs, they become willing to invest. Why do they invest in us and not others? Because they know NVIDIA has the capacity to absorb their output and sell it through our downstream channels. The scale of NVIDIA’s downstream supply chain and demand is so immense that they’re willing to bet on the upstream. Look at GTC—people are amazed by its scale and the caliber of participants; the entire AI universe is there, fully covered in 360 degrees, because they all need to meet each other. I bring them together so the downstream can see the upstream, the upstream can see the downstream, and everyone can witness the latest advancements in AI. Crucially, they can see firsthand what AI-native companies and AI startups are building—validating everything I’ve been telling them. I’ve spent enormous time, directly and indirectly, helping our supply chain, partners, and ecosystem understand the opportunity before us.
Our keynote always includes a segment that feels a bit heavy, like a lecture—yes, that’s exactly my intention. I need to ensure that everyone across the entire supply chain, from upstream to downstream, understands what’s coming, why it’s coming, when it’s coming, and how large the scale will be—and can think through it systematically, just as I do. Regarding the moat you mentioned, we are building for the future. If the industry’s scale over the next few years reaches the trillions of dollars, we have the supply chain capacity to support it. Without our coverage and reach, our business velocity—like cash flow, supply chains also have flow and turnover—no one would build a supply chain for an architecture with low turnover. We can sustain this scale only because our downstream demand is so strong; they’ve seen it, heard it, and it’s all right in front of them.
Host Dwarkesh: I’d like to get more specific about whether upstream supply can keep up. For years, you’ve been doubling annually, with computing power supply increasing by more than threefold each year—doubling at this scale is truly astonishing. You’re the largest customer for TSMC’s N3 node and one of the biggest customers for N2. This year, AI is expected to account for 60% of N3 capacity, and according to SemiAnalysis projections, that will rise to 86% next year. How do you double again when you’re already the majority—and do it year after year?
Jensen Huang:
To some extent, global demand following trends exceeds the supply from upstream and downstream. This is exactly the kind of industry you want—where instantaneous demand surpasses the total industry supply. The opposite, of course, is undesirable: if the gap becomes too large and supply at a specific stage becomes critically insufficient, the entire industry will rush to resolve it. For example, you now hear far less discussion about CoWoS because over the past two years, we aggressively tackled this bottleneck, increasing capacity several-fold. Today, TSMC understands: CoWoS capacity must keep pace with the rhythm of logic and memory demand, and they are expanding CoWoS and future packaging technologies at the same rate as logic. For a long time, CoWoS and HBM were seen as niche products—but now they are no longer niche; people recognize them as mainstream computing technologies, and we are now able to influence a much broader segment of the supply chain. When the AI revolution first began, I was saying exactly what I’m saying today—five years ago. Some believed and invested—like Sanjay and the Micron team—I still clearly remember that meeting, where I laid out exactly what would happen, why it would happen, and my predictions. They truly bet on it, and we deeply collaborated on LPDDR and HBM memory; they made substantial investments, and the results have been excellent for them. Others came later, but now they’ve all caught up. Every bottleneck is receiving intense attention, and now we’re anticipating bottlenecks years in advance. For instance, over the past few years, our collaboration with Lumentum, Coherent, and the entire silicon photonics ecosystem has completely reshaped the supply chain. We’ve built an entire supply chain around TSMC, partnered with them on COUPE, invented an entire suite of technologies, and licensed our patents to the supply chain to keep the ecosystem open. We’ve prepared the supply chain for scaling challenges by inventing new technologies and workflows, developing novel test equipment like double-sided probing, and investing in companies to help them expand capacity.
If we discourage people from becoming software engineers, we’ll face a shortage of software engineers. The same prediction happened ten years ago. Some doomsayers back then said, “Whatever you do, don’t become a radiologist.” You can still find those videos online claiming radiology would be the first profession to disappear and that the world no longer needed radiologists. Guess what we’re short on now? Radiologists.
Host Dwarkesh: How did you achieve a doubling of logic manufacturing each year? Ultimately, both memory and logic are constrained by EUV.
Jensen Huang:
None of these are impossible to scale quickly in the short term—they can all be resolved within two to three years, requiring only a demand signal. These things are not difficult to replicate. As for how far down the supply chain I need to go—some require direct negotiations, some indirect, and others: if I can convince TSMC, ASML will naturally follow. We need to focus on critical nodes, but if TSMC is convinced, you’ll have enough EUV machines within a few years. The point is: none of these bottlenecks last longer than two or three years—not one of them. Meanwhile, each generation of computing efficiency improves by 10 to 20 times; from Hopper to Blackwell, it’s 30 to 50 times. We continuously invent new algorithms using CUDA’s flexibility to drive efficiency gains while scaling capacity. I’m not worried about any of this. What I am concerned about are downstream issues—such as policies that hinder energy development. Without energy, you cannot build an industry; without energy, you cannot advance reindustrialization, bring chip manufacturing, computer manufacturing, and packaging back to the U.S., build EVs and robots, or construct AI factories. More chip capacity is a two- to three-year problem; more CoWoS capacity is also a two- to three-year problem.
Will TPU disrupt Nvidia's dominance in AI computing?
Host Dwarkesh: If you look at TPUs, one perspective is that two of the top three models globally—Claude and Gemini—were trained on TPUs. What does this mean for NVIDIA?
Jensen Huang:
We are doing something entirely different. NVIDIA does accelerated computing, not tensor processing units. Accelerated computing is used across a wide range of fields: molecular dynamics, quantum chromodynamics, data processing, data frames, structured data, unstructured data, fluid dynamics, and particle physics—alongside AI. Accelerated computing encompasses a far broader scope; although everyone is talking about AI today—and AI is indeed critically important—the scope of computing extends far beyond AI. NVIDIA has redefined computing by shifting from general-purpose computing to accelerated computing. Our market reach far exceeds what any TPU or ASIC could possibly achieve. Look at our market position—we are the only company accelerating applications across the board, with a vast ecosystem where every major framework and algorithm runs on NVIDIA. Because our systems are designed for external operators, any operator can purchase our systems. Most in-house systems must be operated internally because they were never designed to be flexible enough for others to run. Since anyone can operate our systems, we are present on every cloud—including Google, Amazon, Azure, and OCI. If you want to profit by leasing, you need a large customer ecosystem spanning multiple industries. If you want to use it internally, we can help you too—just as we helped Elon’s xAI. And because we enable operators across any company or industry to deploy our systems, you can use them to build supercomputers for Lilly’s scientific research and drug discovery, helping them run their own supercomputers to meet diverse needs across drug discovery and biomedical science. There are far too many applications that TPUs cannot address. NVIDIA has turned CUDA into an exceptional tensor processing unit—while also handling every stage of the data lifecycle: data processing, computation, and AI—all of it. Our market opportunity is larger and more comprehensive. Because we support every application in the world, whenever you deploy a NVIDIA system anywhere, you know there will be customers using it—that’s an entirely different proposition.
Host Dwarkesh: This is a long question. You’re generating astonishing revenue—$60 billion per quarter—not from pharmaceuticals or quantum computing, but because AI is an unprecedented technology growing at an unprecedented pace. The question is: What is optimal for AI itself?
I discussed this with my AI researcher friend, who said that when they use a TPU, it’s a massive systolic array specifically optimized for matrix multiplication, whereas GPUs are more flexible and better suited for scenarios with heavy branching or irregular memory access. But AI is essentially nothing but highly predictable matrix multiplication, repeated over and over. TPUs don’t need to sacrifice chip area for warp schedulers or switching between threads and memory banks—they’re deeply optimized for the primary use case driving today’s computational growth. What do you think?
Jensen Huang:
Matrix multiplication is a crucial component of AI, but not the whole story. If you want to invent new attention mechanisms, perform slicing differently, or design an entirely new architecture from scratch—such as a hybrid SSM—you need a general-purpose programmable architecture. If you want to create a model that fuses diffusion and autoregressive techniques, you同样 need a general-purpose programmable architecture. We can run anything you can imagine—that’s our advantage. Programmability makes inventing new algorithms vastly easier, and it’s precisely the invention of new algorithms that drives AI’s rapid progress. TPU and everything else are subject to Moore’s Law—Moore’s Law delivers roughly 25% improvement per year. To achieve true 10x or 100x leaps, the only way is to fundamentally transform algorithms and computation every year. This is NVIDIA’s core strength. When I first announced that Blackwell’s energy efficiency would be 35 times greater than Hopper’s, no one believed me. Then Dylan wrote an article saying I undershot—it’s actually 50 times, something impossible to achieve through Moore’s Law alone. Our solution lies in novel model designs—such as MoE, parallelization across computing systems, slicing, and distributed deployment. Without CUDA to generate new kernels, I wouldn’t even know where to begin. This is the combination of a programmable architecture with NVIDIA’s extreme co-design capabilities. We can offload portions of computation onto the NVLink fabric or to the Spectrum-X network layer. We can drive innovation simultaneously across processors, systems, fabrics, libraries, and algorithms. Without CUDA, I have no idea where to start.
Host Dwarkesh: You generate 60% of your revenue from the five largest cloud providers. In earlier customer phases—such as professors running experiments—they need CUDA and cannot use other accelerators; they must rely on PyTorch and CUDA. But these hyperscalers have the resources to write their own kernels—in fact, they must do so to extract the final 5% of performance on specific architectures. Anthropic and Google mostly run on their own accelerators, such as TPUs and Trainium, while OpenAI, even when using GPUs, employs Triton because they need their own kernels, all the way down to CUDA C++, bypassing cuBLAS and NCCL to build their own stack—and they can even compile it to other accelerators.
If most of your customers can substitute CUDA to some extent, is CUDA still the decisive factor enabling cutting-edge AI to occur on NVIDIA?
Jensen Huang:
CUDA is an incredibly rich ecosystem. If you want to build on any architecture, starting with CUDA is a wise choice because the ecosystem is so extensive—we support every framework. If you want to create custom kernels, we’ve made substantial contributions to Triton—the backend of Triton is filled with NVIDIA technology. We’re eager to help every framework become as excellent as possible. There are many frameworks: Triton, vLLM, SGLang, and now an explosion of post-training and reinforcement learning frameworks like verl and NeMo RL. Across the entire domain of post-training and reinforcement learning, everything is growing explosively. So if you’re building on any architecture, choosing CUDA is the smartest decision because you know the ecosystem is rich, and when issues arise, it’s far more likely to be your code than the underlying mountain of code. When something goes wrong—is it your problem or the machine’s? You want it to always be your problem; you want to trust the machine. Of course, we still have our own bugs, but our system has been refined enough that you can at least build on top of it with confidence. Second, as a developer building software anywhere, the most important thing is having a massive installed base so your software can run on a vast number of other machines. You’re not writing software for yourself—you’re writing it for your entire fleet, or everyone else’s fleets, because you’re a framework builder. The NVIDIA CUDA ecosystem is the true core treasure. We now have hundreds of millions of GPUs deployed out in the wild—from A10, A100, H100, H200, to the L-series and P-series, in every size and form factor. If you’re a robotics company, you want that CUDA tech stack to run directly on the robot itself—we’re everywhere. A massive installed base means that once you develop software or a model, it’s useful everywhere. This value is irreplaceable. Finally, the fact that we’re present on every cloud makes us truly unique. If you’re an AI company or developer and aren’t sure which cloud provider to choose—or where you want to run—you can run anywhere, including your own private deployments. The combination of ecosystem richness, widespread installed base, and diverse deployment options makes CUDA uniquely valuable.
Host Dwarkesh: But I’m curious whether these advantages are equally important for your core customers? Many large clients are capable of building their own software stacks. Especially as AI becomes increasingly powerful in domains with strict verification loops, the question becomes: Can all hyperscalers build their own kernels? NVIDIA still offers excellent value, so they may still prefer NVIDIA—but the question is, will this eventually become a pure competition based on who has the best specs and the most compute and memory bandwidth per dollar? Historically, NVIDIA’s gross margins have exceeded 70% precisely because of the CUDA moat. If most customers can effectively replace CUDA, can those gross margins still be sustained?
Jensen Huang:
We assign an astonishing number of engineers to these AI labs, dedicated solely to collaborating with and optimizing their tech stacks. The reason is simple: no one understands our architecture better than we do. Our architectures aren’t generic like CPUs—CPUs are like Cadillacs: smooth, luxurious sedans that aren’t fast but are easy for anyone to drive, with cruise control and everything just works. NVIDIA’s GPUs, by contrast, are accelerators—more like F1 race cars. Anyone can drive one to 100 km/h, but pushing it to its limits requires deep expertise. We heavily use AI to generate our own kernels. I’m confident that for a very long time, we’ll remain indispensable. Our expertise helps our AI lab partners extract up to 2x more performance from their stacks—often with remarkable ease. After optimizing a specific kernel or the entire stack, it’s routine for their models to accelerate by 2x, 3x, or even 50%. That’s a massive number, especially when you consider the massive installed base of Hopper and Blackwell systems they already have—doubling efficiency directly translates to doubled revenue. NVIDIA’s compute stack is the most cost-effective in the world—without exception. No platform has ever demonstrated to me a better performance-per-total-cost-of-ownership ratio today. None. The InferenceMAX benchmark is out there for anyone to run—TPUs aren’t coming, Trainium isn’t coming. I welcome them to use InferenceMAX to demonstrate their claimed 40% cost advantage—no one wants to. Second, you say 60% of customers are from the top five—but most of those are external customers: the vast majority of NVIDIA on AWS serves external clients, not AWS itself; all Azure customers are external; OCI too. They prefer us because our reach is unparalleled—we bring top-tier customers from every industry. These customers are built on NVIDIA because our coverage and flexibility are unmatched. The flywheel works like this: installed base, architectural programmability, rich ecosystem, combined with tens of thousands of AI companies worldwide. If you’re one of those AI startups, which architecture would you choose? The richest one. The one with the largest installed base. The one with the most extensive ecosystem. That’s how the flywheel spins—plus the highest revenue-per-unit-compute and highest performance-per-watt. If your partner builds a gigawatt-scale data center, and that data center must generate the most tokens and the highest revenue, we have the world’s highest tokens-per-watt architecture. If your goal is to rent out infrastructure, we have the largest customer base globally. That’s why the flywheel keeps turning.
Host Dwarkesh: Anthropic just announced a multi-gigawatt TPU order in collaboration with Broadcom and Google, and most of their compute power comes from TPUs. When I look at these large AI companies, it seems like a significant portion of compute is shifting away from NVIDIA. How do you explain this trend?
Jensen Huang:
Anthropic is an exception, not a trend. Where would TPU growth come from without Anthropic? It’s 100% from Anthropic. Where would Trainium growth come from without Anthropic? Again, it’s 100% from Anthropic. This is widely known and generally understood. It doesn’t mean there’s a flood of ASIC opportunities—there’s only one Anthropic.
I don’t feel threatened when others try different things or explore new directions. If they don’t try alternatives, how would they know how good we are? Sometimes you need a reference point to remind yourself. We must continuously earn our market position. There’s always a lot of hype—just look at how many ASIC projects have been canceled. Building something better than NVIDIA isn’t easy. What would have to go wrong at NVIDIA for people to find this plausible? Because of our scale, our pace of iteration, we are the only company in the world making major strides every single year.
ASIC profit margins are quite high—NVIDIA’s is 70%, and ASIC’s is 65%. From what I can see, ASIC’s gross margin is very strong, and they themselves recognize and take pride in their healthy ASIC gross margins. So, you ask why this is the case. Long ago, we simply didn’t have the capacity to do so. At the time, I didn’t fully appreciate how difficult it was to build foundational AI labs like OpenAI and Anthropic, or how massive the capital investments required from suppliers themselves would be. We weren’t able to invest billions of dollars into Anthropic in exchange for them using our compute. Google and AWS had that capability—they made substantial early investments in Anthropic, which is why Anthropic ended up using their compute. I didn’t have that capacity back then. My mistake was failing to fully grasp that these labs had no other option—venture capital would never invest $5 to $10 billion into an AI lab expecting it to become today’s Anthropic. That was my blind spot. But even if I had understood it at the time, we still may not have had the resources—but I won’t make the same mistake again. I’m thrilled to have invested in OpenAI and to have helped them scale; I believe it was the right thing to do. When we had the capacity and Anthropic came to us, I was delighted to become an investor and help them grow. We just weren’t able to do it back then. If I could rewind time, and if NVIDIA had been as large as it is today back then, I would have been more than happy to do exactly that.
Why doesn't Nvidia become a hyperscale cloud provider directly?
Host Dwarkesh: For years, NVIDIA has been the company making the most money in AI, earning enormous amounts. Now you’ve started investing—reports say you’ve invested $30 billion in OpenAI and $10 billion in Anthropic. Their valuations have surged dramatically and are sure to keep rising. After all those years of supplying them with computing power, you saw the direction clearly—their value was just one-tenth of what it is today a year ago, and even then, you had substantial cash on hand.
One idea is that NVIDIA could either establish its own foundational model lab or make investments like these at a much lower valuation. You had the cash—why not do it sooner?
Jensen Huang:
We did what we could when we could. If I had known earlier, I would have acted. At the time Anthropic needed us to move, we simply didn’t have the conditions, nor the awareness or mindset to do so.
It was a matter of investment scale. At the time, we had never made investments outside the company, and certainly not at such a large amount. We didn’t realize we needed to do so. I always assumed they could raise funds from VCs like any other company. But what they wanted to accomplish couldn’t be achieved through VCs. What OpenAI wanted to do also couldn’t be achieved through VCs. That’s their insight—that’s what made them smart—they recognized early on that they had to take a different path. I’m glad they did. Even though Anthropic had to turn to others because we weren’t involved, I’m still thrilled that it all happened. Anthropic’s existence is a positive thing for the world, and I’m delighted by it.
Host Dwarkesh: You’re still making huge profits, and earning even more quarter after quarter. The question now is: what should NVIDIA do with all this money? One answer is: there’s now an entire ecosystem of intermediaries that convert CapEx into OpEx, offering compute rental services to these labs—because chips are expensive, but as AI models improve, the value they generate over their lifetime keeps growing. NVIDIA has the cash to make CapEx investments; you’ve reportedly backed CoreWeave with up to $6.3 billion and invested $2 billion. So why doesn’t NVIDIA just build its own cloud services and become a hyperscaler, renting out that compute directly? You have more than enough cash.
Jensen Huang:
This is a matter of company philosophy, and I believe it is wise: we should do as many necessary things as possible and as few unnecessary things as possible. This means: the work we do in building our computing platform—if we didn’t do it, I truly believe no one else would. If we hadn’t taken the risks we’ve taken—had we not built NVLink our way, not constructed the entire technology stack, not cultivated the ecosystem our way, not focused on CUDA for 20 years, most of which was unprofitable—if we hadn’t done it, no one else would. If we hadn’t created all those CUDA-X domain-specific libraries—for ray tracing, image generation, early AI work, whether in data processing, structured data handling, or vector data processing—if we hadn’t done it, no one else would. For example, we created cuLitho for computational lithography—if we hadn’t, no one else would have.
So accelerated computing wouldn’t have developed the way it has today if not for us doing it—and therefore, we should do it. We should go all in and concentrate all our energy on making it happen. But the world doesn’t lack cloud services—if we didn’t do it, someone else would. From the philosophy of “do as much as is necessary, and as little as is unnecessary,” if we hadn’t supported CoreWeave, these emerging AI clouds wouldn’t exist. Without helping CoreWeave, they wouldn’t be here. Without supporting Nscale, they wouldn’t be where they are today. Without supporting Nebius, they wouldn’t be where they are today. Now all of them are thriving. Is this a business model—do as much as is necessary, and as little as is unnecessary? That’s why we invest in the ecosystem: because I want our ecosystem to flourish. I want this architecture and AI to connect with as many industries and as many countries as possible, so the world can be built on AI and on the U.S. tech stack—that’s the vision we’re pursuing.
When NVIDIA was first founded, there were 60 3D graphics companies, and we were the least likely to survive. At the time, NVIDIA’s graphics architecture was fundamentally wrong—not slightly off, but in the exact opposite direction; developers simply couldn’t support it. Starting from sound first principles, we derived an incorrect conclusion—anyone would have crossed us off the list. Yet here we are today, so I have enough humility to recognize this: don’t bet on winners. Let them decide for themselves, or give everyone a chance.
Host Dwarkesh: But you listed a whole bunch of emerging clouds that wouldn’t exist without NVIDIA—how does that align with “not picking winners”?
Jensen Huang:
First, they must seek us out for help in order to exist. When they have a vision, a business plan, professional expertise, and passion—of course, they must also possess inherent strength—but if, in the end, they need some investment to get started, we will be there to support them. Once their flywheel starts turning, we want them to become independent as quickly as possible. Your question is, “Do we want to be in the funding business?”—the answer is no. There are specialized firms that do funding; we’d rather collaborate with all of these funders than become one ourselves. Our goal is to focus on what we do best, keep our business model as simple as possible, and support our ecosystem. When a company like OpenAI needs $30 billion in investment before going public—and we firmly believe they will become an extraordinary company—they already are, and they will one day be an incredible one, the world needs them to exist, the world wants them to exist, I want them to exist, their wind is blowing—support them, let them grow. We will make such investments because they need us to. But we’re not trying to do as many things as possible; we’re trying to do as few as possible.
Host Dwarkesh: The topic of GPU allocation—we’ve been in a state of GPU shortage. NVIDIA is reportedly allocating its scarce capacity through a system that isn’t purely auction-based, ensuring that emerging cloud companies like CoreWeave, Crusoe, and Lambda receive some supply. Why is this beneficial for NVIDIA? First, do you agree with this characterization of “dividing up market share”?
Jensen Huang:
No, your assumption is incorrect—we’ve thought very carefully about these matters. First, if you don’t place a purchase order (PO), no amount of discussion matters. What can we do before a PO is placed? So our first priority is to work closely with everyone on forecasting, because these items take a long time to build, and data centers also require significant time to construct. We align supply and demand through forecasting—that’s our top priority. Second, we strive to involve as many parties as possible in forecasting, but ultimately, an order must be placed.
Perhaps for some reason you didn’t place your order in time—what can I do? First come, first served. If you’re not ready—your data center isn’t prepared, or certain components aren’t available to enable deployment—we may decide to serve another customer instead, simply to maximize our factory’s production efficiency and make necessary adjustments. Beyond that, the priority principle is strictly first come, first served, and only a confirmed PO counts. The stories about Larry and Elon begging me for GPUs over dinner? Those are completely untrue. We did have dinner—it was a great meal—but they never begged for GPUs; they simply needed to place an order. Once an order is placed, we do everything we can to secure capacity. We keep it simple.
Host Dwarkesh: So there’s a queue that determines when you receive your hardware based on whether your data center is ready and when you placed your purchase order. But it doesn’t seem to be strictly based on the highest bidder. Why not do it that way?
Jensen Huang:
We never do this, and we never will. It’s poor business practice—you set a price, and people decide whether to buy or not. I’m aware that other companies in the chip industry adjust prices during periods of high demand, but we don’t. This has never been our approach. You can count on us. I’d rather be the dependable infrastructure of this industry—no guessing involved. I give you a price, and that’s the price, no matter how high demand gets. Our relationship with TSMC works the same way: there’s no legal contract, sometimes I benefit, sometimes I lose out, but overall, the relationship is excellent—I fully trust them and rely on them completely. One thing you can always count on with NVIDIA: this year, Vera Rubin will be outstanding; next year, Vera Rubin Ultra will arrive; the year after, Feynman will come. Every year, you can count on us. Go anywhere in the world and find another ASIC team that can say, “I’ll stake my entire business on you—I’m confident you’ll be here every year, driving down token costs by an order of magnitude each year, and I can trust you as reliably as I trust a clock.” You simply cannot say that about any other foundry. But with NVIDIA, you can say it today. Want to buy $1 billion worth of AI computing power? No problem. $100 million? No problem. $10 million—a single rack, a single GPU? No problem. A $100 billion order? Also no problem—we are the only company in the world that lets you say that. The same applies to our relationship with TSMC: whether you buy one chip or a billion, it’s all possible—as long as you follow the process and do all the right things properly. NVIDIA’s position as the global infrastructure for AI was built over decades, requiring immense commitment, dedication, and unwavering stability and consistency from the company.
Should we sell AI chips to China?
Host Dwarkesh: Anthropic just released a preview of Claude Mythos a few days ago—a model they aren’t even publicly releasing, because they say it has such powerful cyberattack capabilities that the world isn’t ready for it until those zero-day vulnerabilities are patched. They claim it discovered thousands of critical vulnerabilities in every major operating system and browser, even finding one in OpenBSD—a system specifically designed to have no zero-day vulnerabilities—that had existed for 27 years.
If Chinese companies, laboratories, and government entities possess AI chips to train models with cyberattack capabilities, such as Claude Mythos, and run millions of instances, would that pose a threat to U.S. companies and national security?
Jensen Huang:
First, Mythos was trained on relatively ordinary computing power and at a relatively ordinary scale, by an extraordinary company. China has abundant computing power and the computational resources required for training, so you must first recognize: chips exist in China—they manufacture 60% or more of the world’s mainstream chips, and this is a massive industry for them. They have some of the world’s top computer scientists, and the majority of AI researchers in these AI labs are of Chinese descent, accounting for 50% of the world’s AI researchers. So the question is: given the assets they already possess—abundant energy, ample chips, and a vast number of AI researchers—if you are concerned about them, what is the best way to create a safer world? Labeling them as victims or turning them into enemies is likely not the best answer. They are competitors, and we want the U.S. to win—but I believe maintaining dialogue and preserving research exchange may be the safest approach. This is an area severely lacking due to our current attitude toward China as a competitor. Communication among AI researchers must be preserved; we must strive to jointly agree on what uses AI should never be permitted for. As for discovering software vulnerabilities, that is precisely what AI should be doing. Will AI find many bugs in software? Of course—it will, because there are countless bugs, and even AI software itself contains massive numbers of bugs. This is exactly what AI was meant to do, and I am delighted that AI has reached this level of capability.
One underappreciated aspect is the entire ecosystem surrounding cybersecurity, including AI cybersecurity, AI safety, and AI privacy. A large number of AI startups are building this future for us: an incredibly capable AI agent surrounded by thousands of other AI agents protecting its security and safe operation. This future is inevitable. The idea of an AI agent roaming freely without any oversight is inherently absurd. We clearly understand that this ecosystem must thrive—and it requires open source, open models, and open technology stacks so that all these AI researchers and outstanding computer scientists can build AI systems that ensure AI safety. Therefore, one thing we must ensure is maintaining the vitality of the open-source ecosystem. This cannot be overlooked, and a significant portion of open-source contributions comes from China. We certainly want the United States to have as much computing power as possible; we are constrained by energy, but many are working to solve this issue, and we cannot let energy become a bottleneck for our nation. At the same time, we want all global AI developers to build on the U.S. technology stack, ensuring that AI advancements—especially open-source progress—are accessible to the U.S. ecosystem. If we end up with two separate ecosystems—one open-source ecosystem running solely on foreign technology stacks and one closed ecosystem running on U.S. stacks—that would be a terrible outcome for the United States, and I will not allow it to happen.
Host Dwarkesh: So the core concern, returning to the gap in computing power and the hacking issue, is this: estimates suggest that, due to their use of 7nm technology without EUV—export controls limiting access to lithography machines—their available computing power is only about one-tenth of America’s. Could they eventually train a model like Mythos? Yes. But because we have more computing power, U.S. labs will achieve it first. Because Anthropic got there first, they said, “We’ll hold it back for a month to let all U.S. companies patch the vulnerabilities before releasing it.” And even if they do train such a model, the ability to deploy at scale is critical—if you have a network of hackers, having a million is far more dangerous than having a thousand. So inference computing power is extremely important. In fact, what’s truly concerning is that they have so many outstanding AI researchers—this is precisely why computing power makes these engineer-researchers so much more effective.
If you speak with any AI lab in the United States, they all say the bottleneck is computing power. Wouldn’t it be better if U.S. companies, with greater computing power, reached the Mythos capability level before China, giving society more time to prepare?
Jensen Huang:
We should always be first, and we should always have more. But to achieve the outcome you describe, you must go to extremes—they must have absolutely no hashing power. If they have any hashing power at all, the question becomes how much. China possesses enormous hashing power; you’re talking about the world’s second-largest computing market. If they want to consolidate hashing power, they have more than enough to do so.
AI is a parallel computing problem, right? So why can’t they simply connect four times or ten times as many chips together? They have abundant energy. They have completely vacant data centers with power already connected. China has ghost cities and abandoned data centers, and they have massive infrastructure capacity. If they wanted to, they just need to stack more chips together—even 7nm ones. Their ability to manufacture chips is among the largest in the world; the entire semiconductor industry knows they dominate mainstream chip production, and their capacity is excessive, more than enough. So the idea that China lacks AI chips is completely absurd.
Of course, if you ask me whether the U.S. would go further if there were no computing power worldwide, that’s not a plausible scenario—it’s not real. They already have ample computing power. The threshold you’re concerned about, they’ve not only crossed but surpassed. I think you misunderstand AI as a five-layer cake, with energy at the bottom. When you have abundant energy, it can compensate for insufficient chips; when you have abundant chips, it can compensate for insufficient energy. For example, the U.S. has limited energy, which is precisely why NVIDIA must continuously advance our architecture and pursue extreme co-design, ensuring that every batch of chips we deliver delivers exceptional throughput per watt—because energy is so constrained. But if your wattage is completely abundant, nearly free, why care about performance per watt? You simply stack older chips. A 7nm chip is essentially equivalent to the Hopper generation. Most models today were trained on Hopper-class hardware—7nm chips are more than sufficient. Abundant energy is their advantage. But what if they could manufacture enough chips? Huawei just had its largest single-year volume in company history, shipping millions of chips—millions, far more than Anthropic owns. Some question how much logic and memory SMIC can produce—I’m telling you the reality: they have ample logic and ample HBM2 memory.
Host Dwarkesh: But the bottleneck in training and inferencing these models is often memory bandwidth. If they’re using HBM2, there could be nearly an order of magnitude difference in memory bandwidth compared to your latest products—that’s a massive gap.
Jensen Huang:
Huawei is a networking company, and you don’t need EUV to produce the most advanced HBM—that’s simply incorrect. You can connect them together, just as we do with NVL72. They’ve already demonstrated connecting all their computing power into a single massive supercomputer using silicon photonics. The truth is, their AI development is progressing quite smoothly. The world’s best AI researchers, constrained by compute, have devised incredibly clever algorithms. Remember what I just said: Moore’s Law improves performance by about 25% per year, but through excellent computer science, algorithmic efficiency can improve by another 10x.
True leverage lies in excellent computer science. MoE is a brilliant invention, without a doubt. Various outstanding attention mechanisms have reduced computational demands. We must acknowledge that most advances in AI come from algorithmic progress, not just raw hardware. Now, if most progress stems from algorithms, computer science, and programming, tell me—aren’t their large teams of AI researchers their most fundamental advantage? We’ve already seen that DeepSeek represents a significant advancement that cannot be ignored. The day DeepSeek was released on Huawei chips was a poor outcome for our country.
Host Dwarkesh: Currently, DeepSeek can run on any accelerator because it is open-source. Why won’t this continue in the future?
Jensen Huang:
If it’s optimized for Huawei’s architecture, our architecture will be at a disadvantage. You described a situation that I think is good news—a company develops software and AI models that perform best on the U.S. tech stack; that’s good news. But you framed it as bad news. Here’s the bad news version: Global AI models are being developed and perform best on non-U.S. hardware—that’s bad news for us.
You take a model optimized for NVIDIA and try running it on other GPUs—they don’t perform better. NVIDIA’s success is proof enough. It’s not hard to understand that AI models are created and run best on our tech stack. Go to the Global South, go to the Middle East—out of the box, if all AI models performed best on someone else’s tech stack, you’d have to explain: is that really good for the United States?
Host Dwarkesh: Suppose Chinese companies reach the next Mythos before the U.S., discovering all the security vulnerabilities in U.S. software—but they do so on NVIDIA chips—and then release the model to the Global South. Running on NVIDIA chips—what’s the advantage there?
Jensen Huang:
That's not good, indeed it's not good, so we can't let this happen.
Host Dwarkesh: Why do you think Huawei will precisely fill this gap if we don’t sell them chips? Aren’t they behind in chips?
Jensen Huang:
The evidence is right in front of us: their chip industry is massive, they can use twice as many chips, they have abundant, readily available energy to power them, and they excel at manufacturing. I believe they will ultimately outproduce everyone else. If these are critical years, then during this period, we must ensure that all global AI models are built on U.S. technology stacks.
Why would you have one layer of the AI industry lose its entire market to benefit another layer? All five layers need to succeed; each one is important.
Host Dwarkesh: Stepping back, China is currently stuck at 7nm, while you’ll keep advancing to 3nm, 2nm, and even 1.6nm Feynman. By the time you reach 1.6nm, they’ll still be at 7nm—even years later. Don’t you think that your growing compute power, combined with their abundant energy, effectively means they’re continuously gaining more compute?
Jensen Huang:
We should always be first and always have more. But to achieve the outcome you described, they must have no hashing power at all. If they have any hashing power, the question becomes how much.
The U.S. has 100 times more computing power than any other place in the world, and the U.S. should be leading. NVIDIA ensures that American labs are always the first to hear about new advancements and the first to have the opportunity to purchase them. If they don’t have enough money, we even invest in them. The U.S. should be leading, and we are doing everything we can to ensure the U.S. stays ahead—that’s the number one priority, do you agree? We are doing everything possible. But how does selling chips to China, enabling them to break through their computing bottlenecks, help the U.S. maintain its lead?
We have prepared Vera Rubin for the United States. First, why don’t we propose a more balanced regulatory approach that enables NVIDIA to win globally, rather than ceding the entire world? Why would you want the United States to give up the global market? The chip industry is part of the U.S. ecosystem, part of U.S. technological leadership, part of the AI ecosystem, and part of U.S. AI leadership. Why does your policy thinking lead the United States to abandon the majority of the global market?
Host Dwarkesh: If that computing power can run a model capable of launching zero-day attacks on U.S. software, how can you say it’s not a weapon?
Jensen Huang:
First, the way to address this issue is to engage in dialogue with researchers, with China, and with all nations to ensure that people do not use technology in that manner. This dialogue must happen. Second, we must also ensure that the United States leads, with Vera Rubin and Blackwell being abundantly available in the U.S., in vast quantities. The U.S. has outstanding AI researchers who are in strong condition, and we must maintain our leadership. However, we must also recognize that AI is not just a model—AI is a five-layer cake, and every layer of the industry is critical. We want the United States to win at every layer, including the chip layer. Abandoning the entire market will not enable the U.S. to win the long-term technological competition, neither at the chip layer nor at the computing stack layer. This is a fact.
Host Dwarkesh: So how does selling chips to them help us win in the long term? Tesla has been selling excellent electric vehicles to China for a long time, and iPhones are sold there too, but neither created lock-in—China’s EVs now dominate, and their smartphones are also dominating.
Jensen Huang:
When we began today’s conversation, you used the term "moat" to describe NVIDIA’s position. For our company, the single most important factor is the richness of our ecosystem, which centers on developers. Fifty percent of AI developers are in China, and the U.S. should not abandon these individuals.
We must continue to innovate—our market share is growing, not shrinking. The assumption that even if we compete in the Chinese market, we’d just hand it over to others? You’re not speaking to someone who sees themselves as a loser. That loser mentality, those loser assumptions, mean nothing to me. We are not a car, something you can easily switch brands from one day to the next. It doesn’t work that way. x86 exists for a reason, and ARM is so hard to replace for a reason. These ecosystems are incredibly difficult to displace—they require immense time and effort, and most people simply don’t want to undertake that. Our job is to keep nurturing that ecosystem, advancing technology, and staying competitive in the market. I cannot accept the idea of abandoning a market based on the premises you’ve described. It makes no sense, because I don’t see America as a loser, and our industry isn’t a loser either. That loser mindset means nothing to me.
If the AI models running on those chips have cyberattack capabilities, or if those chips are training models with cyberattack capabilities and running more instances, it’s not a nuclear weapon, but it does enable a weapon. By the same logic, your argument could also apply to microprocessors and DRAM, or even electricity. In fact, we have extensive export controls on China covering various chip manufacturing technologies, yet we still sell large quantities of DRAM and CPUs to China, and I believe that’s correct.
Host Dwarkesh: I guess this comes back to the fundamental question of whether AI is different. If AI can find zero-day vulnerabilities in software, do we want to minimize China’s capabilities so they never reach that point and cannot deploy it at scale?
Jensen Huang:
We want the U.S. to lead, and we can make that happen. If the chips are already there and they’ve used them to train this model, how do we control the situation? We have massive computing power and a large team of AI researchers—we’re giving it everything we’ve got to move as fast as possible.
Host Dwarkesh: But they have reasons to buy from you—we’ve heard from Chinese company founders that their computing power is a bottleneck.
Jensen Huang:
Because our chips are better. Overall, our chips are better—there’s no doubt about it. Can you acknowledge that Huawei had a record-breaking year without our chips? Can you acknowledge that numerous chip companies went public? Can you acknowledge that we once held a significant share of that market and no longer do? Can you also acknowledge that China accounts for approximately 40% of the global technology industry? To surrender this market for the benefit of a single company is a harm to the United States, a harm to national security, a harm to technological leadership—and it makes no sense to me.
Host Dwarkesh: I’m confused because it seems like you’re saying two different things. First, if we were allowed to compete, we would win against Huawei because our chips would be much better. Second, without us, they would still do the same thing. How can both of these be true at the same time?
Jensen Huang:
This is clearly logical. When there’s no better option, you choose the only one available. What’s illogical about that? The logic is sound. They want NVIDIA chips because they’re better—better in multiple ways: easier to program, a superior ecosystem. Of course, regardless of what “better” means, we’re certainly going to sell them computing power—so what? The fact is, we benefit from it. Don’t forget: we gain from U.S. technological leadership, from contributions by developers working on the U.S. tech stack, and from the fact that AI models spread globally, making the U.S. tech stack the best. We can continue advancing and spreading U.S. technology—it’s positive, and a vital part of U.S. technological leadership. The policies you advocate have effectively driven the U.S. telecommunications industry out of the global market—we don’t even control our own telecom anymore. I don’t think that’s smart; it’s short-sighted and has led to unintended consequences, as I’ve described, which you seem to struggle to grasp. Alright, let’s take a step back.
Host Dwarkesh: The core issue is this: there are potential benefits and potential costs, and we must determine whether the cost is worth paying. Compute power is an input for training powerful models, and such models indeed possess formidable offensive capabilities. It’s beneficial that Anthropic obtained Mythos-level capabilities first—they will hold onto this capability for a while, allowing U.S. companies and the government to patch vulnerabilities and build defenses before release. If China were to acquire more or consolidate greater compute power and thereby train and deploy Mythos-level models earlier and at scale, that would be extremely concerning. The fact that the U.S. has more compute power and companies like NVIDIA is one reason we haven’t seen that scenario unfold. This is the cost of selling chips to China. So, setting aside the benefits for a moment, do you acknowledge this as a potential cost?
Jensen Huang:
I’ll also tell you the potential cost: we allowed the chip layer—one of the most critical layers in the AI technology stack—to lose the entire market, the world’s second-largest market, so that others could build scale and develop their own ecosystem, enabling future AI models to be optimized in ways fundamentally different from the U.S. technology stack. As AI spreads to other parts of the world, their standards and technology stack will surpass ours because their models are open. The U.S. should win every layer of that five-layer cake, including the chip layer. All five layers are vital, and we should win them all.
Host Dwarkesh: But China is the world’s largest contributor to open-source software and open models, yet all of it is built on U.S. technology stacks and NVIDIA.
Jensen Huang:
Correct. Today, it is built on the U.S. technology stack, but if we are forced to leave China, first and foremost, that is a policy mistake—one that has clearly backfired and has already led to the very situations I’ve just described to you: accelerating their chip industry and forcing their entire AI ecosystem to focus inward. It’s not too late yet, but these things have already happened. Over the next few years, you’ll see they won’t stop at 7nm—they excel at manufacturing and will continue advancing beyond 7nm. Is there a tenfold difference between 5nm and 7nm? The answer is no. Architecture matters, networks matter, energy matters. All of these are crucial; it’s not as simple as you’re trying to make it out to be. If you reduce the issue to “whether computing is equivalent to nuclear weapons,” you’ve completely misframed the problem.
AI is not a nuclear weapon; equating AI with various types of weapons only instills fear in everyone, and what benefit does that bring to the United States? If we scare away everyone who wants to become a software engineer by claiming that AI will eliminate all software engineering jobs, we’ll face a shortage of software engineers. So what I’m saying is, when you build such extreme premises—turning everything from zero to infinity—you’re ultimately scaring people in ways that are simply not true. Life doesn’t work that way. Do we want the United States to be first? Absolutely. Do we need to lead at every layer of that technology stack? Absolutely.
We’re talking about Mythos today because it’s important, but in a few years, when we want U.S. technology to spread globally—to India, the Middle East, Africa, Southeast Asia—when our nation seeks to export technology and our standards, I hope you and I will have this same conversation again. I’ll remind you that today’s dialogue, your policies, and your imagination have literally handed over the world’s second-largest market. We shouldn’t give it up. Why would we voluntarily surrender it? No one is advocating for an all-or-nothing approach—no one is calling for unrestricted sales of anything to China at all times. We should always have the best technology in the U.S., the most technology in the U.S., and be first. But we should also strive to compete and win globally—these two goals can happen simultaneously. It requires nuance, maturity, not extremism. The world isn’t black and white.
Why doesn't Nvidia develop multiple different chip architectures?
Host Dwarkesh: We previously discussed the production bottlenecks of TSMC’s N3. If you’ve already secured the majority of N3 capacity, and at some point you’ll become the majority holder of N2—could you potentially leverage the excess capacity of the older N7 node, combined with today’s knowledge of numerical types and other advancements, to redesign a product at the level of Hopper or Ampere? Is this possible before 2030?
Jensen Huang:
There’s no need for that. The reason is that each new architecture involves far more than just transistor scaling—it includes massive engineering innovations in packaging design, chip stacking, numerical formats, and overall system architecture. Going back to use an older node for research and development at this level would be prohibitively expensive—no one could afford it, at least not by going backward. What we can afford is to keep moving forward. Unless one day the world declares that no more capacity will ever be produced—on that day, would I毫不犹豫地 go back to 7nm? Of course, without a second thought.
Host Dwarkesh: I’ve spoken with some people who ask: Why doesn’t NVIDIA simultaneously pursue multiple entirely different chip architectures? For example, Cerebras’s wafer-scale architecture, Dojo’s massive packaging approach, or even a path without CUDA. You have the resources and engineering talent to run these in parallel—so why put all your eggs in one basket? AI could evolve in directions we can’t predict, and architectures could too.
Jensen Huang:
We could do this, but we don’t have a better idea. We could absolutely do all of these things, but none of them are better. We’ve run all of them in simulation, and the results were worse, so we won’t go that route. We’re focusing on the projects we want to do, and if there’s a significant shift in workload—by which I don’t mean algorithmic changes, but real changes in workload patterns, which depend on market structure—we might decide to add other accelerators.
Several years ago, tokens were either free or nearly cost-free. But now, you can offer different pricing for tokens to different customers. Some clients are willing to pay a higher price for them—for example, our software engineers: if I could give them tokens with faster response times to make them more efficient than they are today, I would be willing to pay for that. But this market has only recently emerged. So now I have the ability to segment the same model based on response time. This is why we decided to expand the Pareto frontier and create a segmented market for faster-response inference, even at lower throughput. Previously, higher throughput was always better. We believe there may be a world where extremely high-priced tokens can compensate for lower factory throughput. That’s why we’re doing this. But from an architectural perspective, if I had more money, I would allocate more resources to NVIDIA’s architecture. I think the segmentation of the inference market and the emergence of ultra-high-priced tokens are very interesting directions.
Host Dwarkesh: Final question. If the deep learning revolution hadn’t happened, what would NVIDIA be doing?
Jensen Huang:
Accelerating computation, just as we’ve always done. Our company’s core premise is that Moore’s Law is no longer sufficient for many computing tasks—while general-purpose computing is useful, it’s not ideal for many types of computation. So we’ve combined GPU and CUDA architectures with CPUs to accelerate workloads by offloading different kernels or algorithms from our code onto our GPUs, resulting in applications running 100 to 200 times faster. Where is this applicable? Engineering, science and physics, data processing, computer graphics, image generation, and many other areas. Even if AI didn’t exist today, NVIDIA would still be enormous. The reason is fundamental: the ability to continue scaling general-purpose computing has essentially reached its limit, and the path forward is domain-specific acceleration.
We initially entered the field of computer graphics, but there are many other areas—particle physics, fluid dynamics, structured data processing, and a wide variety of algorithms that benefit from CUDA. Our mission has always been to bring accelerated computing to the world, enabling application types that general-purpose computing cannot handle, and extending them to the level of capability needed to break through bottlenecks in certain scientific fields. Early applications included molecular dynamics, seismic processing for energy exploration, and image processing—all areas where general-purpose computing was too inefficient to be practical.
Without AI, I would be sad, but because of the advances we’ve made in computing, we’ve democratized deep learning, enabling any researcher, any scientist, anywhere, any student to do amazing science with just a PC or a GeForce GPU. This fundamental commitment has never changed—not one bit. If you look at GTC, a large portion of the opening sessions aren’t even about AI—those sessions—computational lithography, quantum chemistry work, data processing tasks—are all unrelated to AI, yet they remain critically important.
