GPT-6 Astra Utilizes Over 100,000 Blackwell GPUs, Highlighting China's Advanced Cluster Gap

iconMetaEra
Share
AI summary iconSummary
AI + crypto news: GPT-6 Astra was trained using over 100,000 Blackwell GPUs, signaling a new phase in AI development. Chinese companies such as ByteDance and DeepSeek are scaling GPU usage but still lag in Blackwell deployment and cluster efficiency. The main challenge for China is not the number of GPUs, but building highly efficient systems. The latest crypto news shows that AI and infrastructure remain key battlegrounds.
GPT-6 Astra publicly disclosed that it was trained using over 100,000 Grace Blackwell GPUs, signaling that the competition for cutting-edge models has entered the era of hyperscale clusters.

Article author and source: ME News



TL;DR:

  • GPT-6 Astra publicly disclosed that it was trained using over 100,000 Grace Blackwell GPUs, signaling that the competition for cutting-edge models has entered the era of hyperscale clusters.
  • China's leading model companies do not lack large-scale computing power, but public information indicates there remains a clear gap in advanced GPU generations and the scale of centralized deployment.
  • Companies such as ByteDance, Kimi, and DeepSeek are expanding their computing power infrastructure, but it is currently more focused on Hopper architecture or domestic alternatives, still lagging behind Blackwell-class clusters.
  • The key to AI competition has shifted from “whether you have GPUs” to “whether you can obtain the latest-generation GPUs and efficiently organize tens of thousands or even hundreds of thousands of chips.”
  • Even though Blackwell is no longer NVIDIA's latest architecture, the trend of Rubin further reducing training costs suggests that the lead gap may continue to concentrate in infrastructure capabilities.

Behind 1 million Blackwell GPUs, the competition in large models has entered the era of computing infrastructure.

When Huang Renxun revealed the training scale of GPT-6 Astra, the public was not only focused on the number “100,000 GPUs,” but on what this figure represents—a new paradigm of AI competition.

Over the past few years, competition among large models has often been viewed as a race in algorithms, data, and engineering capabilities. Innovations in model architecture, optimizations in training methods, and improvements in inference efficiency have indeed determined the ultimate capabilities of AI products. However, entering 2026, a growing trend is becoming clear: as model scales continue to expand, infrastructure capabilities themselves are emerging as a critical factor in defining the technological upper limit.

According to Jensen Huang's public disclosure, the training of GPT-6 Astra utilized over 100,000 Grace Blackwell GPUs connected via NVLink72 to form a high-speed interconnect cluster. He also mentioned that the next batch will bring online 400,000 GPUs, though he did not disclose the specific models or their ownership.

Regardless of subsequent deployment details, this signal is clear: training the most advanced models is shifting from competition involving hundreds of trillions of parameters and thousands of GPUs to competition among GPU clusters numbering in the tens or even hundreds of thousands.

The key here isn't just the quantity.

If we only compare the number of GPUs, Chinese companies are not entirely without large-scale computing resources. The real gap lies in whether they can access the latest generation of high-performance GPUs and whether they can enable these GPUs to work together efficiently within the same training task.

For large models, the peak performance of a single GPU is just the baseline. Training a large model requires continuous exchange of parameters and data among a large number of GPUs. If interconnect efficiency is insufficient, even having many chips may fail to deliver effective computational power.

Therefore, competition in AI infrastructure has evolved from "how many chips to buy" to "how to build the right computing system."

The computing power deployment of China's leading model companies lags behind the Blackwell era by a generational gap.

Current public information shows that the computing power of China's leading model companies is rapidly increasing, but there remains a significant gap compared to the Blackwell-class training clusters represented by GPT-6 Astra.

ByteDance is one of the companies in China with the most advanced public compute infrastructure, approaching the level of global leaders. This year, ByteDance connected approximately 36,000 B200 GPUs through Malaysia. These chips are based on NVIDIA’s Blackwell architecture and belong to the same generation as the GB200 chips used by Astra.

However, it should be noted that this batch of B200s is deployed overseas. When ByteDance publicly describes its AI infrastructure within China, it primarily relies on the China-specific H20 based on the Hopper architecture, as well as previously acquired H800 chips.

What this reveals is not merely a difference in quantity, but a gap in supply chain and deployment environment.

Compared to Hopper, Blackwell has been upgraded in terms of computational power, memory capacity, and interconnect capabilities. In particular, the GB200 is not merely a single GPU, but rather integrates the Grace CPU with the Blackwell GPU, connecting 72 GPUs within a single high-speed interconnect domain via the NVL72 system.

For training large models, the importance of this system-level design is growing increasingly significant. As model sizes increase, the proportion of communication between GPUs rises, and the efficiency of collaboration among chips directly impacts training performance.

Behind Kimi, Moonshot AI has reportedly acquired approximately 20,000 Hopper GPUs through Alibaba. Bloomberg cited sources claiming the specific model is the H200, but Alibaba has denied that the H200 model was involved.

If calculated based on the H200, it still belongs to the previous-generation Hopper architecture, while the B200 has entered the Blackwell era. The two are not simply a matter of new replacing old, but represent different stages of AI infrastructure capability.

The situation with DeepSeek is even more special.

In early 2025, DeepSeek garnered global attention with its highly efficient models, demonstrating that algorithmic optimization and engineering efficiency can partially reduce compute requirements. However, based on publicly available information, the company has not disclosed the full hardware configuration used for training V4.

The leaked transcript of Liang Wenfeng’s investor Q&A session showed that the company had approximately 20,000 H-equivalent compute units in May, most of which had just arrived, and future purchases will be “almost entirely NVIDIA.” Huawei provided DeepSeek with a capacity of about 16,000 units of the 950. Liang Wenfeng stated that this batch of domestic chips is roughly equivalent to 4,000 NVIDIA B-series GPUs—just sufficient for training the current generation of models.

This information reflects a reality: even though Chinese companies already possess a substantial amount of AI computing power, there remains a scale gap compared to the requirement of deploying 100,000 advanced, same-generation GPUs for a single training task.

China's real weakness in AI computing power is not a lack of chips, but the absence of advanced cluster capabilities.

When discussing China's AI computing power, it's easy to fall into two extremes.

One view holds that China completely lacks advanced AI chips and therefore cannot compete; another view argues that as long as the number of domestic chips increases, China can quickly catch up to NVIDIA.

In fact, the situation is more complex.

AI training capability is determined by multiple factors, including chip performance, advanced manufacturing processes, memory capacity, high-speed interconnects, server design, data center power supply, software ecosystem, and large-scale scheduling capability.

The GPU is just the most critical component, but not the whole picture.

NVIDIA’s long-term advantages stem not only from its GPU hardware but also from its CUDA ecosystem, networking technologies, server solutions, and the comprehensive system built with cloud service providers.

During the Blackwell era, the advantages of this system are further amplified.

When training a model requires 100,000 GPUs, the issue is no longer simply procuring tens of thousands of chips, but rather how to build a large-scale computing facility capable of stable operation.

This is also why Chinese companies have not yet disclosed any cases of training a model using 100,000 identical advanced GPUs in public records.

Chinese companies are enhancing their capabilities through various approaches, including increasing overseas compute deployment, expanding the adoption of domestic chips, optimizing model architectures, and improving training efficiency. While all these paths hold value, they are unlikely to fully replace the foundational performance advantages provided by the most advanced GPU clusters in the short term.

Over the past few years, China's AI industry has demonstrated that algorithmic innovation can significantly narrow certain gaps.

The emergence of companies like DeepSeek demonstrates that outstanding engineering teams can achieve breakthroughs under limited computational resources by optimizing model architectures, adjusting training strategies, and improving resource utilization efficiency.

But another fact remains: as global competition enters the next generation of foundational models, the importance of computational scale will continue to rise.

Efficiency optimizations can reduce costs, but they struggle to fundamentally alter the performance boundaries set by advanced hardware and hyperscale clusters.

After Blackwell, the Rubin era could further widen the infrastructure gap.

More importantly, Blackwell is not the end of NVIDIA’s current technology roadmap.

NVIDIA's next-generation Vera Rubin architecture has entered mass production. According to NVIDIA's public estimates, training large MoE models requires approximately one-fourth the number of GPUs compared to Blackwell.

This means a new cycle of AI competition may emerge in the future.

Leading companies have the latest architecture, enabling them to perform larger-scale training with fewer chips; lower training costs encourage more experiments, further enhancing model capabilities.

Companies that can only access legacy hardware may face two issues: reduced training efficiency and difficulty competing with the largest models.

Of course, this does not mean that Chinese AI companies have no opportunities.

AI history has repeatedly shown that technological competition is not solely determined by hardware alone. Model innovation, application scenarios, commercialization capabilities, and engineering efficiency can all lead to new breakthroughs.

However, from an infrastructure perspective, the fact that GPT-6 Astra uses 100,000 Blackwell GPUs reveals a significant shift: the core battleground of global AI competition is moving further from the model level to the compute infrastructure level.

In the coming years, an AI company’s competitiveness will depend not just on its ability to train a high-quality model, but on its capacity to continuously train the next generation of models.

For China's AI industry, what truly needs to be caught up is not a single model or a single release, but the entire system capability spanning chip acquisition, data center construction, cluster orchestration, and software ecosystem development.

The significance of 100,000 Blackwell chips lies not in creating a stunning number, but in reminding the market that artificial intelligence is entering an industrial phase, and industrial competition ultimately comes down to infrastructure.

Source:

  1. NVIDIA's official public materials on Blackwell, GB200 NVL72, Vera Rubin architecture, and AI infrastructure.
  2. Huang Renxun's public speeches and media interviews.
  3. Bloomberg report on Kimi's AI compute procurement and H200-related matters.
  4. ByteDance's publicly disclosed information on AI infrastructure and overseas computing power deployment.
  5. DeepSeek public information and transcript materials from the investor meeting with Liang Wenheng.
  6. Alibaba Cloud's public information regarding AI infrastructure and large model computing power cooperation.
Disclaimer: The information on this page may have been obtained from third parties and does not necessarily reflect the views or opinions of KuCoin. This content is provided for general informational purposes only, without any representation or warranty of any kind, nor shall it be construed as financial or investment advice. KuCoin shall not be liable for any errors or omissions, or for any outcomes resulting from the use of this information. Investments in digital assets can be risky. Please carefully evaluate the risks of a product and your risk tolerance based on your own financial circumstances. For more information, please refer to our Terms of Use and Risk Disclosure.