DeepSeek invests $2.56 billion in Huawei Ascend 950DT for data center, faces challenges in transitioning from CUDA

iconMetaEra
Share
AI summary iconSummary
DeepSeek has ordered $2.56 billion worth of Huawei Ascend 950DT accelerators for a gigawatt-scale data center in Wulanchabu, Inner Mongolia, with deliveries set to begin in Q4 2026. The company is adapting its training software to support Huawei’s hardware, replacing NCCL with HCCL and rewriting key operators such as FlashMLA and DeepGEMM. On-chain data indicates continued reliance on NVIDIA for training due to Huawei’s production delays, particularly in advanced memory components. Inflation trends may impact future hardware procurement strategies.
DeepSeek has invested $2.56 billion to purchase at least 160,000 Huawei Ascend 950DT accelerators for deployment in a gigawatt-scale data center in Ulanqab, Inner Mongolia, with deliveries scheduled to begin in Q4 2026. Liang Wenheng regards training models using domestic chips as one of the company’s “biggest bets,” and is currently training a large model with over 2 trillion parameters, exceeding the scale of the current flagship V4. DeepSeek’s shift to Ascend faces three major challenges: the entire training software stack must be rewritten, communication primitives must be replaced from NCCL to HCCL, and chip production is constrained by U.S. export controls. Currently, Huawei is facing production bottlenecks due to shortages of core components such as advanced memory, meaning DeepSeek cannot fully escape NVIDIA in the short term. Meanwhile, the open-sourcing of TileLang provides an operator development toolkit for domestic chips, reducing adaptation costs.

Author and source: Leiphone

DeepSeek releases an open-source operator toolkit in collaboration with Huawei Ascend, eliminating CUDA dependencies!

A single codebase can run on all chips.

At the investor closed-door meeting, Liang Wenfeng described "training models using Huawei or other domestic chips" as one of the company's biggest bets right now, and explicitly stated that this effort "must succeed."

DeepSeek has placed an order worth $2.56 billion to purchase at least 160,000 Huawei Ascend 950DT accelerators for deployment in its gigawatt-scale data center under construction in Ulanqab, Inner Mongolia. Deliveries for this order are expected to begin as early as Q4 2026. These Ascend 950DT accelerators are primarily intended for inference workloads, while training tasks continue to rely mainly on NVIDIA.

In terms of computing power deployment, DeepSeek has not treated domestic chips as a fallback option to NVIDIA, but rather positioned them at the core of its strategic priorities. Liang Wenhong believes that domestic chips will inevitably reach global prominence—it is only a matter of time—and that DeepSeek should act first to adapt to semiconductor hardware beyond NVIDIA.

DeepSeek has shifted to Ascend, but three major challenges remain.

First section: The entire training software stack has been rewritten.

The training code for DeepSeek-V3 and V4 was originally written for NVIDIA's Tensor Core/Hopper architecture WGMMA instructions. To run on Ascend, every cuBLAS/NCCL call must be replaced with Ascend C/HCCL. Additionally, the proprietary core operators FlashMLA, DeepEP, and DeepGEMM must be individually verified for performance on the Ascend 950.

Second part: Underlying replacement of communication primitives.

Training large models requires synchronizing gradients across thousands of accelerators, with collective communication primitives serving as the foundation of parallel training. NCCL in NVIDIA's ecosystem and HCCL in Ascend's architecture differ fundamentally in API design, performance characteristics, and fault tolerance models. Migrating from NCCL to HCCL is not as simple as changing an import statement—any mismatch in adapting collective communication operations can cause the entire training process to fail silently.

Third: Bet on chip production capacity.

Hidden risks lie in the supply chain. The driving force behind DeepSeek’s shift to domestic chips is U.S. export controls, which also limit Huawei’s ability to scale up production of the Ascend 950. DeepSeek’s bet on 160,000 Ascend chips implicitly assumes that Huawei’s Q4 deliveries cannot falter.

Currently, DeepSeek is training a large model with 2 trillion parameters, exceeding the 1.6 trillion parameters of the current flagship V4 model; Liang Wenheng also revealed that the company has planned an even larger model with 8 trillion parameters.

The larger the model scale, the exponentially greater the computing power required for training and inference. According to Liang Wenhong’s previous estimate, training a large model on par with OpenAI’s would require approximately 50,000 NVIDIA GB300 processors, or equivalently, 200,000 Huawei Ascend 950 chips.

Although DeepSeek is fully committed to transitioning to domestic computing power, the reality is that Huawei still faces supply gaps in its current chip production. Sources close to the matter reveal that Huawei is encountering production bottlenecks due to shortages of critical components such as advanced memory, meaning DeepSeek will remain dependent on NVIDIA hardware in the short term. To address this, DeepSeek has completed multiple rounds of large-scale funding and continues to invest heavily in expanding its computing capacity.

Over the next 12 months, at least four variables will determine the outcome of this bet: Huawei’s Ascend 950 Q4 delivery rate, which dictates DeepSeek’s training schedule; whether the TileLang Ascend backend can match the CUDA backend within six months, determining development efficiency; whether ByteDance, Tencent, and Alibaba follow suit in betting on Ascend, shaping the ecosystem’s scale; and whether U.S. export controls escalate again, setting the floor for this path.

There is still no definitive answer to these questions. But one trend is clear: the open-sourcing of TileLang has transformed the question of "how to build a domestic computing power ecosystem" from an engineering challenge into an ecosystem issue.

Compared to CUDA and TileLang, they represent two entirely different ecosystem paths: CUDA is a proprietary, full-stack, hardware-bound ecosystem dominated by vendors, with deep accumulation but extremely high migration costs; TileLang is an open-source, focused, cross-hardware community tool designed specifically to address the core operator development challenges of large models, offering high flexibility and low porting costs.

The significance of TileLang lies not in replacing CUDA, but in addressing the most urgent gap today—enabling domestic chips to rapidly gain top-tier operator capabilities for large model training without having to build a complete software ecosystem from scratch. As more large model vendors develop operators based on TileLang, the adaptation cost for domestic chips will decrease exponentially, and the entire ecosystem’s iteration speed will accelerate significantly.

In a sense, DeepSeek has demonstrated that world-class supercomputing clusters can be built using domestically produced chips.

As the competition among large models enters the deep waters of computational cost and supply chain security, "defining compute through software" is no longer just a slogan—底层 tool innovations like TileLang are quietly reshaping the landscape of China's computational ecosystem.

DeepSeek releases an open-source operator toolkit in collaboration with Huawei Ascend, eliminating CUDA dependencies!

Disclaimer: The information on this page may have been obtained from third parties and does not necessarily reflect the views or opinions of KuCoin. This content is provided for general informational purposes only, without any representation or warranty of any kind, nor shall it be construed as financial or investment advice. KuCoin shall not be liable for any errors or omissions, or for any outcomes resulting from the use of this information. Investments in digital assets can be risky. Please carefully evaluate the risks of a product and your risk tolerance based on your own financial circumstances. For more information, please refer to our Terms of Use and Risk Disclosure.