Chinese AI models face high costs despite high usage

iconMetaEra
Share
AI summary iconSummary
Chinese AI models remain popular despite rising operational costs, with token usage leading for nine consecutive weeks. However, high expenses and unreliable service undermine their cost-effectiveness compared to GPT Codex. Domestic models such as GLM-5.2 and Kimi face quota restrictions and throttling during peak hours, while Zhipu AI reported losses of 3 billion yuan by 2025. Inflation data underscores the pressure on firms, as China lags in computing infrastructure, with only 449 data centers compared to 5,427 in the U.S. The AI developer fear and greed index reflects growing concern over long-term sustainability.
Although China’s domestic large AI models lead the U.S. in usage volume by nine weeks, their actual costs in high-intensity programming scenarios far exceed those of GPT Codex plans. Zhipu recorded losses exceeding RMB 3 billion in 2025, while Kimi and other plans face usage limits and throttling during peak hours. The primary reason is China’s insufficient computing infrastructure (449 data centers versus 5,427 in the U.S.), preventing the launch of high-cost-performance Codex-style plans. Currently, developers must balance performance, price, and stability; Chinese models only approach international top-tier models in certain performance metrics and continue to face the challenge of being well-received but poorly adopted.

Article author and source: Zuihua FunTalk

Over the past month, Silicon Valley’s AI community has unveiled a series of major releases, including GPT-5.6, Fable5, and Grok-4.5.

For a time, domestic developers and programmers were abuzz, with widespread discussions and evaluations of cutting-edge models like GPT-5.6.

Although new domestic models like Kimi-K3 have emerged this month, overall, domestic large models continue to be met with praise but limited adoption among top-tier developers.

According to the latest data, in terms of usage: From June 22 to 28, Chinese models reached 2.039 trillion tokens in weekly usage, while U.S. models reached 425 billion tokens, with Chinese models leading for nine consecutive weeks.

In contrast, the domestic AI "Six Startups" have performed poorly in terms of revenue: public financial reports show that Zhipu reported revenue of RMB 724 million in 2025, but incurred losses exceeding RMB 3 billion; MiniMax reported revenue of $79 million (approximately RMB 570 million); the remaining four companies, including Moonshot, have not yet disclosed their full annual revenues.

Even leading companies still have revenues at the hundreds of millions level.

In comparison, Anthropic, a leading AI company, has achieved an ARR of approximately $60–70 billion, primarily driven by B2B API and AI coding use cases.

This raises a sharp question:

Why, despite being touted as "cost-effective" and "high-volume," have domestic models still failed to gain widespread adoption and remain unable to directly compete with top international models in high-value scenarios such as AI coding?

Recently, after in-depth discussions with multiple frontline developers, I discovered a counterintuitive fact:

Domestic large models are, to some extent, quite expensive.

It is precisely this high cost that prevents certain domestically developed models with performance breakthroughs from truly showcasing their strengths.

01 The Illusion of "Low Price"

When people think of domestic large models, they often think of representatives like DeepSeek and Kimi, one of the most closely associated labels being low cost and affordability;

Looking solely at token pricing, domestic models indeed offer significant advantages compared to top-tier models like GPT-5.6 and Fable 5. For example, GLM-5.2 charges $1.4/$4.4 per million tokens for input/output, while DeepSeek-V4-Pro is priced at $0.435/$0.87; in contrast, GPT-5.6 Sol costs $5/$30, and Claude Fable 5 is priced at $10/$50.

In reality, this impression of “low price” is a long-standing misconception among outsiders.

In particular, under high-intensity programming scenarios, the unit price of domestic large models not only lacks an advantage but is actually several times higher than GPT and Claude with bundled plans.

Li Guang (a pseudonym) is a researcher in the field of AI who, due to the demands of his work, consumes billions of tokens daily in AI coding scenarios. Under such usage volumes, relying on domestic large models like GLM-5.2 would result in prohibitively high overall costs.

Li Guang's token transaction statement

On the day of the conversation, Li Guang consumed nearly 4 billion tokens. If he used GLM-5.2, the approximate cost would be: (86.39 million regular inputs × $1.4 + 3.38 billion cached inputs × $0.26 + 19.06 million outputs and inferences × $4.4) = about $1,084. At an exchange rate of 7.2 RMB per USD, this equals approximately RMB 7,800.

The reason Li Guang feels so free to burn tokens is that he subscribed to the GPT CodeX plan.

Here’s a simple explanation of the CodeX plan: In short, it’s a package designed specifically for programming scenarios. Strictly speaking, it’s not truly unlimited—instead, it integrates Codex into the ChatGPT subscription: Plus costs $20/month, while Pro starts at $100/month, offering approximately 5 to 20 times the quota of Plus. Once you reach your limit, you can purchase additional Credits to continue using the service.

Similarly, Anthropic’s Claude Code offers a comparable set of plans: Claude Pro at $20/month, Max 5x at $100/month, and Max 20x at $200/month; the Claude web interface and Claude Code share a quota of 5 hours every 5 hours and weekly limits, with additional usage available at API rates after reaching the cap.

Under this package structure, the substantial Token costs are effectively offset.

For developers in China, consuming billions or even tens of billions of tokens daily under intensive AI coding scenarios is a common norm.

Liu Xiao is also a heavy user of AI coding; he actually spends 300 RMB per week on Codex, even without company reimbursement.

But even at this cost, it still appears significantly cheaper compared to domestic models, as the same 300 yuan could not buy an equivalent number of tokens for GLM 5.2.

Xiao Liu's Token Statement

From the perspective of domestic developers, GPT’s Codex plan is currently the cheapest token available—several times more affordable than Zhipu or Kimi. This may seem counterintuitive, but the fact remains that domestic models cannot match its cost-effectiveness in high-intensity coding scenarios.

It should be noted that Codex is not truly an unlimited plan; at the highest tier of $200, the weekly limit is approximately $2 billion.

However, since some developers use models through intermediaries or other methods, the actual cost of certain models can be significantly lower than through official channels, resulting in situations where you can buy more tokens at a lower price—hence the phenomenon of being able to purchase 12 billion tokens for just $300.

In fact, domestic models have previously offered similar packages, but upon closer examination, these packages impose significant restrictions on limits compared to Codex.

02 The Pain of Domestic Pricing

The current state of domestic large models in pricing can be summarized in one sentence:

You thought you were walking into a buffet, but in the end, the vendor still charged you by weight.

For example, regarding subscription plans, domestic services like Kimi and GLM-5.2 offer four pricing tiers: 49, 99, 199, and 699 RMB, with quotas refreshed weekly. For GLM Coding Plan, the Lite, Pro, and Max tiers provide approximately 80, 400, and 1,600 prompts every 5 hours, or roughly 400, 2,000, and 8,000 prompts per week, respectively. During peak hours, GLM-5.2 deducts usage at three times the normal rate.

Taking Alibaba’s Qoder as an example, according to a friend of the author, named Ami (pseudonym), who works in B2B customer operations: “Qoder was once the most user-friendly AI IDE in China, but its new pricing strategy eliminated its advantages.”

Regarding plans: Qoder CN Personal Pro and Pro+ are priced at RMB 59 and RMB 169 per month, including 2,000 and 6,000 credits respectively; Teams and VPC versions are priced at RMB 99 and RMB 199 per user per month, both including 3,000 credits, with the VPC version available for a minimum of 50 users.

"After they merged Lingma and adjusted their pricing this year, at least twenty to thirty companies, as far as I know, have stopped renewing their subscriptions," A Mi later explained.

According to Ami's testing, the Qoder Enterprise Plan at $199 is exhausted in just three days with his usage—and would likely be depleted in a single day with heavy developer usage.

Similarly, regarding packages, GLM-5.2 has left domestic users with mixed feelings.

According to developer Xiao Wei: "The GLM-5.2 plan is essentially nonexistent; even if a plan exists, traffic is still throttled during peak hours, and there are instances of users being betrayed."

According to 36Kr: The GLM-5.2 plan was harder to secure than concert tickets, refreshing punctually at 10 a.m. daily—whether you got one depended entirely on your speed.

In January 2026, Zhipu announced that due to limited computing power, it had reduced its daily available quantity to just 20% of the original amount, releasing a limited allocation each day at 10:00 AM, which sold out within minutes.

Ultimately, the reluctance of domestic models to offer generous packages and low prices stems from a lack of computing power.

According to 2025 data from China Academy of Information and Communications Technology and the Ministry of Industry and Information Technology, China has 449 existing data centers, while the United States has 5,427. In terms of total computing power, China stands at 1,053 EFLOPS, compared to the United States’ 2,400 EFLOPS.

But the key point here is that in the U.S.'s 2,400 EFLOPS, high-end GPUs make up a very large share; in China's computing power, a significant portion consists of domestic chips such as Huawei Ascend and Cambricon, which still lag behind the H100 in single-card performance.

To address rising computing power demands, starting in March 2026, Zhipu, Alibaba Cloud, Tencent Cloud, and Baidu collectively increased their token prices by between 5% and 460%.

In short, the era of "low prices" and "price wars" for domestic models has long passed.

It’s not that domestic models don’t want to offer buffet-style pricing—it’s because the numbers don’t add up: at today’s compute costs and current loss levels, offering a monthly subscription means betting a fixed fee against users exhausting the service. Given how domestic developers and programmers routinely consume billions, even tens of billions, of tokens in a single day, this business model would quickly breach its cost threshold.

OpenAI and Anthropic’s programming packages target high-net-worth enterprise clients—customers like Cursor and GitHub Copilot generate $1.4 billion in annual revenue for Anthropic.

With the support of high computational power, this business model essentially exchanges a fixed monthly fee for long-term customer retention, resulting in high-quality customers and relatively high ARPU, which helps to offset costs effectively.

In comparison, although China has cloud giants like Alibaba, the issue is that Alibaba Cloud's adjusted EBITA margin for FY2026 is only 7.5%; although the Cloud Intelligence Group is experiencing rapid revenue growth, its profits are not substantial.

Moreover, traditional cloud providers can offer "unlimited cloud servers" because users don't operate at full load 24/7. A single ECS instance may have an actual CPU utilization of only 10%, allowing the cloud provider to oversell one physical server to ten users.

But large model inference is different: each token generated requires actual GPU computation power. If users continuously generate tokens 24/7, the cloud provider’s HBM memory, GPU cores, and electricity costs become rigid expenses with no possibility of overbooking.

03 The "Golden Triangle" of the Model

In addition to price considerations, the compelling programming performance of Codex and Claude Code is undoubtedly why domestic developers choose them.

However, in terms of performance, domestic models are not as some stereotypes suggest.

After speaking with multiple developers, a common observation emerged: domestic large models, particularly the latest ones like GLM-5.2, are now capable of handling auxiliary and secondary tasks such as testing and fixing bugs.

However, for more complex and long-running asynchronous tasks, domestic models still lag significantly behind top-tier models like GPT and Claude.

Developer Xiao Zhang used a domestic model for programming in March, comparing GLM, Qwen, and GPT on the same task—only GPT completed the task successfully and met the page requirements. The project involved multi-level, multi-column content comparison and result generation.

Another developer, Xiao Wei, said that based on her experience, Kimi 2.7 and Doubao 2.1 Pro perform reasonably well, but they are generally one tier lower (belonging to the second tier), yet still outperform Claude Sonnet 4.6 overall.

Currently, the general consensus among developers is that GLM-5.2 is the first domestic model to reach Opus-level performance.

It is available and well-established for executing real, complex, long-running asynchronous tasks.

Although it still lags significantly behind the true Opus model in many aspects, such as world knowledge, this does not prevent it from representing a new milestone for domestic models.

But in reality: this has just begun to emerge, with only a few "top performers," yet they are held back by other factors unrelated to performance.

Specifically, factors such as low cache efficiency and continued rate limiting during peak times result in a poor developer experience.

This introduces another dimension in current model comparisons: stability.

Among frontline developers and programmers, model selection is no longer a simple comparison of performance—but rather—

The golden triangle of performance, price, and stability.

In these three dimensions, domestic models only barely achieved some metrics in the first dimension.

Recently, China's Kimi released its latest flagship model, K3: an open-weight flagship model with 2.8 trillion parameters, native multimodality, and a 1 million-token context window, primarily designed for long-range programming, knowledge work, and deep reasoning.

Among numerous evaluations, it scored 57 on Artificial Analysis’s intelligence index, ranking third globally; it topped the Arena front-end programming leaderboard with a score of 1679, and there are even rumors that its performance rivals that of Claude Fable 5.

However, behind the impressive scores and rankings, developers remain most concerned about one word: value for money.

After the release of K3, many developers expressed that Kimi is still expensive, offering poor value compared to Codex, with insufficient quotas. Even looking solely at API pricing, K3 cannot be considered "budget-friendly": the cost per million tokens for cache, input, and output is 2, 20, and 100 RMB respectively; GPT-5.6 Sol is $0.5, $5, and $30, meaning K3 is cheaper than Sol but significantly more expensive than GPT-5.6 Luna’s $0.1, $1, and $6.

During the use of K3, developers reported experiencing API disconnections, including one instance that lasted an entire morning.

Currently, domestic developers are not unwilling to pay for domestic models; rather, due to limited computing resources and high costs, they cannot afford cost-effective plans like Codex, and can only opt for so-called "Dianping plans"—apparently all-you-can-eat but actually limited in usage (similar to Zhipu’s Coding Plan).

The inconsistent experience and unclear value proposition of the packages make it difficult to retain users after initial use. This is, at its core, a rational choice by users given the high cost of the current computing power infrastructure.

None of the above is meant to boost others' confidence or diminish our own; we simply wish to highlight a risk.

Currently, domestic models, while catching up with GPT and Claude, tend to heavily emphasize performance. However, the journey from development to application for domestic large models involves climbing far more than just the mountain of performance.

What we need is patience—waiting until domestic computing power infrastructure reaches large-scale mass production, when technology, pricing, and stability come together in harmony, allowing our product to meet users’ expectations.

Disclaimer: The information on this page may have been obtained from third parties and does not necessarily reflect the views or opinions of KuCoin. This content is provided for general informational purposes only, without any representation or warranty of any kind, nor shall it be construed as financial or investment advice. KuCoin shall not be liable for any errors or omissions, or for any outcomes resulting from the use of this information. Investments in digital assets can be risky. Please carefully evaluate the risks of a product and your risk tolerance based on your own financial circumstances. For more information, please refer to our Terms of Use and Risk Disclosure.