Kimi K3 Launch Triggers 'Circuit Breaker' as Demand Exceeds Expectations

iconMetaEra
Share
AI summary iconSummary
A surge in demand for Moonshot AI’s Kimi K3 triggered a 'circuit breaker' within 48 hours, forcing the team to suspend new user signups. The 2.8 trillion-parameter model employs a KDA hybrid linear attention mechanism, reducing KV cache by 75% and increasing throughput sixfold at a 1 million context length. Kimi K3 scored 57 out of 60 on global AI benchmarks, ranking third behind Claude Fable 5 and GPT-5.6. The launch caused an $111 billion drop in NVIDIA’s market capitalization and elevated Moonshot’s valuation to $31.5 billion within six months. Altcoins to watch may benefit as the Fear & Greed Index responds to AI-driven market shifts.
Within 48 hours of its release, Moonshot’s Kimi K3 suspended new user subscriptions due to demand far exceeding projections, forcing a computational “circuit breaker.” The K3 model has a total parameter scale of 2.8 trillion and employs the KDA hybrid linear attention mechanism, reducing the computational complexity of attention layers from quadratic to linear, cutting KV cache usage by up to 75%, and increasing decoding throughput sixfold at a context length of 1 million. This architecture is regarded as a practical implementation of the “Tao Law” in AI, achieving performance leaps through systemic architectural innovation. On global AI model benchmarks, K3 scored 57 points overall, ranking third behind Claude Fable 5 and GPT-5.6. In financial markets, NVIDIA’s market capitalization dropped by approximately $111.1 billion on the day of K3’s release, while Moonshot’s valuation surged from $4.3 billion to $31.5 billion within six months.

Article author and source: 36Kr

Kimi also has its own "Taolaw."

K3 is booming, but Kimi must hit pause.

In the late hours of July 19, the Moonshot AI Kimi team announced that, due to user demand far exceeding projections within 48 hours of the Kimi K3 launch, and given limited computing resources, they have temporarily paused new user subscriptions, prioritizing existing subscribers for available computing capacity.

The membership system has been split into Kimi primary benefits and Code benefits to allocate resources more precisely. According to the official response that followed, Kimi did not adjust the membership system but merely rebalanced the user experience under limited computing capacity.

K3 has triggered its computing power "circuit breaker," so we can only politely decline new customers and focus on serving existing users.

Kimi's "Law of Tao"

Currently, major model providers are aggressively expanding their user bases through free services, price cuts, and other cost-effective strategies to attract new users. In contrast, Kimi’s opposite approach seems almost effortlessly elite.

However, the stronger the model, the longer the context, and the more frequent the user requests, the higher the demand for inference computing power. Under the condition that “K3 request volumes are approaching the current computing cluster’s capacity limit,” Kimi is also facing the same issue as “Doubao”: the more C-end users there are, the less noticeable the revenue contribution, but the GPUs can’t keep up.

As the first 3-trillion-parameter MoE model from Moonshot AI, K3 has a total parameter scale of 2.8 trillion. Under conventional understanding, a model of this scale consumes computational power exponentially with each output. However, supported by its unique core architecture, K3 achieves the effect of Huawei’s “Tao’s Law” in the field of chips.

This architecture is also regarded as a super engine, implementing three key technologies: KDA (Kimi Delta Attention) hybrid linear attention, attention residuals, and high-sparsity MoE, maximizing the efficiency of every unit of computing power in enhancing model performance.

For a long time, the semiconductor industry relied on Moore’s Law to improve performance by shrinking transistor sizes. However, as advanced processes approach physical limits and costs surge, the traditional model of simply stacking hardware has hit a bottleneck.

This year, Huawei introduced the "Taolaw Principle," moving beyond the traditional approach of spatial miniaturization toward temporal miniaturization: by leveraging logical folding, 3D stacking, and end-to-end architectural optimization, it reduces signal transmission latency and minimizes redundant data movement, achieving energy efficiency gains comparable to advanced process nodes on mature process technology.

Its core methodology is: achieving performance breakthroughs through system-level architectural innovation under hardware limitations or cost constraints.

In my view, the KDA hybrid linear attention mechanism adopted by K3 is a practical application of the "Tao Law" in the field of AI.

Keep in mind that the Transformer architecture is most computationally intensive in its attention mechanism. Standard attention scales quadratically with sequence length, and in million-token context tasks, the majority of computational resources are consumed by maintaining the attention matrix for long sequences. The KDA mechanism reduces the computational complexity of most attention layers from quadratic to linear.

According to the technical report previously released by Moonshot AI, using a mixed ratio of KDA and MLA on experimental models reduced KV cache occupancy by up to 75% and increased decoding throughput by six times with a 1 million context length. Combined with attention residual and high-sparsity MoE technologies, training efficiency improved by approximately 25%, with additional costs under 2%.

This is a reimagining of "model production efficiency." If the mainstream Silicon Valley Transformer approach is like a gasoline-powered car that relies on brute force to achieve results, KDA is more like a hybrid engine optimized for peak thermal efficiency.

SemiAnalysis's report indicates that KDA improves inference speed by 300% while maintaining performance close to Transformer models in long-context processing. This means that, with the same computational investment, K3 can generate more effective intelligence.

Real-world feedback from the developer community shows that K3 excels in long-form agent tasks and code generation, demonstrating competitive strength against leading international models.

At its core, K3 still adheres to the Scaling Law, but targets inefficient algorithmic waste. While Silicon Valley AI companies continue to hoard chips and pile on computational power, Kimi achieves a balance between performance and efficiency by reengineering its algorithmic pipeline and eliminating computational redundancy—paving a path that maximizes intelligent output per watt.

This surprised Wall Street observers and investors, who, like their skepticism toward DeepSeek, once again questioned whether the hundreds of billions of dollars invested in AI by Silicon Valley would yield corresponding returns, and whether America’s lead in AI has been diminished.

Meanwhile, enterprise users and developers are also beginning to pay attention to AI usage costs. Some developers have already adopted "model routing" services, which automatically select the most cost-effective and efficient AI model for each task—including Chinese models such as Kimi.

Yang Zhilin reaches for the height

Looking back, Kimi’s computing power being overwhelmed was an unintended side effect of technical success: improvements in model efficiency reduced the cost per use, which in turn triggered a surge in overall demand.

Notably, on July 11, Zhipu’s founder Tang Jie announced the "Touch High" initiative, stating outright, "If we don’t reach the summit, it’s a failure," and clearly emphasized that over the next two years, the company will not pursue short-term monetization of applications but will instead concentrate resources on advancing four core areas of AGI, with plans to raise funds and allocate billions of resources to break through mechanical interpretability technology.

If Zhipu's "reaching for the heights" is a make-or-break move after its listing, it aims to solidify its narrative as the world's first publicly traded large model company and alleviate genuine concerns by pushing the boundaries of technology.

I also see Kimi's "reaching higher" in three dimensions within Yang Zhilin's decisions:

First, prioritize computing power allocation by shifting from a traffic-based mindset to a value-based mindset. Public information shows that Kimi’s API business now accounts for over 70%, a significant departure from its previous reliance on consumer subscription models. By retaining existing users and rejecting new ones, Kimi is directing limited computing resources toward high-value use cases, preventing mass consumer traffic from crowding out B2B business segments that are now core revenue drivers.

The cost is also evident. For example, the criteria for implementing "prioritizing existing users" are not transparent; if certain requests are prioritized while others are deprioritized, and mismanagement occurs, user trust will be eroded.

In the long term, Kimi must strengthen its infrastructure—including elastic compute pools, tiered response mechanisms, and dynamic scheduling capabilities—to transition from a technology standout to a mature leading AI enterprise.

Second, pricing is being pushed to its limits, as Kimi undergoes stress testing from a low-cost buffet model to a tiered-value structure. The announcement’s mention of “separating Kimi’s core rights from Code rights” serves as a short-term adjustment to address compute shortages and may also signal the beginning of a broader pricing system overhaul.

In the past, users paid the same fee and consumed the same amount of computing power whether they were writing code, researching, or chatting. As the proportion of users consuming high computing power for coding increases, this model inevitably leads to structural misalignment.

Kimi is taking this opportunity to break down membership benefits and reprice them according to usage scenarios: light users enjoy basic services, while heavy users pay a premium for high-value features.

This is an attempt to gain pricing power for model products through user segmentation by value. Currently, K3’s pricing has reached $2.3 per million tokens, which, while lower than top overseas models, has set a new high for domestic models.

Perhaps Kimi also wants to move away from low-cost models and convey a message to the market: high-performance models have pricing power. Moving forward, competition among large models will shift from “competing on parameters” to “competing on infrastructure” and “competing on cost pricing.”

This step is harder than reaching the top of the leaderboard and closer to the essence of business.

Third, Kimi's elevation from a geopolitical variable to a pricing anchor.

K3 quickly garnered global attention after its release. In the latest "Intelligence Index" published by the independent AI model evaluation agency Artificial Analysis, it achieved a composite score of 57, ranking third globally, behind Claude Fable 5 (60) and GPT-5.6 Sol (59), and ahead of Claude Opus 4.8 (56).

This achievement also inadvertently shattered the industry consensus that open-source projects lag behind proprietary ones by half a generation.

Meanwhile, K3's influence has extended beyond the tech community. A U.S. tech investment firm has identified it as a "turning point" in AI development, asserting that high-quality open-source models are accelerating the diffusion of capabilities. A computer science professor at the University of California, Berkeley, claims that K3 has narrowed the gap between Chinese open-source models and advanced U.S. models to just two to three months.

The capital market reacted just as sharply: on K3's first trading day, NVIDIA's market value dropped by approximately $111.1 billion in a single day; JPMorgan Chase strategists referred to it directly as "DeepSeek 2.0" in their report.

This acceleration is what changes the pricing logic.

In addition, Moonshot AI's annualized recurring revenue (ARR) had approached $300 million by June, with overseas growth reaching 400% and revenue increasing despite a 60% price hike. These three metrics together indicate that Kimi’s “open-source + efficiency” model can be scaled profitably.

Adding to the news, Moonshot is reportedly seeking to list in Hong Kong within as little as six months, with its valuation surging from $4.3 billion to $31.5 billion in six months—an increase of more than sevenfold.

These are capital markets affirming a non-Silicon-Valley mainstream path. By betting on SOTA models with a viable unit economic model, they demonstrate that Kimi has developed its own valuation logic, no longer needing to benchmark itself against OpenAI or Anthropic to prove its worth.

Dark Side of the Moon, image source: online resources

But K3 still has a long way to go before reaching AGI. For example, in benchmark scores, although K3 is only 3 points lower than Fable 5, the underlying gap remains significant.

Additionally, Kimi acknowledged two deployment-level flaws in its technical report: first, K3 retains a complete thinking history during training; if the caller fails to properly pass back the previous reasoning process or switches to K3 from another model within a conversation, the output quality will significantly degrade.

Second, K3 tends to be "overly proactive" when tasks are ambiguous, preferring to make decisions on behalf of the user. This tendency to act independently can trigger cascading errors in engineering processes that require precise execution.

Another easily overlooked but critically important issue: hallucinations. Artificial Analysis specifically noted that K3’s hallucination rate has increased compared to the previous generation, K2.6.

To some extent, K3’s shine and underlying concerns reflect the industry-wide imbalance between computing power supply and demand, as well as the collective challenges faced by China’s AI industry under the heavy pressures of computing power bottlenecks, uncompleted commercialization, and overly high user expectations.

Each overload and each "circuit breaker" event by these AI companies is a hurdle China's AI sector must overcome in its pursuit of higher goals. One thing is certain: a new competitive cycle has begun, where large models will shift from competing on capability to competing on computing power.

Those who can build advantages in algorithms, computing power reserves, and other infrastructure will ultimately prevail and define the survival standards for the next generation of AI companies.

Reference materials:

Letter Ranking, "K3 Released in 48 Hours"

Brocade, It's All Kimi's Fault

This article is from the WeChat public account "Tang Chen's Notes," authored by Tang Chen, and is published by 36Kr with permission.

Disclaimer: The information on this page may have been obtained from third parties and does not necessarily reflect the views or opinions of KuCoin. This content is provided for general informational purposes only, without any representation or warranty of any kind, nor shall it be construed as financial or investment advice. KuCoin shall not be liable for any errors or omissions, or for any outcomes resulting from the use of this information. Investments in digital assets can be risky. Please carefully evaluate the risks of a product and your risk tolerance based on your own financial circumstances. For more information, please refer to our Terms of Use and Risk Disclosure.