After the launch of DeepSeek V4 Flash, it topped the list of API usage on OpenRouter, as global developers are rapidly migrating to Chinese AI models. Platforms like OpenCode have seen daily usage increase by 30%, while overseas AI startups have reduced their monthly inference costs by 60% to 85% after switching. The total API calls for Chinese models have consistently exceeded the combined usage of U.S. proprietary models for multiple weeks, with developers from the U.S. and Europe accounting for nearly half. This global reshuffling of large models, driven by cost-effectiveness, is breaking Silicon Valley’s long-standing dominance over pricing power.Article author and source: Phoenix Tech
Summary: For a long time, Silicon Valley companies have dominated global pricing power for large models. DeepSeek has created a squeezing effect in the overseas developer market by leveraging its dual advantages of performance and price.
In recent days, an AI-generated comic image has gone viral on social media. Although the figure depicted is Liang Wenheng, it has been stylized in a comic format, with his appearance particularly resembling characters like Superman. Netizens have dubbed this the portrayal of Liang Wenheng as seen by overseas developers.
This comic illustrates that, just last weekend (July 31), the official release of DeepSeek V4 Flash transformed the global landscape of AI usage with its exceptional value for money.
After hands-on evaluations of hundreds of large models, the independent overseas large model evaluation community Artificial Analysis developed a chart, with DeepSeek-V4-Flash emerging as a dividing line on it.

The latest weekly statistics from OpenRouter show that DeepSeek-V4-Flash has topped the platform's model invocation leaderboard, while V4-Pro remains firmly within the top five, as developers worldwide continue to migrate large-scale production inference traffic to this Chinese AI company's model interfaces.
On August 1, OpenCode CEO Jay (Jaya Kumar) posted that, since the launch of DeepSeek V4 Flash, daily usage increased by 30%, and new subscriptions for OpenCode Go also rose by 30%.
For a long time, the global large model market has been shaped by Silicon Valley giants. OpenAI, Anthropic, and Google have set the performance benchmarks and maintained tight control over pricing, creating high barriers to entry for AI startups due to persistently expensive inference costs. Overseas developers seeking stable, high-quality large model services have seemed forced to accept the closed-source pricing offered by these major companies.
The industry landscape has reached an irreversible turning point. Leveraging continuously evolving model capabilities and ultra-competitive cloud API pricing, DeepSeek has established a clear “kill line” in the global developer market.
DeepSeek V4 Flash regains the top spot, with Chinese large models securing the top five positions.
The so-called "kill line" refers not merely to a performance threshold, but to the critical balance point between performance and price: models priced significantly higher than DeepSeek within the same capability tier will continuously lose small and medium-sized developers; products with inferior performance to DeepSeek yet higher costs will see their market space rapidly shrink.
The latest OpenRouter public token usage data shows that DeepSeek-V4-Flash ranks first in single-model API calls on the platform, followed by Xiaomi MiMo-V2.5, Tencent Hy3, DeepSeek V4 Pro, and Zhipu GLM5.2 in second through fifth places. When aggregating the total usage of all Chinese-developed models, they have consistently surpassed the combined usage of U.S. proprietary models for multiple consecutive weeks. Notably, nearly half of the traffic comes from developers in the U.S. and Europe, with numerous agent automation projects, code assistance tools, and long-text processing applications proactively switching their default model to DeepSeek.

In stark contrast, the growth rate of mainstream versions of Claude and Gemini’s flagship models has continued to slow, as the market share of traditional overseas proprietary leading models steadily shrinks. Several overseas AI startup founders have publicly reviewed their experiences on X: after migrating their entire teams to DeepSeek, their monthly inference costs dropped by 60% to 85%. For small AI companies that are not yet profitable and facing tight cash flow, such a significant cost difference can determine whether a project survives.
The market is gradually splitting into two segments: large multinational corporations, due to data compliance and supply chain risk concerns, still prefer to purchase services from local providers such as OpenAI and Google; however, smaller developers, who require greater flexibility and are highly cost-sensitive, have begun a large-scale migration.
Many industry observers believe that OpenRouter’s data has lifted a veil on the industry: in the past, many teams chose models from major Silicon Valley companies not because they believed these models were irreplaceable, but because there had long been a lack of cost-effective alternatives. Now that low-cost models with top-tier capabilities have emerged, demand has immediately shifted.
Of course, data from third-party aggregation platforms has inherent limitations. Traffic volume does not equate to a company’s revenue scale, as most requests are concentrated on low-cost, lightweight Flash versions; additionally, platform samples cannot represent the enterprise and government client market. However, it is undeniable that developers are the source of innovation in the AI industry—by engaging developers, you secure the foundation for the future growth of application ecosystems.
Still the king of value for money
On the same day DeepSeek released V4 Flash, Artificial Analysis tested 128 large models. The term "cut-off line" originated from this, with developers generally agreeing that merely low pricing lacks competitiveness, and merely high performance struggles to disrupt existing markets; only models that achieve near-top-tier performance while offering a significant price advantage can continuously exert pressure on competitors.
Reviewing the evolution of DeepSeek’s pricing strategy reveals a clear progression in building a price barrier. In May 2026, DeepSeek-V4-Pro announced a permanent 75% price reduction, converting a limited-time promotion into a long-term baseline price and undercutting the global floor for high-end inference models. Subsequently, V4-Flash was officially launched, incorporating a conversation caching mechanism and peak-valley differential pricing, further reducing the input token cost in scenarios where cache hits occur.
A horizontal comparison of current mainstream large models' publicly available cloud API pricing (per million tokens, in USD) reveals a clear gap: DeepSeek V4-Flash charges $0.28 per million tokens for output; even after multiple price adjustments, GPT-5.6 Luna charges approximately $1.20 per million tokens for output. Competing flagship models like Claude Sonnet and Gemini are priced several times—up to nearly ten times—higher than DeepSeek.
This data has led the overseas developer community to reach a consensus: for most B2C and small-to-medium B2B commercial scenarios, bearing ten times the inference cost for marginal performance improvements is not commercially viable. Any competitor falling above this “kill line” has only two options: actively lower prices significantly, or completely abandon the mass developer market and focus exclusively on enterprise customers with ample budgets and a strong emphasis on supply chain stability.
Amid ongoing traffic fragmentation, OpenAI urgently lowered pricing for its GPT series models; European vendor Mistral revised its commercialization strategy, retaining its developer base with an open-source free model paired with cloud discounts; numerous smaller open-source model providers have been forced to recalculate costs and abandon the illusion of profiting through high API margins.
More importantly, DeepSeek has long been known in the industry for its low pricing. Rather than relying on subsidy-driven price wars, it continuously reduces per-token inference costs through fundamental engineering optimizations, such as an efficient inference architecture, dynamic KV caching, traffic peak-valley scheduling, and long-text optimization solutions. Inside the company, Liang Wenhong stated that this low-price strategy enables more developers to use DeepSeek—a prospect that has been highly encouraging.
The cost-performance threshold brings a new industry imperative: the core metric for large model competition is shifting. Previously, the industry competed on parameters, training compute power, and benchmark scores; now, developers first evaluate the metric of "how much effective output can be obtained per unit of cost."
The 48 hours that stirred social media
Performance and pricing data ultimately translate into genuine community reputation. X, the Reddit r/LocalLLaMA subreddit, HackerNews, and the LMSYS Arena form the primary battleground for overseas public opinion. On X, numerous frontline developers and AI entrepreneurs continuously post real-world test results, driving sustained momentum for DeepSeek, with public sentiment clearly stratified.
Interestingly, the day before the official release of DeepSeek V4 Flash, OpenAI announced an 80% price reduction for the GPT-5.6 Luna API. The next day, someone posted a screenshot of DeepSeek in the comments, while others thanked Chinese large models, saying they had prompted OpenAI to take notice and benefited developers.
Many individual developers have also expressed strong affection for DeepSeek, not only because of its high cost-performance ratio but also due to its fast response times and smooth user experience.
Wall Street analysts hold slightly divergent views. Optimists believe that DeepSeek’s price competition accelerates the adoption of AI applications and lowers the barrier to global innovation; pessimists worry that the ongoing price war compresses profits across the industry and undermines companies’ long-term funding for model development.
The cost-performance threshold set by DeepSeek essentially represents a direct clash between two global approaches to large model development.
One path, represented by Silicon Valley giants, focuses on heavily investing in training proprietary flagship models, maintaining high profit margins, targeting large enterprise clients, generating revenue through high-value B2B contracts, prioritizing peak capability, and responding relatively slowly to pricing changes in the broader developer market. The other path, exemplified by DeepSeek, relies on engineering optimizations to reduce inference costs, balances open-source ecosystems with cloud services, adopts a low-margin, high-volume strategy, prioritizes capturing the developer community, leverages economies of scale to reduce costs, and rapidly expands its global market share through superior value for money.
Neither path is absolutely right or wrong, but DeepSeek’s rise proves a core truth: the competition among large models is no longer solely about who spends the most money. While computational investment is foundational, inference efficiency, commercial pricing, and ecosystem strategy can equally reshape market share.
The global reshuffling of large models, initiated by cost-performance advantages, is far from over. Price competition is merely the opening act; long-term competition in model capability iteration, global service delivery, compliance solutions, and sustainable monetization models will determine who ultimately stands firm. This past Monday (August 3), MiniMax officially open-sourced its general-purpose video model H3, aiming to replicate DeepSeek’s impact in the multimodal space. On the same day, Alibaba’s Qwen3.8-Max large model was officially launched, boasting a total of 2.4 trillion parameters—the most powerful model in the Qwen family to date. The model weights will be open-sourced next week.
For millions of AI developers worldwide, the most obvious change has already occurred: more choices, lower costs, and a continuously lowering barrier to innovation.
