DeepSeek has released its latest model, V4-Flash, with a cost of only about $0.03 per task—the lowest among mainstream large models globally. V4-Flash is priced at $0.14 per million input tokens and $0.28 per million output tokens, resulting in a significantly lower overall cost compared to competitors—approximately just 1/105th of Anthropic’s Claude 3 Haiku. DeepSeek is preparing for its IPO, with its API cloud service generating annual recurring revenue of $400 million to $500 million and a gross margin exceeding 50%. Competition in the AI industry is shifting from model performance to inference cost and commercial efficiency, with the marginal cost of enterprises deploying AI applications expected to drop by an order of magnitude.Article author and source: Wall Street Journal
DeepSeek has released its latest model, V4-Flash, once again bringing the cost competition in AI models to the forefront.
According to the latest analysis by research firm Artificial Analysis, the average cost per task completed by V4-Flash is only about 3 cents—the lowest among today’s leading global large models, roughly 1/105 the cost of Anthropic’s Claude Fable 5. This suggests that, for many general-purpose tasks, the marginal cost for businesses deploying AI applications could decline further, reigniting market discussions about the return on massive AI capital expenditures.
Notably, the release of V4-Flash comes as DeepSeek accelerates its commercialization efforts. According to reports citing informed sources, DeepSeek is preparing for a potential IPO. Its annual recurring revenue (ARR) from AI model API cloud services has reached $400 million to $500 million, with a gross margin exceeding 50%.
The newly launched V4-Flash continues the company’s consistent focus on low cost and high efficiency, further strengthening its strategy of driving market competition through price and efficiency.
V4-Flash costs just 3 cents per task, leading the world in cost-performance among mainstream models.
According to Artificial Analysis, V4-Flash is priced at $0.14 per million input tokens and $0.28 per million output tokens.
However, the institution noted that "Cost per Task" is a more accurate indicator of a model's actual usage cost compared to mere list pricing. This metric takes into account the total number of input and output tokens consumed when completing the same task, making it more meaningful for evaluation.
At this rate, the average cost per task for V4-Flash is approximately $0.03, the lowest among leading global models.
For comparison, Moonshot AI's Kimi K3 is approximately $0.86, OpenAI's GPT-5.6 Sol is about $1.86, and Anthropic's Claude Fable 5 is around $3.15. The cost of V4-Flash is only about 1/105 of the latter.
Artificial Analysis further emphasizes that models with lower per-token pricing may still incur higher overall costs if they require generating more content or undergoing more interaction steps to complete a task. Therefore, "cost per task completed" is a more accurate indicator of a model’s value for money than a single price metric.
Breaking through with low prices, AI competition shifts from "competing on performance" to "competing on value for money."
Compared to performance rankings, V4-Flash has garnered more market attention for its cost advantage.
Over the past year, the focus of competition in the AI industry has gradually shifted from model capabilities to inference costs and commercial efficiency. As enterprise AI applications continue to scale, model operational costs have become a critical factor in procurement decisions.
When a model can complete a task for approximately 3 cents, while some high-end models cost several dollars, businesses deploying it in high-frequency scenarios such as customer service, office automation, and code assistance may see a drastic reduction in overall usage costs.
This again touches on Wall Street’s ongoing concern: whether massive capital investments in AI will ultimately generate sufficient business returns.
In early 2025, after DeepSeek released the R1 model, the market briefly reassessed the necessity of global AI infrastructure investment, triggering significant volatility in tech stocks. Today, V4-Flash once again highlights its extremely low operational cost as its core selling point, signaling that competition in the AI industry is shifting further from "which model is stronger" to "who can deliver sufficiently strong capabilities at a lower cost," potentially reigniting discussions around AI return on investment.
