Tuesday evening this week, DeepSeek The Open Platform has announced plans to adjust the pricing of its large model API once again—this time, a welcome price reduction.

Off-peak pricing is uniformly: input (cache hit) 0.02 yuan, input (cache miss) 1 yuan, output 4 yuan (all units in yuan per million tokens). Peak-hour pricing remains twice the off-peak rate: cache hit 0.04 yuan, cache miss 2 yuan, output 8 yuan.

This adjustment applies only to the V4-Flash series; the price of the higher-tier V4-Pro remains unchanged. The new prices will take effect on September 10 at 12:00 Beijing Time.
The official also specifies that peak hours are Beijing Time Monday through Friday, 9:00–12:00 and 14:00–18:00; all other times are off-peak.
Previously, DeepSeek The price increase has sparked widespread discussion. The official explanation is to allocate resources more reasonably and encourage users to adjust their task execution times. Outsiders generally believe it is DeepSeek V4 Flash 0731 offers excellent value, attracting a large volume of calls from users worldwide, and DeepSeek Insufficient computing power.
Compared vertically, before the price increase at 00:00 on August 17, Flash had a uniform pricing of 0.02 yuan (cache hit) / 1 yuan (cache miss) / 2 yuan (output). Here is the new pricing included:

In other words, the two input prices have returned to their pre-increase levels, while the output remains twice the pre-increase level. When comparing peak period rates, the difference is even greater: the new peak prices of 0.04 / 2 / 8 yuan versus the pre-increase prices of 0.02 / 1 / 2 yuan represent increases of 2x, 2x, and 4x, respectively.
Therefore, whether users can truly return to a state of “unrestricted usage” depends on three variables specific to their application: cache hit rate, input-to-output ratio, and the degree to which tasks can be staggered. Scenarios with high cache hit rates—such as long contexts, fixed system prompts, or Agent cyclic calls—will show significant reductions; whereas scenarios primarily focused on generating output will experience limited impact. As is well known, DeepSeek The API cache hit rate is relatively high, so the feedback received so far has been quite positive.
This price reduction is not an isolated event; yesterday afternoon, DeepSeek The official community has announced the start of internal testing for the V4.1 Flash intermediate release: the model name is deepseek-v4.1-flash-expires-on-0910, the base_url remains unchanged, with a limit of 20 concurrent requests per account, and billing remains the same as for V4 Flash.
In addition to native multimodality and extremely fast inference speed, the official also stated that cost reduction has been achieved through architectural improvements. This demonstrates that cost efficiency can be enhanced through engineering and technical optimizations of the large model itself. DeepSeek We have been continuously working to improve the issue of inference costs.
The model name "expires-on-0910" indicates that the new model will automatically expire and be taken offline on September 10—the same day the price reduction takes effect.
This month, leading AI companies updated their large models. DeepSeek If the official V4.1 release delivers a significant leap in capability, combined with this price reduction, it could spark another major surge.
This article is from the WeChat public account "Machine Heart" (ID: almosthuman2014), authored by Machine Heart focusing on large models, edited by the Machine Heart editorial team.
