DeepSeek Announces a 500% Increase in API Pricing and Launches the Open-Source Agent Framework Harness

iconTechFlow
Share
AI summary iconSummary
DeepSeek increased its API prices by up to 500% effective August 17, 2026, with peak rates reaching 12 times current levels. The company also open-sourced its Agent framework, DeepSeek Harness, on GitHub. Market responses were mixed, with some observers noting the move aligns with growing open interest in AI-driven tools. Fear and Greed Index readings indicate cautious optimism among developers, who see potential in the transition from token sales to Agent productivity. Despite the increase, DeepSeek’s rates remain below those of global competitors, according to Morgan Stanley.

Author: Claude, Deep潮 TechFlow

DeepChao Summary: Just one day after releasing the official V4 Pro version, DeepSeek announced a price increase plan: starting August 17, it will implement peak-off-peak pricing, with maximum increases of up to 500%, and peak-hour rates reaching up to 12 times the current price. On the same day, the developer preview version of DeepSeek Harness, the agent execution framework, was open-sourced and launched. With one hand raising prices and the other offering tools, DeepSeek’s strategy is clear: it’s no longer just selling tokens—it’s now selling agent productivity.

On the evening of August 13, DeepSeek announced a price adjustment for its API: the new pricing will take effect at 00:00 Beijing time on August 17, implementing peak and off-peak rates. Peak hours are daily from 9:00 to 12:00 and 14:00 to 18:00, with off-peak rates set at half the peak rate. This announcement came less than 24 hours after the official launch of V4 Pro, and the "affordable window" we highlighted in yesterday’s report closed sooner than expected.

Price Increase Details: Cache-hit input price increased by 500%; peak price is up to 12 times the current price.

According to Beijing Business Daily and The Paper, the new pricing for the flagship model V4 Pro is: 0.15 yuan per million tokens for cached input during off-peak hours, 4.5 yuan for uncached input, and 13.5 yuan for output; during peak hours, these prices double to 0.3 yuan, 9 yuan, and 27 yuan, respectively. The V4 Flash has also been adjusted accordingly, with off-peak prices at 0.05 yuan, 1.5 yuan, and 4.5 yuan for the three categories, and peak-hour prices at 0.1 yuan, 3 yuan, and 9 yuan.

There are two metrics for the price increase, both worth noting. Compared to the off-peak rate, the maximum increase is 500% (the cache hit input price rose from ¥0.025 to ¥0.15 for V4 Pro); compared to peak-hour rates, the cache hit input price is now 12 times the current rate, and the output price is 4.5 times the current rate. Regular users can still use V4 Pro for free on the web and app—the groups truly affected are developers and enterprise customers accessing the model via API.

Developers are outraged: “This Pro version offers no value for money”—but it’s still a fraction compared to overseas prices.

Developers on Weibo responded directly: “It’s no longer cheap—prices have caught up to those of similar models.” “The price hike during peak times was too steep; this Pro version offers no value for money.” Others turned their hopes elsewhere, calling for the rapid rollout of Ascend’s new chips to give DeepSeek an opportunity to lower prices again. DeepSeek previously promised a significant price reduction for the Pro model after the bulk launch of Ascend 950 super nodes in the second half of the year—now, that promise carries even greater weight.

However, when the ruler is moved overseas, the financial picture isn’t nearly as bad. Even at the peak-hour output price of 27 yuan—approximately $3.80—it’s still only a fraction of Claude Fable 5’s $50 output price. The change lies in the pricing logic: the price gap between cache hits and misses has widened to 30 times, turning off-peak scheduling and caching from “optimizations” into “cost-critical factors.” Moving batch tasks to overnight idle periods cuts costs in half—that’s the clear signal left for developers in the announcement.

On the same day, Harness launched: DeepSeek no longer wants to just sell tokens.

Almost simultaneously with the price increase announcement, the developer preview of DeepSeek Harness opened for testing, with the full source code publicly available on GitHub and garnering over 560 upvotes on Hacker News. The official page clearly defines its positioning: “An Agent equals a Model plus a Harness”—the model is the soul, while the Harness is the body that enables the Agent to understand context, invoke tools, and operate continuously in real-world environments.

Harness is the “office system” for Agents, responsible for assigning tasks, coordinating tools, logging activities, and managing workflows—with a design where everything is a plug-in: models, tools, sandboxes, storage, scheduling, and interfaces can all be swapped and recombined, while every step leaves a traceable record. The Science and Technology Daily’s analysis reveals the business strategy: DeepSeek is no longer just selling tokens, but selling Agent productivity—raising model prices while offering a free, open framework to make Agents truly usable, making ecosystem lock-in more compelling than pricing alone.

Industry Signal: Morgan Stanley says China's large models are moving beyond price competition into a battle of intelligence.

This price increase is not an isolated event. According to a Morgan Stanley report released on August 9, titled "Goodbye to Price Wars, Welcome to the War of Intelligence," the average API input price for large models in China rose to RMB 4.9 per million tokens in the second quarter of 2026, while the output price increased to RMB 21.9 per million tokens—up approximately 48% and 80% respectively from the first quarter of 2025. Companies such as ByteDance, Alibaba, Baidu, Tencent, and Moonshot AI are among those affected.

In other words, DeepSeek was merely the last player to put down its "price weapon," and even after the price increase, it is still likely to remain one of the cheapest in its class. For AI application entrepreneurs and compute chain investors, the three key metrics to track are: actual developer migration after August 17 takes effect (whether they stagger their move or switch to competitors); whether the delivery timeline for the Ascend 950 can deliver on the promise of "further price cuts"; and whether Harness’s plugin ecosystem can replicate the rapid adoption speed of DeepSeek’s models. Yesterday’s "pennies" were the cost of acquiring customers; today’s price is the business.

Disclaimer: The information on this page may have been obtained from third parties and does not necessarily reflect the views or opinions of KuCoin. This content is provided for general informational purposes only, without any representation or warranty of any kind, nor shall it be construed as financial or investment advice. KuCoin shall not be liable for any errors or omissions, or for any outcomes resulting from the use of this information. Investments in digital assets can be risky. Please carefully evaluate the risks of a product and your risk tolerance based on your own financial circumstances. For more information, please refer to our Terms of Use and Risk Disclosure.