DeepSeek V4 Pro Launches with Agent Performance Approaching Claude 3 Opus at a Fraction of the Cost

iconTechFlow
Share
AI summary iconSummary
On-chain news: DeepSeek V4 Pro was launched on August 13, delivering Agent performance nearing that of Claude Opus 5 at a lower cost. The model supports up to 1 million tokens in context and 384,000 tokens in output, with pricing significantly cheaper than overseas competitors. The update garnered attention on Hacker News and Zhihu. DeepSeek also announced upcoming increases to its API pricing. Token launch news continues to underscore the rapid advancements in competitive AI development.

Author: Claude, Shenchao TechFlow

DeepChao Overview: At凌晨 on August 13, DeepSeek unexpectedly launched the official version of V4 Pro: the same API name, but the agent capability surged from the preview’s “barely functional” to nearly matching Anthropic’s flagship Claude 3 Opus, surpassing it on some benchmarks. Even more striking is the price—just a fraction of what comparable models charge overseas. For developers and AI entrepreneurs, the entry cost for flagship agents has been shattered again. The only catch? The official price increase notice has already been posted.

This early morning, the DeepSeek official website quietly updated: the DeepSeek-V4-Pro model has been upgraded to the official 0813 version. Developers do not need to modify any code—calling deepseek-v4-pro automatically uses the latest model. The new model supports up to 1 million tokens in context and a maximum output of 384,000 tokens, with thinking mode enabled by default. It is fully compatible with OpenAI, Anthropic, and Responses API formats, allowing seamless integration with leading agent tools such as Codex, Claude Code, and OpenCode. The announcement has simultaneously topped the homepage of Hacker News (over 700 upvotes) and Zhihu’s trending list (over 20 million views), becoming today’s biggest consensus topic in both Chinese and English AI communities.

Midnight raid: The same API name, which couldn't do the job yesterday, can do it today.

The core of this update is the Agent's capabilities. According to official evaluations, the V4 Pro final version's score on the company's own software engineering benchmark, DeepSWE, surged from 12.8 in the preview version to 62.7; NL2Repo for natural language generation of code repositories increased from 38.5 to 61.5; and DSBench-Hard, the challenging subset for full-stack development, doubled to 67.2.

According to a real-world test by a domestic tech media outlet: “The same API name—yesterday it couldn’t even handle the task, but today it can go head-to-head with the world’s top-tier code agents.” In AnFanEr’s testing, when given vague instructions like “collect recent tech news and create a homepage mimicking traditional media,” V4 Pro autonomously broke down the requirements and executed them step by step, achieving a very high level of completion. Developers on Hacker News have also shared real-world cost reports: running a full day of traffic simulation and distributed physics engine optimization consumed about 2 billion tokens for roughly $12.50, resulting in “significant performance improvements without introducing any new issues.”

It surpasses Fable 5 in security defense, but it’s not yet the top contender when it comes to open-source origins.

According to official benchmarks, V4 Pro scored 83.3 on the AI security defense benchmark Cybergym, surpassing Claude Fable 5’s 83.1; it also outperformed on the workflow agent benchmark AutomationBench; on the challenging reasoning benchmark HLE (with tools enabled), it achieved 60.0, just below Fable 5’s 63.0; and on Terminal Bench 2.1, it scored 87.9, higher than Claude Opus 4.8’s 85.0.

But two points need to temper the enthusiasm. First, it still lags slightly behind Moonshot’s Kimi K3 (87.9 vs. 88.3 on Terminal Bench 2.1), meaning the top spot in the open-source community hasn’t changed hands yet. Second, third-party metrics are less optimistic than the official claims: OpenRouter, citing Artificial Analysis’s Composite Intelligence Index, rated it at 45.3, placing it “better than 70% of models”—a noticeable gap from the official narrative of “matching Fable 5.” The official benchmarks were measured using their own Harness framework, and independent verification will take time.

3 yuan for $10, the flagship agent’s “penny pricing”

What truly puts competitors at a disadvantage is pricing. The official release of V4 Pro maintains the preview version’s pricing: 0.025 yuan per million cached token inputs, 3 yuan per uncached input, and 6 yuan per output. For comparison, Kimi K3’s API pricing is $3 per input and $15 per output; Claude Haiku 3 is $10 per input and $50 per output. Even accounting for currency differences, DeepSeek’s pricing is only a fraction of that of overseas equivalent models.

OpenRouter's data shows that, due to a cache hit rate as high as 86%, the weighted input price users actually pay is as low as $0.064 per million tokens. As the saying goes in China’s crypto circle: DeepSeek’s dominance remains “cheaper than those better than me, and better than those cheaper than me.”

Price increase notice has been posted; API output described as "clunky" draws criticism

Cheap pricing may be temporary. On August 6, DeepSeek announced it would “soon implement a significant overall price increase for its API” and plans to introduce peak-off-peak pricing: during peak hours daily from 9:00 to 12:00 and 14:00 to 18:00 Beijing Time, the cost for unmatched inputs for V4 Pro will rise from ¥3 to ¥6, and outputs from ¥6 to ¥12—doubling the current rates. Since the official release, no price adjustment has taken place, effectively granting developers a widely understood grace period. Another counterbalancing factor is DeepSeek’s earlier promise during the preview launch that prices for Pro will be significantly reduced after the bulk rollout of Ascend 950 super nodes in the second half of the year. The gap between these potential increases and decreases depends on the timing of domestic compute capacity deliveries.

Additionally, developers on X have reported a noticeable difference in chain-of-thought performance between the V4 Pro web interface and API calls, with API outputs sounding "stiff and awkwardly phrased." DeepSeek has not yet responded. Teams relying on the API for their products should test it thoroughly before launch.

After the collapse of the flagship agent costs, the next focus is on Harness.

Put this into the bigger picture:

This week, research on "stealing inference traces from closed-source model APIs" sparked renewed debate, revealing cracks in the proprietary models' competitive moat. Now, DeepSeek has demonstrated that the performance gap is narrowing, offering top-tier results at a fraction of the cost. The pricing power of closed-source flagship models has weakened another notch.

For AI application entrepreneurs and investors focused on compute chains, the most promising hint lies in the official documentation. DeepSeek’s V4 Flash changelog reveals that its Agent execution framework, DeepSeek Harness, is set to launch, with recruitment and official WeChat public account channels already in place. The industry consensus is that as model inference capabilities approach their limits, the upper bound of Agent performance will be determined by frameworks like Harness—the “operating systems” for AI agents. With the models already in place, if the framework delivers, the next DeepSeek moment won’t be far off.

Disclaimer: The information on this page may have been obtained from third parties and does not necessarily reflect the views or opinions of KuCoin. This content is provided for general informational purposes only, without any representation or warranty of any kind, nor shall it be construed as financial or investment advice. KuCoin shall not be liable for any errors or omissions, or for any outcomes resulting from the use of this information. Investments in digital assets can be risky. Please carefully evaluate the risks of a product and your risk tolerance based on your own financial circumstances. For more information, please refer to our Terms of Use and Risk Disclosure.