DeepSeek Surpasses Its Own Pro Model in Agentive Benchmarks with V4-Flash-0731 at $0.14 per Million Tokens

Launched on July 31, V4-Flash-0731 scores 82.7 on TerminalBench 2.1, compared to 72.1 for V4-Pro-Preview, maintaining the same architecture as the April preview and costing $0.14 per million input tokens.
DeepSeek released DeepSeek-V4-Flash-0731 on Hugging Face and moved the official API to public beta on July 31, 2026. The company was straightforward in its model documentation: the architecture and checkpoint size are identical to the April preview. The model retains a total of 284 billion parameters, with 13 billion activated per token via sparse Mixture-of-Experts, a context window of 1 million tokens, and a maximum output of 384 thousand tokens. What changed was the post-training, redone with a focus on coding, reasoning, tool usage, and agentive flows.
For product teams already using the deepseek-v4-flash endpoint, the transition was automatic: applications inherited the upgrade without code modification. The model now natively supports the Responses API format and entered beta in the Codex API, the same interface standard adopted by OpenAI, making direct comparisons between the two tool environments easier without stack migration.
The Numbers from DeepSeek are Substantial
According to data published by DeepSeek itself, V4-Flash-0731 outperforms V4-Pro-Preview across all nine of the company's agentive benchmarks. In TerminalBench 2.1, the new model registers 82.7, compared to 72.1 for the Pro-Preview and 61.8 for the April Flash Preview. In DeepSWE, the difference is even more pronounced: 54.4 on the 0731 versus 7.3 in the preview, a variation of 7.4 times. The intelligence index from Artificial Analysis places the model at 50 points, 10 above the previous Flash.
A caution is necessary: all these numbers come from DeepSeek's internal harness. No independent laboratory had reproduced the metrics by the time of this article's publication. This does not invalidate the data but keeps it within the realm of vendor statement while awaiting external verification.
Price is the Argument that Doesn't Depend on Benchmark
V4-Flash-0731 is listed on OpenRouter at $0.14 per million input tokens for cache misses and $0.28 per million output tokens. Cache hits drop to $0.0028. Anthropic charges $2.00 per million input tokens on Sonnet 5, an introductory price valid until August 31, 2026, and $5.00 on Opus 5, according to price tables published by the company in July 2026. The difference in the case of Anthropic's top-tier model is 36 times.
In TerminalBench 2.1, Anthropic's Opus 4.8 scores 85.0, according to the comparative table published by DeepSeek itself, 2.3 points above V4-Flash-0731. Sonnet 5, which Anthropic positions as performing close to Opus 4.8, costs 14 times more than V4-Flash-0731 at the introductory price. If DeepSeek's benchmarks hold up under external evaluation, the distance between the Chinese model and the top of the Western market is smaller than the gap between their prices.
From China to Silicon Valley and Bengaluru
DeepSeek is operated by High Flyer Capital Management, a quantitative firm based in Shenzhen. The model distributes open weights via Hugging Face, where the hosted inference has reported a throughput of 400 tokens per second at $10 per hour, and is available globally via API, including through OpenRouter and DeepSeek's own endpoint.
In the United States, the launch coincides with the peak adoption of coding agents in development environments. Software companies and automation startups assessing inference costs at scale now have a concrete price argument: $0.14 per million input tokens, regardless of the model's geographic origin.
In India, Wipro, Infosys, and HCL Technologies manage software automation contracts valued in the tens of billions of dollars. For solution architects at these firms, a model with agentive benchmarks of this magnitude, for less than a tenth of Sonnet 5's introductory price, requalifies the calculus between adopting the model infrastructure preferred by Western clients and building their own stack with lower-cost alternatives.
What the market awaits in the coming weeks is straightforward: independent reproduction. If external labs confirm the agentive benchmarks of the 0731, the model presses Western providers to justify margins that currently rely, in part, on the absence of verified alternatives in this price range. If the numbers shrink under third-party evaluation, the episode hardens the growing skepticism about self-published benchmarks in an industry where the incentive to inflate metrics is structural and the capacity for external verification remains slow.