DeepSeek Releases V4 Pro and Raises API Prices by up to 1,100%, Focusing on Agents

Chinese startup has officially launched the V4-Pro-0813 with a context of 1 million tokens and is repricing the API effective August 16, marking the end of the aggressive pricing phase that brought it into the spotlight in 2025.
DeepSeek released the production version of the V4-Pro-0813, its flagship model, this Thursday, and in the same announcement stated that the API pricing will increase by up to 1,100% starting at 16:00 UTC on August 16. This is the first time the Chinese startup has openly acknowledged that the aggressive pricing model that put it on the map in 2025 was unsustainable.
The V4-Pro-0813 features a context window of 1 million tokens and allows outputs of up to 384,000 tokens, with the option to operate in thinking mode or directly. The declared focus is agents: using tools, executing code, and orchestrating multi-step tasks without human supervision. This positions the product against Anthropic's Claude Opus 4.7 and the Fable 5 family, models that have dominated the business agent vertical in recent months.
Price Changes the Conversation
According to Caixin Global, the entry token cache-miss price for the V4-Pro rises from $0.435 per million to $1.32 per million during peak times, a jump of 200%. In some pricing tiers, the variation reaches 1,100%. DeepSeek has segmented the day into windows: peak hours between 01:00-04:00 and 06:00-10:00 UTC; outside of peak hours, customers pay half the peak price. This follows the same dynamic pricing logic that cloud providers use to flatten GPU peaks.
For the European CIO who ran a POC last summer with the promise of costs ten times lower than American competitors, the calculations have changed. The gap between the V4-Pro during peak hours and a Claude Fable 5 at equivalent volume is no longer a chasm; it is now a marginal discount. For the Chinese hyperscaler that resold DeepSeek tokens, the margin compression is immediate.
Preview Outperformed by Flash
An embarrassing detail, which the company did not hide in the press material, is that the V4-Flash launched in July has outperformed the preview version of the V4-Pro in several independent benchmarks since April. DeepSeek spent three months tuning the V4-Pro on top of a Flash that was already delivering results. The 0813 addresses the gap, but the market perception has changed: the capital expenditure return curve for training has flattened even in-house.
None of this is abstract for banks and consultancies. JPMorgan released this week a projection of $1.2 trillion in AI capital expenditure by the end of 2027, and most of that figure is based on the thesis that the cost per token would continue to fall. DeepSeek has just introduced a public disagreement with this curve.
Market Readings
In the United States, the repricing legitimizes the argument by OpenAI and Anthropic that the floor price for inference is not zero. Both companies have been trying for months to convince investors that the gross margin in production is defensible, and they promoted a new cost metric this week, according to Bloomberg. A Chinese competitor validating the price floor helps their narrative.
In India, the epicenter of offshoring development and where DeepSeek has become the default in several POC teams due to a combination of price and latency, the pass-through will hit the P&L of software factories. TCS, Infosys, Wipro, and several new AI-native firms structured business proposals in the last two rounds based on July's pricing. Multi-year contracts signed on this basis will have pressure on margin until the next renegotiation.
For the Big 4 and MBB, the pass-through is even more direct. Accenture, which announced in May the cutting of 11,000 jobs within an $865 million restructuring attributed to AI, has assembled business agent offerings in recent months with price structures anchored on the V4-Pro's launch cost. Deloitte and Bain have followed the same path. If the pass-through remains at 200% even outside of peak times, the contractual margin of these offerings will decrease before the next fiscal turnaround, and none of them have the ability to renegotiate client by client in a short window.
In Brazil, the effect is indirect: companies that conducted proof of concept tests with V4-Pro due to cost now need to redo their TCO spreadsheets to decide whether to migrate workloads to a domestic provider, remain with DeepSeek outside of peak hours, or close volume contracts with Anthropic or OpenAI. Itaú and Bradesco, which opened RFPs for LLM usage in customer service and credit, are likely to request new business conditions in the coming weeks.
DeepSeek has not announced a corresponding price increase for the V4-Flash at the same level. The company seems to be positioning Pro as a premium product and reserving Flash for volume retention. If this strategy works, it is a signal that the Chinese market is following the same tier segmentation path that the American market has trodden over the last two years.