DeepSeek Launches V4.1 Flash with 1M Context Tokens

DeepSeek announced V4.1 Flash on September 10, with 552B MoE, 1M context tokens, and automatic routing to Flash from Pro starting September 14.
A New Flash, A New Price Ceiling
DeepSeek announced on Tuesday, September 9, the V4.1 Flash model, available from 12:00 PM Beijing time on September 10 (04:00 UTC). The new pricing took effect at the same time. At off-peak rates, V4.1 Flash charges $0.15 per million non-cached input tokens, $0.003 per million cached input tokens, and $0.60 per million output tokens. During peak hours, from 1 AM to 4 AM and between 6 AM and 10 AM UTC on weekdays, these rates double. Starting September 14, any call to deepseek-v4-pro will be automatically routed to Flash, charged at the lower rate.
What's Under the Hood
Asymmetric Causal Encoder-Decoder architecture with a total of 552 billion parameters, only 8 billion active for input processing, and 16 billion for output generation. Context window of up to 1 million tokens. Native multimodal understanding of text and image. It is the smallest model in the new family of DeepSeek architecture, designed for cheaper inference and later scaling. According to public evaluation from AlphaSignal, V4.1 Flash surpasses the company's previous flagship at one-fourth the memory cost.
In third-party compiled benchmarks, Flash either ties or remains within shooting distance of GPT-5.6 Sol and Claude Opus 5 in coding tests and agent usage. In intensive knowledge assessments like HLE, and in some terminal and coding tests, the two western models still lead by a clear margin. DeepSeek has not yet published an official technical report or benchmark table, so some readings must await the next wave of independent replications.
What Changes for AI Buyers in the Next Two Weeks
The price floor for a frontier model with 1 million context tokens and multimodal capability has dropped to $0.60 per million output tokens. Anthropic charges $10 per million input and $50 per million output in Fable 5.1. OpenAI charges comparably in Astra. The difference is by an order of magnitude.
The immediate cut affects three conversations. The first is about mass agent work. Mid-sized consultancies that built their orchestration layer on top of Sonnet or Fable will redo the spreadsheet in three scenarios next week. An operation routing 20% of the volume to DeepSeek for lower-risk tasks (summarization, extraction, text transformation) changes the monthly infrastructure run rate significantly. Much enterprise work in LLM does not require the top-tier offerings from the western provider, and V4.1 Flash is made for that segment.
The second conversation pertains to sovereignty and procurement. European banks operating under the AI Act and private data sales under GDPR need contractual assurance from the provider, which is currently harder to obtain from DeepSeek. Enterprise buyers in the U.S. are evaluating jurisdictional restrictions on models of Chinese origin in pipelines connected to sensitive data, a discussion that has gained traction following the Biden administration's package on advanced computing and chip export restrictions. Brazilian multinationals operating in the three markets will need to run more than one model in production, with routing policies by data type and jurisdiction.
The third conversation is about the roadmap. The automatic routing of deepseek-v4-pro to Flash beginning September 14 is the format that Anthropic and OpenAI have avoided so far. Transparent downgrading by cost margin, with the client's contract being fulfilled at the new rate. If successful, it alters market expectations regarding what enterprise buyers can anticipate in continuity when a provider launches a cheaper generation of its own line.
What the Board Needs to Hear
The decline in token prices will not stretch indefinitely. Training costs continue to rise, and each new round of infrastructure compresses the inference cost at the end, not at the beginning of the cycle. Several analysts view part of the training capex as transitional, but Anthropic and Google continue to operate with margin, which separates the reading from what applies to a laboratory burning venture capital. What Flash 4.1 demonstrates is that the price ceiling for frontier work in multimodal tasks with 1 million tokens dropped in a week. Those looking to sign multi-year contracts in October should request an automatic review clause for reference price drops. It is the same clause commodity buyers sought in the 1990s. It makes sense again for AI.