Google Cuts 17% Off Price of Gemini 3.6 Flash and Starts Pre-training of Gemini 4

Google launched Gemini 3.6 Flash at $1.50 per million input tokens and $7.50 for output, generating 17% fewer tokens. On the same day, DeepMind confirmed the start of the largest pre-training in history.
Google released Gemini 3.6 Flash on Tuesday with a direct price cut: $7.50 per million output tokens compared to $9.00 for version 3.5, and $1.50 per million for input. The context cache is priced at $0.15 per million, with storage at $1.00 per million per hour. According to the company, the model consumes 17% fewer tokens at output to complete the same multi-step agential flow, stacking a second cost reduction on top of the lower price.
Timing matters more than absolute value. Logan Kilpatrick, Senior Product Manager at Google DeepMind, announced the start of the company's "most ambitious pre-training" for Gemini 4 on the same day. This is the first time Google has publicly layered an active commercial model with a confirmed successor in the works, forcing OpenAI and Anthropic to decide whether to respond with price changes now or wait for the next generation of their rival to emerge.
The Numbers Anthropic and OpenAI Will Read First
In the benchmarks released by Google, the 3.6 Flash scores 49% on DeepSWE (up from 37% for 3.5 Flash), 63.9% on MLE Bench (previously 49.7%), 83.0% on OSWorld-Verified (up from 78.4%), and 1421 on GDPval-AA v2 (up from 1349). These gains are concentrated in three areas that the enterprise agent market pays dearly to resolve: software engineering, operating system usage, and multi-step execution with external tools.
At $7.50 per million output tokens, the Flash models compete in the same range as Claude Haiku 4.5 and GPT Luna, the cheapest tier of the trio Sol/Terra/Luna released by OpenAI on July 9. Google’s perspective is clear on the pricing page: the Flash is the default model for code pipelines, data extraction, and service agents where the cost per call dominates product economics.
What Changes for Those Already Purchasing Tokens at Scale
For a CIO running a support agent with 5 million conversations per month and 3,000 output tokens per conversation, the gross price drops from $135,000 to $112,500 annually based on the 3.5 Flash. Applying the 17% cut in generated tokens, the annual bill comes out to about $93,000, which is a difference of approximately 31% compared to the previous generation in the same scenario. This is the same math that drove Klarna to standardize Flash for ticket triage in 2025, according to a presentation by the fintech at Google Cloud Next.
The effect is not localized. In Accenture's delivery hubs in Bengaluru and Capgemini’s in Krakow, where most enterprise agent MVPs currently run on Vertex AI, Flash 3.6 becomes the inexpensive baseline reference that any architectural proposal must justify. In the UK, where Cohere has positioned its Command model for the sovereign market, the new price for Flash tightens the "European AI" narrative to defend margin against a hyperscaler that has become 31% cheaper in less than six months.
The Shadow of Gemini 4 is the Real Point
Kilpatrick’s announcement did not come with a timeline, benchmarks, or architecture details. It’s a message for the purchasing committee: anyone signing a three-year contract with a competitor will need to renegotiate before it expires if Gemini 4 delivers what marketing suggests. Sundar Pichai has already stated in the Q2 call that Gemini 3 Pro has "decisively" outperformed the rest of the market in the company's internal benchmark, and the next wave comes in the largest compute training that Google has ever financed.
It will be interesting to see how Anthropic and OpenAI respond. Anthropic promised during the press briefing for Claude Opus 4.7 in May that it "would not engage in a price race" and would prioritize margin over quality. OpenAI segmented the 5.6 line into three names precisely to avoid cannibalizing its own Sol when Luna needs to compete with Flash. If Gemini 3.6 Flash starts winning deals from SaaS providers that currently use GPT Luna as an agent backend, the segmentation will operate against OpenAI’s revenue rather than favor it.
The question Google did not answer in the announcement: if Gemini 4 takes 18 months to arrive, how to sustain the economic incentive of Flash 3.6 throughout that gap? The answer will come from the next pricing cycle, not from the next model.