Strategy7 minNewsroom

Google Launches Gemini 3.8 Flash at $0.75 per Million Tokens and Reserves Cyber Variant for Selected Clients

Palco vazio de lançamento com emblema Gemini projetado em luz azul contra parede escura, púlpito no primeiro plano.

The third Flash version in six weeks comes with a promotional price until December 31 and outperforms Claude Opus 5 in three of the five benchmarks released by Google itself. Immediate pricing effect on agents.

Google launched the Gemini 3.8 Flash on September 2, marking the third iteration of the Flash line in six weeks, and also released a hardened variant called 3.8 Flash Cyber, with restricted access for clients who pass a specific screening process. The key point for those setting autonomous agent budgets is the pricing table: $0.75 per million input tokens and $3.75 per million output tokens valid until December 31, 2026, with automatic doubling to $1.50 and $7.50 on January 1, 2027. Batch and Flex are priced at half those rates, while Priority is charged at 1.8x.


The commercial design is aggressive. Flash 3.7 was already a cost benchmark for large-scale pipelines, and a successor Flash priced below the announced cap of the previous tier is a clear move to capture workload away from Claude Fable 5.1, whose caching read rate dropped to $0.25 per million tokens at the same start in September. The promotional pricing window until December gives operators three months to migrate workloads before repricing, pushing the architectural decision to the fourth quarter instead of the first quarter of 2027.


The Comparison Google Chose to Publish


The launch table positions the 3.8 Flash against its own 3.7 Flash and Claude Opus 5 across five benchmarks. The numbers published by Google: DeepSWE v1.1 at 73.7% versus 65.3% for Opus 5, Terminal-bench 2.1 at 89.4% versus 85.8%, OSWorld-2.0 at 59.0% versus 50.6%, Vals Finance Agent v2 at 61.4% versus 59.0%, and HLE-Verified at 54.9% versus 53.6%. In three of the five (DeepSWE, Terminal-bench, and OSWorld), the gap is significant enough to fall outside the noise margin; in Vals Finance Agent and HLE-Verified, the advantage is statistically marginal and will be contested as soon as Anthropic and OpenAI publish their own tables.


Google emphasizes that the 3.8 Flash was built on top of 3.7 Flash, not on a new base model. The company states in its own documentation that the gain comes from greater token usage during reasoning and recommends that those optimizing for pure cost remain on 3.7 Flash. This is the first time Google has taken this segmentation approach based on consumption patterns, signaling commercial maturity: Flash shifts from a single tier to a family with distinct curves.


What Changes for Consulting Firms and Banks


The reading must be made in at least two markets. In the United States, integrators like Accenture Federal, Deloitte Consulting US, and IBM Consulting already operate a vertical of 'agentic finance' under federal and commercial contracts, and the pricing table of Gemini 3.8 allows for repricing proposals in the current RFP cycle with gross margins 30% to 40% higher than was possible with Flash 3.7 at the same throughput. In European banks, the dominant variable is regulatory: BaFin and the French ACPR have been requiring that any model operating in the credit decision chain undergo evaluations under Articles 13 and 14 of the AI Act, which delays the adoption of 3.8 Cyber in production despite its competitive pricing. The practical result is asymmetry: American banks accelerate, while European banks revert to pilots.


In India, TCS and Infosys, which repriced AI-managed contracts in dollars during the second quarter, gain space for aggression. The Flash 3.8 in the promotional table fits into a 24x7 support SLA without exceeding budget, reducing the price floor at which these firms can compete with Western integrators. In Brazil, providers like CI&T and Stefanini will have two quarters to explore the same window before the table doubles in January, and the decision to keep workloads on-cloud or move them to dedicated on-premise inference returns to the top of the technology agenda.


The Counterpoint That Needs to Stand


Benchmarks published by the vendor are always tailored in favor of the vendor. DeepSWE, Terminal-bench, and OSWorld are useful but have high variance between executions, and HLE-Verified is a closed Google benchmark. Independent analysts like Artificial Analysis and Aider typically publish replicas within two weeks of the launch; until then, technical architects conducting vendor-neutral comparisons should not replace Claude or GPT solely based on the launch table.


The Underlying Signal in the Announcement


Three Flash launches in six weeks indicate a product cadence that only makes sense if Google is trying to offload excess inference capacity before it becomes a burden on the accounting books at the end of the fiscal year. The promotional table until December reinforces this hypothesis. Those setting budgets for 2027 should assume that the price in January is the real price, and that the discount from September to December is a temporary migration subsidy, not a new baseline. Annual contracts signed now with a lock until December are a good hedge; contracts extending to the second quarter of 2027 without an explicit lock carry risk.

The week's analysis, by email

One weekly edition with what matters to people who decide. No ads, no sponsorship.

One-click cancellation, at any time.

Strategy