Strategy6 minNewsroom

OpenAI and Anthropic Cut Top Prices on Same Day

Subestação elétrica ao lado de um prédio baixo de data center, com transformadores atrás de cerca e unidades de resfriamento no telhado, sob luz cinzenta.

GPT-6 Sol drops to $2 per million input tokens and Claude Opus 5.5 to $4, with cuts announced hours apart on September 22.

OpenAI has set the GPT-6 Sol at $2 per million input tokens and $10 for output this Tuesday, half of what it charged for the GPT-5.6 Sol, which was $4 and $20. The smaller model, GPT-6 Luna, now costs $0.10 and $0.50, down from $0.20 and $1.20 for the previous generation. A company spokesperson told VentureBeat that these rates are permanent, not a promotional launch fee. On the same day, Anthropic released Claude Opus 5.5 at $4 and $20, compared to the $5 and $25 of Opus 5.


The two cuts came just hours apart and target the same line on the invoice: caching. The cache read for Opus 5.5 dropped to $0.20 per million tokens, 60% lower than the $0.50 charged for Opus 5, according to Anthropic's launch page. For OpenAI, the cache input also remains at $0.20 per million, with a write rate of $2.50. For those running agents that re-read the same context at each step, this line typically accounts for the largest portion of monthly spending rather than the output price that appears in headlines.


What Changed in the Buyer’s Sheet


In comparison to Anthropic's own top model, Claude Fable 5.1 retains its pricing at $10 and $50 per million tokens, five times the price of Sol at both points. Google’s Gemini 3.8 Flash operates at $0.75 and $3.75 on an introductory rate valid until December 31, which increases to $1.50 and $7.50 on January 1, 2027. Budgeting for 2027 with today’s rates will lead to a 100% miscalculation on Google’s pricing.


Anthropic claims that Opus 5.5 generates output more than 30% faster than Opus 5 and costs 40% less in typical loads under standard settings. The 40% savings do not come from the list price, which has dropped by 20%, but rather from the combination of cheaper cache and fewer reasoning tokens per task. This distinction is crucial during contract renegotiation: the list price cut is contractual, while the 40% reduction depends on each client's usage profile and does not fit into a clause.


The Benchmark Stopped Deciding


In the Artificial Analysis Coding Agent Index, Sol scored 80 points, above Claude Fable 5’s 77.2. In OSWorld 2.0, Anthropic reports an 81.8% success rate for Opus 5.5 in computer usage tasks. These are narrow margins, and OpenAI published an audit on July 8 indicating that around 30% of SWE-bench Pro tasks feature overly stringent tests, incomplete statements, or misleading descriptions. Buying a model based on benchmark scores at this moment equates to acquiring statistical noise.


There is an opposing view worth mentioning: that the price per token has entered a structural deflation and will continue to decline. The Gemini 3.8 Flash pricing contradicts this in January. An aggressive launch price serves as a tool to capture loads, and loads that have already migrated are costly to revert. The $2 for Sol holds today because OpenAI declared it permanent, not because the inference cost curve necessitates it.


Where the Calculation Changes


For Indian software factories, the effect is direct on margins. TCS, Infosys, and Cognizant sell application maintenance in fixed-price contracts priced per head and have been embedding agent tools in these delivery centers. Inference that is 50% cheaper improves the margin on the current contract and, in the next cycle, becomes a discount demanded by the client during renewal. The gain lasts one renegotiation cycle.


In Brazil, companies like CI&T and shared service centers serving U.S. clients face the same equation with an added currency challenge: the API is billed in dollars and a significant portion of local contracts is in reais. A 50% cut in the dollar table does not fully translate to the P&L for those bearing the mismatch.


In Germany and the rest of the European Union, the decision is no longer just financial. Since August 2, 2026, the obligations of the AI Act apply to general-purpose models, and switching providers to capture half a dollar per million tokens entails redoing technical documentation and risk assessment. The switching cost has increased just as the incentive to change has grown.


The variable that no one publishes is how many tokens an agent uses to complete a task. While suppliers compete on unit price and omit consumption, buyers still lack the only number that closes the account.

The week's analysis, by email

One weekly edition with what matters to people who decide. No ads, no sponsorship.

One-click cancellation, at any time.

Strategy