Strategy5 minNewsroom

OpenAI Unveils Jalapeño Benchmarks, Claims Efficiency Edge Over Nvidia

Corredor de data center ao amanhecer com dois corredores de racks lado a lado e um técnico solitário caminhando entre eles.

First public numbers from OpenAI's chip in partnership with Broadcom show up to 1.9x more work per watt than Nvidia's Blackwell in InferenceX. Volume won’t arrive until 2027, and Nvidia announces results today.

OpenAI published its first benchmark results for Jalapeño, an inference chip developed in collaboration with Broadcom, on Tuesday, August 25. Measured in InferenceX, a public benchmark maintained by SemiAnalysis, the chip delivered between 1.5 to 1.9 times more work per watt than Nvidia's Blackwell systems, and reduced end-to-end latency by 1.7 to 3.6 times. In conversational workloads like those generated by ChatGPT, the results showed up to a 4.1 times advantage.


The announcement was made by Sam Altman, CEO of OpenAI, and Hock Tan, CEO of Broadcom. Charlie Kawwas, President of Semiconductor Solutions at Broadcom, physically handed over the first unit to Altman and Greg Brockman, President of OpenAI. The system was designed by OpenAI based on the requirements of its models, kernels, and serving systems; Broadcom and Celestica handled the industrialization, rack integration, high-performance networking, and production volume.


What Changes for Those Funding Inference


Jalapeño is an inference chip, not a training chip. This distinction matters. Training is a concentrated expense occurring over short windows, with investments amortized over years. Inference is a recurring expense, executed in milliseconds, which grows every time a user queries ChatGPT. If each query costs a smaller fraction of watts, the product's operational margin changes.


According to Tan, the collaboration points toward "gigawatt-scale data centers with Microsoft and other partners starting in 2026." OpenAI expects to deploy Jalapeño in small volume by the end of this year, with larger scale in 2027. The company did not disclose the unit cost. Independent analysts who replicated the tests indicated significant savings in cost per query, although the final figure depends on the price OpenAI decides to allocate to the chip on an internal accounting basis.


Altman's comment was succinct: "We made the chip, and it’s fast."


Four Nvidia Clients Now Making Their Own Chips


Google has been running TPUs for a decade. Microsoft announced Maia 200 this year. Meta is accelerating the MTIA for internal inference. Amazon runs Trainium and Inferentia on AWS. With Jalapeño, OpenAI joins this club, doing so while continuing to be the largest individual buyer of H200 and Blackwell GPUs on the planet.


The short takeaway is that Nvidia is under attack. The more accurate reading is different. The four hyperscalers have built their own silicon for specific use cases, not to fully replace Nvidia. Google runs most of its inference workloads on TPUs but still buys H100 in volume. Microsoft uses Maia for selected OpenAI workloads but allocates the bulk of its capacity to Nvidia GPUs. The emerging pattern is one of a mix, not replacement.


Nvidia Reports Today


Nvidia will release its fiscal second-quarter 2027 results after the close of the U.S. market on Wednesday. Analyst consensus points to $92 billion in revenue, up 97% year-over-year, with the data center segment nearing $85.7 billion. Microsoft, Alphabet, Amazon, and Meta collectively spent $166 billion in capex in the June quarter, an 87% increase year-over-year and 27% from the previous quarter. The portion of this money that translates into Nvidia revenue is the figure that Wall Street will focus on first.


Jalapeño arrives on the eve of this reading. The question analysts should ask in the call: what share of OpenAI’s inference currently runs on Nvidia and will migrate to Jalapeño by 2027? What remains on the revenue chart for the chipmaker, and what exits? And whether Broadcom, currently classified as "networking and custom silicon" in its P&L, will begin to compete directly with Santa Clara for dollars that used to flow entirely to there.


What Inference Economics Now Reprices


For the CIO who allocates the AI infrastructure budget, Jalapeño reshapes the short-term decision. Not because there is a chip available for purchase in 2027 outside of OpenAI, but because it proves, with public benchmark numbers, that the hypothesis of Nvidia's efficiency monopoly in inference does not hold. Service providers, especially in Europe and Japan, are already discussing how to sell inference in sovereign clouds with competitive margins. Data centers in Brazil designed today for AI workloads need to consider this option in the upcoming refresh cycles.


The operational point is that the cost per token has dropped as a structural phenomenon. Not as a commercial campaign.

The week's analysis, by email

One weekly edition with what matters to people who decide. No ads, no sponsorship.

One-click cancellation, at any time.

Strategy