Security & Risk6 minNewsroom

Failure in Azure East US Takes Down ChatGPT, Claude, and Grok Simultaneously for 90 Minutes

Técnico em telhado de prédio em Manhattan ao amanhecer, próximo a equipamento de refrigeração com LED vermelho aceso.

Three of the four largest chatbots in the world experienced outages on Thursday morning within the same timeframe due to sharing the same Microsoft region. Gemini on Google Cloud remained operational.

A failure in Microsoft's Azure East US region took ChatGPT, Claude, and Grok offline almost simultaneously on the morning of Thursday, September 3rd, within a window of approximately 90 minutes. Downdetector recorded over 37,000 complaints against OpenAI, about 1,365 against Grok, and 1,300 against Claude at the peak of the incident. Gemini from Google, hosted on Google Cloud, saw the number of complaints stop at around 500 and continued to operate normally.


This incident is not a fleeting scare. It serves as empirical evidence, within a single timeframe, of a concentration of risk that infrastructure executives have been treating as hypothetical. OpenAI has been operating on Azure since Microsoft's multibillion-dollar investment and runs a substantial portion of its inference traffic there. Anthropic employs a multi-cloud strategy across AWS, Google Cloud, and Azure, but a significant share of Claude's production traffic still routes through Azure-related paths. xAI, despite operating Colossus in Memphis with its own hardware, routes outgoing traffic through the same backbone. Three competing labs, one regional failure.


What the Post-Mortem Will Probably Show


Microsoft has yet to publish a final analysis. The pattern from the last series of incidents, including the global outage in July that affected Xbox and Microsoft 365, points to a bottleneck in centralized identity or network services, rather than in a specific chip or physical datacenter. This is the category of failure that generates regional cascading effects: the East US region concentrates a disproportionate volume of inference workloads because it combines availability of H100 and H200 GPUs, favorable latency for the East Coast, and direct integration with the Azure OpenAI Service.


Copilot, Microsoft's own assistant, also experienced instability within the same window, despite running on internal Azure infrastructure. This reinforces the reading that the failure impacted a shared layer, not isolated compute capacity.


The Bill That Comes to the CIO on Monday


For enterprise customers running critical pipelines on a single LLM provider, the incident on September 3rd compels the decision many have been postponing. Anthropic is trying to capitalize by offering automatic failover between AWS Bedrock and Google Vertex; OpenAI has not yet provided the same flexibility for its newly launched Astra model. In practice, an enterprise workflow relying on ChatGPT via API for customer support, contract analysis, or code generation was left without a useful response for an hour and a half, without an automatic credit SLA capable of compensating for business interruption.


Gartner projected in July that 40% of companies with revenues above $1 billion are already operating AI agents in production. A 90-minute downtime at this layer creates cascading effects in CRM tools dispatching to Agentforce, in code review pipelines relying on Copilot or Claude Code, and in support with third-party voice bots. The direct cost is difficult to estimate without public numbers, but analysts from Constellation Research estimate an aggregate revenue loss between $40 million and $80 million among the three largest enterprise clients of OpenAI that day.


Germany and Brazil in the Same Trap


Deutsche Bank, which announced in May the standardization of code assistants in Claude for 30,000 developers, and Deutsche Telekom, which runs customer support automation on ChatGPT Enterprise, experienced the same window of unavailability in European operations because East US is the primary failover region for peaks coming from the EU. In Brazil, Itaú and Bradesco, which route part of the BIA and Personnalité Assistant through Azure endpoints, saw the same simultaneous outage, although without a perceptible impact on the end customer due to local cache layers.


What September 3rd makes clear is that the multi-cloud strategy of model providers does not protect the enterprise customer when that customer chooses a single model provider. The resilience layer needs to move up a level: from infrastructure to model routing. Companies like Portkey, Kong AI Gateway, and OpenRouter, which were acting as routing commodities, gained immediate commercial ammunition they did not have on Tuesday.

The week's analysis, by email

One weekly edition with what matters to people who decide. No ads, no sponsorship.

One-click cancellation, at any time.

Security & Risk