On September 3, at 10:30 AM ET, the outages began. OpenAI alone received over 37,000 outage reports at its peak. The outage lasted roughly 90 minutes. Four rival AI labs, billions in valuation, and a shared infrastructure dependency that nobody wants to admit exists.
At 10:30 AM ET, ChatGPT error reports jumped from 5,000 to more than 22,000 in ten minutes. Claude peaked at 1,400 reports. Grok hit 1,400. Gemini logged hundreds more. By noon ET, services had mostly stabilized. But the root cause remains unclear.
The outage is not unique. What made it unusual was the coordination: three major AI services failed within roughly 90 minutes of each other, with no cloud provider declaring a fault. That pattern is hard to explain by coincidence.
37,000 Outage Reports in 90 Minutes—No One Knows Why
Three companies, three architectures, three distinct ownership structures. One common factor.
OpenAI is deeply integrated with Microsoft Azure. Anthropic is backed by Amazon and Google, but also leases compute from xAI's Colossus cluster in Memphis. Grok runs on xAI's own infrastructure. Yet all four services experienced simultaneous disruption.
Cardano founder Charles Hoskinson offered a provocative theory: "It looks like a nation state hit the three AI models at once." His reasoning: they all rely on Nvidia chips. Grok, Claude, and ChatGPT all utilize significant portions of Colossus 1 in Memphis—a single data center with roughly 500,000 Nvidia GPUs. Google uses its own TPUs and appears to have been the least affected, suggesting a possible link to Nvidia-dependent infrastructure.
Others pointed to Microsoft Azure, which experienced a concurrent spike in outage reports at the same time. OpenAI, Anthropic, and xAI all depend on Azure to varying degrees.
500,000 GPUs in One Data Center. That's the Risk.
The outage exposed a risk that has been quietly accumulating: the AI industry's infrastructure is highly concentrated.
Anthropic draws over 300 megawatts across 220,000 Nvidia chips at Colossus 1, just under half of xAI's roughly 500,000-GPU fleet. When two competitors sit on the same physical racks, a single data center issue can take down both.
Cloudflare also experienced issues that previously caused simultaneous xAI and ChatGPT outages in November 2024. The pattern is consistent: dependencies accumulate, redundancy is expensive, and concentration is cheap—until everything stops at once.
Only One Lab Published Details. The Others Stayed Quiet.
OpenAI was the only lab to confirm anything, logging elevated errors across ChatGPT and Codex and listing 19 affected components. Anthropic blamed its Opus models. xAI declared no incident, yet Grok told users its model was unavailable. The transparency gap is notable: one company published details, another posted vague status updates, and a third said nothing.
The incident raises a harder question for enterprises that rely on these services: if the providers themselves can't explain why they went down, how can customers trust the infrastructure?

The September 3 outage was not a catastrophic failure. Services recovered within 90 minutes. But the pattern—simultaneous failure across rival services, shared infrastructure, and opaque root causes—is a warning.
The AI industry is building the next wave of digital infrastructure on a foundation that is less distributed than it appears. One data center, one cloud provider, or one chip supplier can bring down multiple services at once. Concentration stays cheap until everything stops at once. And when it does, the providers may not be able to tell you why.
P.S. The outage came just days after OpenAI announced Astra—its first "Critical" cybersecurity model capable of autonomously finding zero-day vulnerabilities. The juxtaposition was not missed. One model finds vulnerabilities. A shared infrastructure failure exposed one of its own. The question is not whether AI can hack. It's whether the infrastructure running it can stay online.
Frequently Asked Questions
Q: What happened on September 3?
A: ChatGPT, Claude, Grok, and Gemini all went down at roughly the same time. OpenAI alone received over 37,000 outage reports at its peak. The outage lasted about 90 minutes before services stabilized.
Q: How many outage reports were there?
A: ChatGPT peaked at over 37,000 reports. Claude peaked at about 1,400. Grok hit about 1,400. Gemini logged hundreds more. The reports surged within a 10-minute window.
Q: What caused the outage?
A: The root cause remains unclear. No cloud provider declared a fault. Theories include a shared dependency on Nvidia GPUs or a single data center (xAI's Colossus 1 cluster), or potential infrastructure issues with Microsoft Azure.
Q: What is the "shared infrastructure" theory?
A: OpenAI, Anthropic, and xAI all depend on a shared pool of Nvidia GPUs—including xAI's Colossus 1 data center in Memphis, which contains roughly 500,000 GPUs. Google uses its own TPUs and was the least affected, suggesting the issue may be tied to Nvidia-dependent infrastructure.
Q: Were all four labs affected equally?
A: No. OpenAI was hit hardest. Claude and Grok had moderate spikes. Gemini was the least affected, suggesting Google's TPU-based infrastructure may be more resilient to this particular type of failure.
Q: What did the labs say about the outage?
A: OpenAI was the only lab to publish detailed status updates, listing 19 affected components. Anthropic blamed its Opus models. xAI declared no incident, even though Grok told users its model was unavailable. Google did not provide a public explanation.
Q: Is this the first time this has happened?
A: No. Cloudflare experienced issues that previously caused simultaneous xAI and ChatGPT outages in November 2024. The pattern is consistent: shared infrastructure dependencies create concentrated risk.
Q: What is the "single point of failure"?
A: The AI industry's infrastructure is highly concentrated. One data center, one cloud provider, or one chip supplier can bring down multiple services at once. Concentration is cheap until everything stops at once.
Q: What does this mean for enterprises?
A: The outage raises a harder question: if providers themselves can't explain why they went down, how can customers trust the infrastructure? Enterprises relying on AI services should consider redundancy and multi-provider strategies.
Q: What is the significance of the timing?
A: The outage came just days after OpenAI announced Astra—its first "Critical" cybersecurity model capable of autonomously finding zero-day vulnerabilities. The juxtaposition highlights the gap between AI's capabilities and the infrastructure supporting it.
