stackwitness

Blog

Three AI outages on one morning are still three incidents

On 3 September 2026, ChatGPT, Claude, and Grok failed in overlapping windows. How to write the customer update without inventing one shared root, and without calling the whole industry down.

ChatGPT is erroring. Claude Code will not start. Grok says the model is overloaded. Downdetector is a wall of red. Is AI down? That question is too coarse. You need three clocks, three named causes, and the one dependency your product actually calls.

A recent case: 3 September 2026

Anthropic opened first. At 13:26 UTC on 3 September 2026, Claude's status page began investigating elevated errors on Mythos 5.1, Fable 5.1, and Opus 5. The incident listed claude.ai, the Claude API at api.anthropic.com, Claude Code, and Claude Cowork. A spokesperson later called it an infrastructure issue. Impact ended at 16:16 UTC. Independent writeups, including Ars Technica, put the first public report at 13:23 UTC.

xAI opened four minutes later. At 13:30 UTC, Grok's status page said Grok was experiencing issues. The incident ran 3 hours and 37 minutes, resolving at 17:07 UTC. xAI later named an outage at its Memphis compute center and apologized to compute partners, as reported by The Verge and The Register.

OpenAI opened more than an hour after the other two. At 14:43 UTC, OpenAI's status page began investigating elevated errors across ChatGPT and Codex, 15 ChatGPT components and 4 Codex components. A spokesperson told The Register the cause was a routing error starting around 14:43 UTC, with a solution in place by 15:17 UTC. The page marked resolved at 16:55 UTC.

Those are three vendor incidents with three start times and three stated causes. Overlap on a clock is not a shared-root verdict.

What a customer of those vendors actually did

Cursor felt Grok immediately. At 13:41 UTC, eleven minutes after xAI opened, Cursor posted its own incident: service degradation affecting all Grok models, Automations, Cloud Agents, Grok Bot, and Review Agents. They named the model family. They resolved at 17:07 UTC, the same minute xAI said traffic was healthy again. That is the operator move a headline like "AI is down" will not make for you.

Why overlapping windows fool people

Crowd-report charts spike together. Slack fills with "is the internet broken?" Someone names Azure. Someone names Cloudflare. The Register asked both questions in public: Cloudflare said it was not experiencing significant disruptions, and the AWS, Google Cloud, and Azure status pages showed no relevant incident. 9to5Google reported Microsoft saying Azure was not the cause. Gemini user reports also spiked; Google never opened an official incident.

xAI named Memphis and apologized to compute partners. Anthropic's own page said infrastructure issue and did not name a site. OpenAI named a routing error. Treating those three sentences as one cloud outage is how you page the wrong people and write a customer update nobody can check an hour later.

A single host probe of api.openai.com or api.anthropic.com would not have told you whether ChatGPT conversations, Codex, Claude Code, or a specific model family worked. Host reachability is host answer, not model health. On StackWitness truth pages (see /status, for example /status/openai and /status/anthropic), the witnessed column sits next to vendor self-report on purpose, and it is still a host-level signal. We do not invent probe outcomes for that past morning here.

Copy you can adapt

The weak version, which could be pasted onto any incident:

We are aware of a widespread AI outage affecting multiple providers. Third-party services are investigating. All other systems remain operational. We apologize for any inconvenience.

That paragraph costs the morning. On-call debugs your inference gateway. Support tells customers to retry. Nobody fails over to a model that is still answering, because the headline said the industry was dark.

The version that names each clock:

As of 13:41 UTC our Grok-backed agents are failing. xAI reports a Grok outage since 13:30 UTC. Anthropic reports elevated errors on Claude Code and the Claude API since 13:26 UTC; we are checking those paths separately. OpenAI has not yet opened an incident on ChatGPT or Codex. We will not mark this operational until the model families we call recover, and we will not blame a shared cloud we have not seen named.

A reader who was not on the bridge can check that against each vendor page and against your own error rates by model family. The weak paragraph cannot be checked, so it cannot be trusted an hour later.

What not to do

Operator checklist

  1. Name the model family and product surface you actually call (Grok agents, Claude Code, ChatGPT, Codex), not the industry.
  2. Open each vendor status page. Write three start times, not one.
  3. Copy the cause the vendor named, or write "cause not yet named." Do not borrow a cause from a neighboring vendor.
  4. If you are a downstream product, open your own incident and list the upstream components. Cursor's Grok-named incident is the pattern.
  5. Keep a timestamp trail: when your errors started, when each vendor opened, when you failed over, when each page resolved.

StackWitness measures host reachability and reads vendor self-reports so teams can attribute faster. Three overlapping outages are still three outages. Name the vendor you lost.

We measure reachability and read vendor self-reports. We never invent calm green or issue a compliance verdict.

Start free See live dependency truth