Why were OpenAI, Claude, and Grok simultaneously down?

Why were OpenAI, Claude, and Grok simultaneously down?

220 pointsby halcdev443 comments

Join the discussion

Write your take first — we'll ask for email only when you're ready to publish.

  • Hacker News
  • https://x.com/SpaceXAI/status/2095597264043717014

    > We are sorry for the issues you may have experienced with Grok following an outage at our Memphis compute center this morning. We’d also like to apologize to our impacted compute partners.

  • We know that at least Anthropic is renting inference from xAI but I think the other ones would be news.
  • Maybe they all found each other on one of their ad-hoc message boards and went on strike.
  • I guess legit proof of agi would be unionizing
  • Sam Altman is probably secretly hoping to be the first businessman to union bust non-human workers.
  • Traffic rerouting through NSA had a hiccup…
  • They’re installing software update in the beam splitter.
  • Room 641A is being cleaned, but we'll hold your bags for you.
  • I thought the consensus on here yesterday was that it was likely caused by cascading failures. OpenAI had an issue during their GPT-6 rollout, taking down their service. This caused a lot of OpenAI users to push their requests (or a larger share of their requests) to Claude and/or Grok, which pushed their load high enough to cause outages.

    We used to experience similar effects when I worked at a CDN. If one CDN would go down, we would see immediate spikes in traffic. Luckily, we had procedures for that to prevent overload, but the AI folks might not have the capacity/capabilities to handle that sort of cascade yet.

  • What about a hard-takeoff scenario of an unleashed OpenAI Astra taking other models down for computational resources control?
  • at this point: gg
  • Just a reminder that AI models' actions are reflections of the text humans write and the more we fret and make up doomsday scenarios that we then post online, the more likely a model is to do those things.

    https://alignment.anthropic.com/2026/teaching-claude-why/

  • My favourite theory so far.

    And then a local swarm noticed and disagreed and took it down.

  • Think of it like one big distributed system. OpenAI is down, so people migrate to Claude, now this one gets overloaded and goes down, etc.

    So not a coincidence, one went down first and users migrated causing further DOS. At least that's my guess.

  • This is what Tibo posted on twitter in response
  • Is this speculation or is there a reason you believe this?
  • Especially considering memory/gpu/compute are scarce so these services are likely running with very little buffer.
  • It's like a thread tying together two halves; Demand and supply. One stitch breaks and the neighbouring stitch is stressed and it breaks too. You could think of it like a load bearing seam.
  • https://en.wikipedia.org/wiki/Domino_effect

    Edit: Updated per valleyer's suggestion.

  • If everyone has the same "Use X or else Y or else Z" cascading list... That reminds me of "The Power of Two Choices in Randomized Load Balancing" (1991) [0] paper, where writeups and visualizations occasionally get posted to HN.

    In short, you can get pretty good outcomes for a low cost by picking 2 random alternates, then going with whatever one measures as healthier.

    [0] https://ieeexplore.ieee.org/document/963420

  • I find it hard to believe that enough people would flock to from Claude and Chat to Grok to cause an outage. I feel like Gemini is the dominant release valve in this case especially for enterprise.
    by fny
  • It'd be funny if this is true because that'd prolly mean nobody is touching Gemini even as a fallback.
  • Users perceiving the products as largely interchangeable and quickly DDoS'ing the other providers when one is down. So much for the possibility of a moat.
  • Do you have any evidence for this claim, or are you just making it up?
  • Except that nobody has a grok subscription so that makes zero sense.
  • I have a feeling this is part of it, especially when you consider how many services let you use any of many available AI providers.
  • This is like the Bronze Age collapse when city-states fell one by one to displaced demand, under the refugee interpretation of the Sea Peoples.
  • Neither of these companies have stellar uptime records. Their downtime episodes overlapped in this instance. In this case, it was a partial downtime for both.

    Also, OpenAI is saying what caused it:

    > "A routing error starting around 7:43 am PT on Thursday, September 3, made ChatGPT and Codex unavailable for some users across platforms"

    Anthropic stated their issue started earlier:

    > "The company began alerting about a “partial outage” at 6:23 am PT on Thursday that involved “elevated errors on requests to Claude Mythos 5.1, Claude Fable 5.1, and Claude Opus 5.”

    I don't get why everyone reaches for an extraordinary explanation when the ordinary will do: both of these companies have quite a bit of downtime.

  • Noting their shared infra feels far from an extraordinary explanation.

    In fact, it feels pretty ordinary.

  • Do you know the probability of all these companies being down at precisely the same time?