stackwitness

Blog

When the replicas detach and the host still answers

GitHub's 23 September 2026 incident started as detached database replicas: organization creation, the API, and Projects degraded while Git Operations stayed off the affected list. How to keep those clocks separate when the component board goes quiet before the incident closes.

git clone still returns. The REST call to create an organization times out. Project search is yesterday's board. Support is asking if GitHub is down. That question is too coarse. You need the component that failed, and whether the host answering is the same fact.

A recent case: detached replicas, 23 September 2026

On 23 September 2026 at 10:11 UTC, GitHub opened incident 8zc63m64hy36 for degraded API Requests. At 10:20 UTC they named the failure: database replicas had detached. Organization creation, the GitHub API, and Projects were degraded. Git Operations was not on the affected-component list.

Replicas were restored at 10:57 UTC. API Requests moved back to operational at 10:58 UTC. Projects recovery started at 11:38 UTC, with stale search while indexing caught up. At 13:35 UTC search was still stale. At 17:01 UTC, issue-label updates into Projects were delayed by about 10 minutes, with a few hours of backlog still to process. The incident stayed in investigating. GitHub's own page still listed it unresolved. isinternetup.com's 23 September digest, written around 12:00 UTC, had the same clocks: API mitigated at 10:58, Projects recovering at 11:38, incident still open.

An independent tracker that splits those questions, lagcheck's GitHub page, recorded 23 September so far as 100% up (the host answered) and 62% clear (the vendor was reporting a problem on the remaining checks). Unreachable, and a vendor saying a component is degraded, are different columns. Treat them as different columns.

What the quiet board hides

At 10:58 UTC the API Requests component returned to operational. From that minute, a skim of the component board looks like a quiet day. The incident was still investigating six hours later because Projects search was stale and label updates were delayed. Closing your customer incident when the board goes quiet is how you tell a PM that GitHub is fine while their board is missing this morning's labels.

The affected-component chips stayed on API Requests even after the 10:20 update named Projects. Reading only the chips misses the body. On 23 September 2026 at 17:10 UTC, after the API mitigation and while the Projects backlog was still processing, the StackWitness GitHub truth page showed host reachability as reachable and vendor self-report as operational. There was no independently measured reachability incident in the prior week. See /status/github, and the index at /status. Reachable means the host we probe answered. It is not a measurement of API error rates, organization creation, or Projects search. We did not invent probe results for the 10:11 to 10:58 UTC window beyond that: the host-level column stayed quiet through the replica event.

Three surfaces, not one GitHub

Git Operations, API Requests, and Projects are different products that share a brand. A clone succeeding does not mean POST /orgs succeeded. A 200 from github.com does not mean Project search is current. Folding those into one GitHub-is-down banner, or one GitHub-is-up banner, is how you page the git layer for an API replica fault, then close the ticket when the board turns operational while search is still yesterday. The shape is cousin to a control-plane incident (see when the control plane is down and the data plane is up). Today's unit is GitHub's replica set, not a cloud API versus running VMs.

Copy you can adapt

The weak version, which could be pasted onto any incident:

We are aware some users may be experiencing issues with GitHub. A third-party is investigating. All other systems remain operational. We apologize for any inconvenience.

That paragraph costs the morning. On-call pages the git layer. Support tells customers to retry clone. Nobody writes down that organization creation and the API degraded from detached replicas at 10:20 UTC, so the postmortem has a GitHub anecdote and no component. At 10:58 the board looks quiet, the ticket closes, and the PM spends the afternoon asking why Project labels are ten minutes late.

The version that names the component and the clock:

As of 10:20 UTC GitHub reports detached database replicas. Organization creation, the GitHub API, and Projects are degraded. Git Operations is not on the affected-component list. Independent host probes of github.com still answer; that is host reachability, not API health. We will not mark this recovered when the API Requests chip returns to operational if Projects search or label updates are still delayed.

A reader who was not on the bridge can check that against GitHub's incident page, against your own error rates by path, and against an outside host view. The weak paragraph cannot be checked, so it cannot be trusted an hour later.

What not to do

Operator checklist

  1. Name the component you actually call (API Requests, organization creation, Projects search, Git Operations), not the brand.
  2. Read the incident body, not only the component chips. Projects can be named in prose while the chip list stays on API Requests.
  3. Keep a second clock for catch-up work: replica restore, search indexing, label backlog. Mitigation of the API is not the end of the incident.
  4. Check an outside host view. Reachable means the host answered. Put it next to the vendor self-report, not in place of it.
  5. Write the customer update so a reader can check it against the vendor incident and against your own error rates by path.

StackWitness measures host reachability and reads vendor self-reports so teams can attribute faster. A detached replica is still an outage for anyone who needed the API or Projects. Name the component you lost.

We measure reachability and read vendor self-reports. We never invent calm green or issue a compliance verdict.

Start free See live dependency truth