Infrabench / Outage replay
Tool 04 — Incident communication

The dashboard was green. How long had it been wrong?

A status page is a published spec like any other, and it fails the same way: optimistically. Every one of these six incidents was already breaking customers before the provider posted a word, and every one of them is documented well enough to measure the gap. Set the threshold your own monitoring pages at and see which outages you would have heard about first. The arithmetic is in the method.

The replay

Incident to replay
Replay .
My own monitoring pages someone minutes after users start failing.

The detection gap
What was happening against what the status page showed
What the console said, minute by minute
All six, ranked by how long they stayed dark
The decision
Why the tool exists

A status page is a press release with a clock on it.

01

It is written, not measured

Nothing on a status page is emitted by the failing system. A human decides that the blast radius is now large enough to admit to, writes a sentence, and gets it approved. That pipeline has a floor of several minutes on its very best day.

02

It often shares fate with the outage

In December 2021 AWS posted that the incident was "affecting some of our monitoring and incident response tooling, which is delaying our ability to provide updates." The instrument was downstream of the thing it was measuring. That is not rare.

03

You are on the clock before it moves

Your users started failing at the first timestamp, not the second. Every minute of that gap is a minute you spent deciding whether it was you, and the only instrument that could have told you sits inside your own account.

Median gap before the page moved
Longest documented gap
Total impact replayed here
How it works

Two clocks, one subtraction.

01 · The first clock

When customers started failing. Taken from the provider's own retrospective, which is written after the fact with access to internal telemetry and is consistently the earliest honest number anyone publishes. AWS puts its October 2025 start at 11:48 PM PDT. Cloudflare puts its November 2025 start at 11:20 UTC.

Nobody outside the provider could have known this time on the day. That is the point of using it.

02 · The second clock

When the status page first said anything at all. Not when it was accurate, not when it named a root cause, just the first public acknowledgement that something was wrong. It is the friendliest possible reading of the provider's behaviour, and the gap is still measured in tens of minutes.

03 · The gap

Second clock minus first clock. That is the window in which your dashboard was green, your users were not, and no external source would have backed you up. Everything else on the page is context around that one number.

04 · Your threshold against it

Your own paging delay, subtracted from the gap. A positive margin means you would have been first and the status page later confirmed what you already knew. A negative margin means you were still waiting when they told you, which is the only outcome where the status page was worth anything as an alerting tool.

Worked example — Google Cloud, 12 June 2025
The incidents

Six that are documented well enough to argue with.

Chosen because each one has a published retrospective that states when impact began, which is the number nobody can reconstruct from outside. Incidents where that timestamp was never published are not in here, however famous they are.

Every incident replayed by this tool, with its detection gap, total impact and source
IncidentImpact beganPage movedGapTotalWhat broke
Method

Every timestamp, and where it came from.

The model5 lines
gapfirst public status post − first customer impact
margingap − your paging threshold
darkimpact → first post. Console green, users failing
openfirst post → mitigation. Acknowledged, unresolved
recoveringmitigation → all clear. Backlog draining
What this does not measure
  • Whether the post was true. The gap closes the moment anything is admitted, however vague. A page that says "investigating elevated error rates" during a total regional failure counts as having moved.
  • Your actual blast radius. Two accounts in the same region had wildly different days. This measures information, not impact.
  • Malice. Every gap here has a mundane explanation, and in at least two cases the tooling needed to update the page was inside the failure. Slow is not the same as dishonest.
  • All times are UTC, converted where the source published local time. The provider's own wording is quoted verbatim in the log where it was published; where a post was paraphrased in secondary coverage it is marked as reported rather than quoted.
  • Impact start always comes from the retrospective, never from a third-party detector. Downdetector spikes and traffic-graph dips appear in the log as context because they are what an outsider actually saw, but they never set the clock.
  • Two acknowledgement timestamps are weaker than the rest and are labelled as such in the table: the AWS December 2021 first dashboard post, and the Azure Front Door first advisory. Both are taken from secondary reporting rather than a surviving status-page archive. If you have the primary record for either, the whole page is one edit away from being corrected.
  • Nothing here is scraped live. These are historical incidents with closed retrospectives. The page ships the data it draws, so it renders identically offline and cannot drift when a vendor reorganises its status history.
Built by
Adrian Mucha

Product designer for the software most designers avoid: operator consoles where a wrong click costs real money.

Focus
Data centers · AI tooling
Base
Los Angeles, CA
Status
Open to new work
Copied