acme · incident INC-2208 · post-mortem

The outage lasted 26 minutes. The wrong theory lasted 19.

acme/checkout returned 500s to roughly one cart in three for 26 minutes. The cause was a database connection pool cut from 20 to 5 by a config merge in acme/identity. Nobody looked at it until minute 22, because the only recent change on the checkout deploy board was a payment provider adapter — and rolling that back at 14:19 changed nothing at all.

SYSTEM · acme/checkout 5xx FROM 14:06 TO 14:32
● HEALTHY 0.02% 5xx
✗ 26 MIN DEGRADED · 41% 5xx PEAK
● RECOVERED

14:19 — wrong fix shipped. 5xx flat at 39%.

14:09 → 14:28
○ 3 MIN UNSEEN · PAGE AT 14:08
19 MIN · WRONG CAUSE
◆ CAUSE FOUND 14:28
✓ POOL 5→20 AT 14:30
RESPONDERS · 2 on call ON THE REAL CAUSE FROM 14:28

Both tracks run on the same clock, 14:00 to 14:40 UTC. The hatched band is not missing data: it is 19 minutes in which two engineers worked continuously, on the wrong service. The two solid segments on the right are everything that actually ended the incident.

Degraded
26min

14:06:12 to 14:32:04.

On the wrong cause
19min

73% of the outage, spent inside acme/checkout.

Everything else
7min

Detection, the real diagnosis, the fix and recovery, added together.

DB pool after merge
5was 20

Shared by checkout and identity. 48 workers wanted one.

Failed carts
1,412

Checkout attempts that got a 500 in the window.

Both clocks, minute by minute

14 entries · UTC

Left column is what the system did. Right column is what people did. They are the same minutes, which is the point.

TimeSystemResponders
13:41Deploy #4471 to acme/checkout: payment adapter v2.3. No error change.—
13:52PR #2208 merged in acme/identity, touching the shared DB chart: pool 20 → 5.—
14:04Checkout pods roll on a routine autoscale event and pick up pool = 5.—
14:06Connection wait timeouts. 5xx 0.02% → 18%.—
14:085xx 38%.Page fires: checkout-5xx-burst.
14:095xx 39%.r.okonjo acks, opens the checkout deploy board.
14:115xx 39%.Narrows to deploy #4471. Payment adapter becomes the theory.
14:14acme/storefront-web shows a generic retry card; support queue +38 tickets.Reading adapter logs for provider timeouts. None found.
14:165xx 41%, the peak.m.haas joins. Both engineers on the adapter.
14:195xx 39%. Unchanged.Adapter rolled back to v2.2.
14:225xx 40%.Rollback confirmed complete. Theory kept anyway.
14:285xx 40%.m.haas greps merges outside acme/checkout. Finds pool = 5.
14:30Pool restored to 20; identity and checkout redeploy.Fix applied.
14:325xx 0.02%. Recovered.Incident closed.

Why the wrong theory held for 19 minutes

3 reasons

The only change anyone could see was in the service that was failing

cost: 19 min

The deploy board a responder opens from the page is scoped to acme/checkout. The change that caused the incident was merged in acme/identity and never appeared on it. From 14:09 the search space excluded the answer.

Why it matters: the adapter deploy was 25 minutes old and completely innocent. It was the best available suspect only because it was the sole visible one.

The rollback disproved the theory at 14:19 and the theory survived it

cost: 9 min

Error rate was 39% before the adapter rollback and 39% after it. That is a clean negative result at minute 13 of 26. The next nine minutes went to confirming the rollback had really landed, then to re-reading the adapter logs.

Why it matters: this is the cheapest 9 minutes on the page. No new tooling was needed to recover them, only the habit of dropping a theory the evidence has already killed.

Pool saturation was at 100% the whole time, on a dashboard nobody opened

not visible

Connection pool utilisation is graphed in the identity dashboard, not the checkout one. It went to 100% at 14:06:12 and stayed there for 26 minutes, four seconds before the first 500.

The actual change

acme/identity · PR #2208

One line in a chart that acme/checkout imports and acme/identity owns. Merged 13:52, titled chore: align pool defaults with the sample config, reviewed by one person, labelled config-only. It was correct for identity, whose pods are small and many. Checkout runs 48 workers per pod.

  charts/_shared/db/values.yaml

- DB_POOL_MAX: 20
+ DB_POOL_MAX: 5
Workers wanting a connection at peak 48
Pool before the merge 20
Pool after the merge 5

Checkout survived from 14:04 to 14:06 on cached connections. The pool was exhausted the moment the first sustained traffic arrived.

What would have cut the 19 minutes

3 actions

Put every merge touching a shared chart on the checkout deploy board

−17 min

PR #2208 would have been the second entry on the board at 14:09, one line under the adapter deploy, with the diff visible. Owner: platform. Two days of work.

Graph pool saturation beside the 5xx rate, not in another service's dashboard

−12 min

A flat line at 100% next to the error rate names the failure mode without naming the change. Owner: checkout. One afternoon.

Treat a fix that does not move the metric as a disproof, on a 3-minute timer

−9 min

The adapter rollback landed at 14:19 and the error rate did not move. Make that a checkpoint in the incident template: if the metric is flat three minutes after a fix, the theory is dead and the search widens. Owner: incident process. Costs nothing.