acme · incident INC-2208 · post-mortem
The outage lasted 26 minutes. The wrong theory lasted 19.
acme/checkout returned 500s to roughly one cart in three for 26 minutes.
The cause was a database connection pool cut from 20 to 5 by a config merge in
acme/identity. Nobody looked at it until minute 22, because the only
recent change on the checkout deploy board was a payment provider adapter — and
rolling that back at 14:19 changed nothing at all.
14:0014:1014:2014:3014:40
SYSTEM · acme/checkout
5xx FROM 14:06 TO 14:32
● HEALTHY 0.02% 5xx
✗ 26 MIN DEGRADED · 41% 5xx PEAK
● RECOVERED
14:19 — wrong fix shipped. 5xx flat at 39%.
○ 3 MIN UNSEEN · PAGE AT 14:08
19 MIN · WRONG CAUSE
◆ CAUSE FOUND 14:28
✓ POOL 5→20 AT 14:30
RESPONDERS · 2 on call
ON THE REAL CAUSE FROM 14:28
Both tracks run on the same clock, 14:00 to 14:40 UTC. The hatched band is not missing
data: it is 19 minutes in which two engineers worked continuously, on the wrong service.
The two solid segments on the right are everything that actually ended the incident.
Degraded
26min
14:06:12 to 14:32:04.
On the wrong cause
19min
73% of the outage, spent inside acme/checkout.
Everything else
7min
Detection, the real diagnosis, the fix and recovery, added together.
DB pool after merge
5was 20
Shared by checkout and identity. 48 workers wanted one.
Failed carts
1,412
Checkout attempts that got a 500 in the window.
Both clocks, minute by minute
14 entries · UTC
Left column is what the system did. Right column is what people did. They are the same
minutes, which is the point.
| Time | System | Responders |
| 13:41 | Deploy #4471 to acme/checkout: payment adapter v2.3. No error change. | — |
| 13:52 | PR #2208 merged in acme/identity, touching the shared DB chart: pool 20 → 5. | — |
| 14:04 | Checkout pods roll on a routine autoscale event and pick up pool = 5. | — |
| 14:06 | Connection wait timeouts. 5xx 0.02% → 18%. | — |
| 14:08 | 5xx 38%. | Page fires: checkout-5xx-burst. |
| 14:09 | 5xx 39%. | r.okonjo acks, opens the checkout deploy board. |
| 14:11 | 5xx 39%. | Narrows to deploy #4471. Payment adapter becomes the theory. |
| 14:14 | acme/storefront-web shows a generic retry card; support queue +38 tickets. | Reading adapter logs for provider timeouts. None found. |
| 14:16 | 5xx 41%, the peak. | m.haas joins. Both engineers on the adapter. |
| 14:19 | 5xx 39%. Unchanged. | Adapter rolled back to v2.2. |
| 14:22 | 5xx 40%. | Rollback confirmed complete. Theory kept anyway. |
| 14:28 | 5xx 40%. | m.haas greps merges outside acme/checkout. Finds pool = 5. |
| 14:30 | Pool restored to 20; identity and checkout redeploy. | Fix applied. |
| 14:32 | 5xx 0.02%. Recovered. | Incident closed. |
Why the wrong theory held for 19 minutes
3 reasons
▲
The only change anyone could see was in the service that was failing
cost: 19 min
The deploy board a responder opens from the page is scoped to acme/checkout.
The change that caused the incident was merged in acme/identity and never
appeared on it. From 14:09 the search space excluded the answer.
Why it matters: the adapter deploy was 25 minutes old and completely innocent. It
was the best available suspect only because it was the sole visible one.
◆
The rollback disproved the theory at 14:19 and the theory survived it
cost: 9 min
Error rate was 39% before the adapter rollback and 39% after it. That is a clean
negative result at minute 13 of 26. The next nine minutes went to confirming the
rollback had really landed, then to re-reading the adapter logs.
Why it matters: this is the cheapest 9 minutes on the page. No new tooling was
needed to recover them, only the habit of dropping a theory the evidence has already
killed.
■
Pool saturation was at 100% the whole time, on a dashboard nobody opened
not visible
Connection pool utilisation is graphed in the identity dashboard, not the checkout one.
It went to 100% at 14:06:12 and stayed there for 26 minutes, four seconds before the
first 500.
The actual change
acme/identity · PR #2208
One line in a chart that acme/checkout imports and
acme/identity owns. Merged 13:52, titled
chore: align pool defaults with the sample config, reviewed by one person,
labelled config-only. It was correct for identity, whose pods are small and many.
Checkout runs 48 workers per pod.
charts/_shared/db/values.yaml
- DB_POOL_MAX: 20
+ DB_POOL_MAX: 5
0connections48
Workers wanting a connection at peak
48
Pool before the merge
20
Pool after the merge
5
Checkout survived from 14:04 to 14:06 on cached connections. The pool was exhausted
the moment the first sustained traffic arrived.
What would have cut the 19 minutes
3 actions
✓
Put every merge touching a shared chart on the checkout deploy board
−17 min
PR #2208 would have been the second entry on the board at 14:09, one line under the
adapter deploy, with the diff visible. Owner: platform. Two days of work.
✓
Graph pool saturation beside the 5xx rate, not in another service's dashboard
−12 min
A flat line at 100% next to the error rate names the failure mode without naming the
change. Owner: checkout. One afternoon.
✓
Treat a fix that does not move the metric as a disproof, on a 3-minute timer
−9 min
The adapter rollback landed at 14:19 and the error rate did not move. Make that a
checkpoint in the incident template: if the metric is flat three minutes after a fix,
the theory is dead and the search widens. Owner: incident process. Costs nothing.