Gregario voice turn latency 10,000 turns
Speech-out to first-audio-back, measured on device across 10,000 real voice turns. p50 is 950 ms and reads as a reply. p95 is 1,940 ms and is survivable. p99 is 3,310 ms, and at three seconds a rider assumes the thing did not hear them and says it again — which starts a second turn, on a bike, at 34 km/h. The mean for this same set is 1,081 ms. It sits 131 ms from the median and hides all of it.
Reads as a reply. This is the number that made us confident.
Just inside the 2 s budget. Riders wait, and they notice waiting.
Past this they repeat themselves and the turn is spent twice.
4.5% of all turns. Worst single turn: 11,204 ms.
Two independent causes, and they are not spread evenly — both cluster on the first turn after the rider has been quiet, which on a real ride means every climb and every descent. Ordered by how many of the 100 slowest turns each one explains.
The synthesis provider parks a voice model after 90 s idle in
eu-west-1. A rider climbing for eight minutes says nothing,
and the first thing they say afterwards pays a cold load. Median TTS
first-chunk is 190 ms; on these turns it is 870 ms.
Why it hid: our staging loop talks every few seconds, so the voice was never cold and this cause has a measured rate of zero in pre-production. It only exists on rides.
The Hono handler retries a rate-limited completion after a 400 ms backoff and re-sends the whole prompt rather than resuming the stream. Median LLM first-token is 340 ms; on a retried turn it is 1,980 ms — the backoff plus a second cold prompt.
Why it hid: the retry succeeds, so the turn is logged 200 OK and never appears in an error rate.
Cellular handover mid-turn, mostly. No pattern worth chasing until the first two are gone.
Supporting detail, not the finding. The point of putting these side by side is that the slow turn is not uniformly slow: three of the five stages barely move, and two carry 2,320 ms of the 2,360 ms of difference.
● Median turn · 950 ms total
▲ p99 turn · 3,310 ms total
| Bucket | Turns | Share | Cumulative |
|---|---|---|---|
| 200–400 ms | 14 | 0.14% | 0.14% |
| 400–600 ms | 486 | 4.86% | 5.00% |
| 600–800 ms | 2,180 | 21.80% | 26.80% |
| 800–1,000 ms | 3,010 | 30.10% | 56.90% |
| 1.0–1.2 s | 1,820 | 18.20% | 75.10% |
| 1.2–1.4 s | 1,010 | 10.10% | 85.20% |
| 1.4–1.6 s | 552 | 5.52% | 90.72% |
| 1.6–1.8 s | 300 | 3.00% | 93.72% |
| 1.8–2.0 s | 178 | 1.78% | 95.50% |
| 2.0–2.2 s ▲ | 116 | 1.16% | 96.66% |
| 2.2–2.4 s ▲ | 78 | 0.78% | 97.44% |
| 2.4–2.6 s ▲ | 55 | 0.55% | 97.99% |
| 2.6–2.8 s ▲ | 40 | 0.40% | 98.39% |
| 2.8–3.0 s ▲ | 30 | 0.30% | 98.69% |
| 3.0–3.2 s ▲ | 22 | 0.22% | 98.91% |
| 3.2–3.4 s ▲ | 15 | 0.15% | 99.06% |
| 3.4–3.6 s ▲ | 12 | 0.12% | 99.18% |
| 3.6–3.8 s ▲ | 10 | 0.10% | 99.28% |
| 3.8–4.0 s ▲ | 9 | 0.09% | 99.37% |
| 4.0–4.2 s ▲ | 8 | 0.08% | 99.45% |
| 4.2–4.4 s ▲ | 7 | 0.07% | 99.52% |
| 4.4–4.6 s ▲ | 6 | 0.06% | 99.58% |
| 4.6–4.8 s ▲ | 6 | 0.06% | 99.64% |
| 4.8–5.0 s ▲ | 5 | 0.05% | 99.69% |
| 5.0 s and over ▲ | 31 | 0.31% | 100.00% |
None of this touches the median, and that is the point — the median was never the problem and any work aimed at it would have moved a number nobody feels. Keep a synthesis warm for the duration of a ride rather than for 90 s of chatter. On a 429, retry with backoff against a warm prompt cache or a fallback model, so the retry does not pay the full setup cost again. Then re-run this same chart: the shape to look for is the red smear shortening, not the hump moving.