Gregario voice turn latency 10,000 turns

The median said ship it. One turn in a hundred is 3.3 seconds long.

Speech-out to first-audio-back, measured on device across 10,000 real voice turns. p50 is 950 ms and reads as a reply. p95 is 1,940 ms and is survivable. p99 is 3,310 ms, and at three seconds a rider assumes the thing did not hear them and says it again — which starts a second turn, on a bike, at 34 km/h. The mean for this same set is 1,081 ms. It sits 131 ms from the median and hides all of it.

  1. 200–400 ms: 14 turns
  2. 400–600 ms: 486 turns
  3. 600–800 ms: 2,180 turns
  4. 800–1,000 ms: 3,010 turns, the mode
  5. 1.0–1.2 s: 1,820 turns
  6. 1.2–1.4 s: 1,010 turns
  7. 1.4–1.6 s: 552 turns
  8. 1.6–1.8 s: 300 turns
  9. 1.8–2.0 s: 178 turns
  10. 2.0–2.2 s: 116 turns, over budget
  11. 2.2–2.4 s: 78 turns, over budget
  12. 2.4–2.6 s: 55 turns, over budget
  13. 2.6–2.8 s: 40 turns, over budget
  14. 2.8–3.0 s: 30 turns, over budget
  15. 3.0–3.2 s: 22 turns, over budget
  16. 3.2–3.4 s: 15 turns, over budget
  17. 3.4–3.6 s: 12 turns, over budget
  18. 3.6–3.8 s: 10 turns, over budget
  19. 3.8–4.0 s: 9 turns, over budget
  20. 4.0–4.2 s: 8 turns, over budget
  21. 4.2–4.4 s: 7 turns, over budget
  22. 4.4–4.6 s: 6 turns, over budget
  23. 4.6–4.8 s: 6 turns, over budget
  24. 4.8–5.0 s: 5 turns, over budget
  25. 5.0 s and over: 31 turns, worst 11,204 ms
0.2 s 2.7 s 5 s +
10,000 turns in 200 ms buckets, shortest on the left. The tallest bucket is 800–1,000 ms and holds 3,010 turns. Bar height is square-root scaled — on a linear count axis every bucket past 2 s is a single pixel and the finding disappears. The red buckets are the 450 turns past the 2 s budget: 4.5% of the ride, and the only part of this chart anybody complained about. Nothing lands under 200 ms because the endpoint decision alone costs 180.
● p50
950ms

Reads as a reply. This is the number that made us confident.

◆ p95
1,940ms

Just inside the 2 s budget. Riders wait, and they notice waiting.

▲ p99 — breaks it
3,310ms

Past this they repeat themselves and the turn is spent twice.

■ over 2 s
450turns

4.5% of all turns. Worst single turn: 11,204 ms.

What the tail actually is

100 slowest turns

Two independent causes, and they are not spread evenly — both cluster on the first turn after the rider has been quiet, which on a real ride means every climb and every descent. Ordered by how many of the 100 slowest turns each one explains.

Cold TTS voice after a silence gap

61 of 100 · fix now

The synthesis provider parks a voice model after 90 s idle in eu-west-1. A rider climbing for eight minutes says nothing, and the first thing they say afterwards pays a cold load. Median TTS first-chunk is 190 ms; on these turns it is 870 ms.

Why it hid: our staging loop talks every few seconds, so the voice was never cold and this cause has a measured rate of zero in pre-production. It only exists on rides.

LLM 429, retried once, full re-request

31 of 100 · next

The Hono handler retries a rate-limited completion after a 400 ms backoff and re-sends the whole prompt rather than resuming the stream. Median LLM first-token is 340 ms; on a retried turn it is 1,980 ms — the backoff plus a second cold prompt.

Why it hid: the retry succeeds, so the turn is logged 200 OK and never appears in an error rate.

Everything else

8 of 100 · watch

Cellular handover mid-turn, mostly. No pattern worth chasing until the first two are gone.

Where the time goes

same scale, 0–2,000 ms

Supporting detail, not the finding. The point of putting these side by side is that the slow turn is not uniformly slow: three of the five stages barely move, and two carry 2,320 ms of the 2,360 ms of difference.

● Median turn · 950 ms total

VAD endpoint 180 ms
Speech to text 210 ms
LLM first token 340 ms
TTS first chunk 190 ms
Playback start 30 ms

▲ p99 turn · 3,310 ms total

VAD endpoint 180 ms
Speech to text 240 ms
LLM first token ▲ 1,980 ms
TTS first chunk ▲ 870 ms
Playback start 40 ms
All 25 buckets, as numbers
BucketTurnsShareCumulative
200–400 ms140.14%0.14%
400–600 ms4864.86%5.00%
600–800 ms2,18021.80%26.80%
800–1,000 ms3,01030.10%56.90%
1.0–1.2 s1,82018.20%75.10%
1.2–1.4 s1,01010.10%85.20%
1.4–1.6 s5525.52%90.72%
1.6–1.8 s3003.00%93.72%
1.8–2.0 s1781.78%95.50%
2.0–2.2 s ▲1161.16%96.66%
2.2–2.4 s ▲780.78%97.44%
2.4–2.6 s ▲550.55%97.99%
2.6–2.8 s ▲400.40%98.39%
2.8–3.0 s ▲300.30%98.69%
3.0–3.2 s ▲220.22%98.91%
3.2–3.4 s ▲150.15%99.06%
3.4–3.6 s ▲120.12%99.18%
3.6–3.8 s ▲100.10%99.28%
3.8–4.0 s ▲90.09%99.37%
4.0–4.2 s ▲80.08%99.45%
4.2–4.4 s ▲70.07%99.52%
4.4–4.6 s ▲60.06%99.58%
4.6–4.8 s ▲60.06%99.64%
4.8–5.0 s ▲50.05%99.69%
5.0 s and over ▲310.31%100.00%

What we are changing

None of this touches the median, and that is the point — the median was never the problem and any work aimed at it would have moved a number nobody feels. Keep a synthesis warm for the duration of a ride rather than for 90 s of chatter. On a 429, retry with backoff against a warm prompt cache or a fallback model, so the retry does not pay the full setup cost again. Then re-run this same chart: the shape to look for is the red smear shortening, not the hump moving.