gregario local-first sync ride 2026-07-29
One ride, 46 minutes without signal, 47 rows replayed from the outbox on reconnect. Forty-two applied cleanly, five had been edited on both sides, and last-write-wins picked correctly four times out of five. The fifth compared a device clock running 41 seconds slow against a server clock, decided a dictated rider note was the older write, and replaced it with a machine-generated summary. No error, no tombstone, no retry.
14:02:11 to 14:49:03. Rider never noticed.
Across 5 tables, one push, 214 KB.
No server-side edit since the row left.
Edited on the phone and on the server.
coach_notes#4471. Human write, lost to a cron job.
Phone behind server. Never corrected on the client.
outbox. Nothing on the ride screen blocks on the network.
updated_at from the device clock at write time.
POST /sync/push, 47 rows, 214 KB, 312 ms round trip, HTTP 200.
updated_at per row, last-write-wins
Drizzle upsert with a WHERE excluded.updated_at > target.updated_at guard. The comparison is between two clocks that were never reconciled.
notes.summarize rewrote body from the last 20 minutes of telemetry. Machine text, regenerable, runs again every ride.
Ordered by volume. Waypoints dominate because the client samples GPS every three seconds while recording; they are also the least contested, since only one writer produces them.
Ordered by how close the two writes were in real time, widest gap first. Both timestamps are as stored: the phone column is device-clock, the server column is server-clock, and nothing in the push reconciles them.
| Row | Phone updated_at |
Server updated_at |
True gap | LWW kept | Sound? |
|---|---|---|---|---|---|
settings#12 |
14:05:20 | 14:41:36 | 35m 35s | server | ✓ yes |
laps#8819 |
14:44:02 | 14:20:55 | 23m 48s | phone | ✓ yes |
laps#8812 |
14:31:07 | 14:22:19 | 9m 29s | phone | ✓ yes |
waypoints#33190 |
14:26:44 | 14:31:02 | 3m 37s | server | ✓ yes |
coach_notes#4471 |
14:39:04 | 14:39:33 | 12s | server | ✕ inverted |
True gap is the real-time distance between the two writes after adding the 41 seconds back onto the device clock. It is not stored anywhere; it was reconstructed from the server's receipt log.
Bars are on a log scale, because the gaps run from 12 seconds to 36 minutes and a linear axis would render the interesting one as a hairline. The dashed bar is not data: it is the 41-second skew window. Any fork whose two writes fall closer together than that bar can be decided the wrong way round. Exactly one did.
Cadence held 82–86 rpm through the climb.
Normalised power 241 W over 18m 40s.
Consider a longer warm-up before the next
threshold block.
Regenerated from telemetry on every ride close. Reproducible: re-running the job restores it byte for byte.
left knee started clicking near the top of
the mortirolo — cut the interval block, keep
it z2 for the rest of the week
Voice capture, transcribed on device, never left SQLite in any other form. Unrecoverable: the outbox row was acked and deleted at 14:49:03.
For four of five forks, yes, and the architecture deserves the credit: offline capture worked for 47 minutes on a mountain with no bars, the replay was a single request, and reinstalling no longer costs the rider their history. The failure is not that LWW is wrong. It is that LWW is being asked to order two events using two clocks that have never spoken to each other, and the only table where a machine writes while the client is offline is also the only table where the loser is a human sentence nobody else can reconstruct.
The exposure is bounded and countable: a fork can only invert if both
writes land within the skew window, and only coach_notes
has a server-side writer that fires during a ride. On this ride that is
1 row in 47. On a ride with no dictated notes it is 0.
coach_notes append-only
A rider's dictated note and a generated summary are not two versions
of one row; they are two rows with a source column. No
comparison, no winner, nothing to lose. Costs one migration and a
list view that already exists.
The push already knows when it started. Sending the device's own
Date.now() alongside it lets the server compute the skew
per batch and normalise every updated_at in it before
comparing. Fixes the ordering without changing the resolution rule.
An HLC per row makes causality explicit and stops a slow device clock deciding the winner. It cannot rank two genuinely concurrent edits, so conflicts like this one have to be kept rather than resolved. It is the right long answer and the wrong first move: it touches all five tables to fix a defect currently confined to one.