A field manual for the day your earlier edit won
The Clock.
Every machine keeps its own time, and every one of them is lying a little — by milliseconds, sometimes by seconds, occasionally backward. You will watch an earlier edit erase a later one because its clock lied, build a clock that needs no time at all, and meet the database that solved it with atomic clocks and a promise to wait. This is the machinery under last-write-wins — and the reason “just order by timestamp” is a trap.
two clocks, one wire, and a disagreement nobody notices until it costs something.
BEGIN12:00:00 is an opinion.
Every computer keeps time with a vibrating quartz crystal — and every crystal vibrates at its own pace, drifting by seconds per day. Network Time Protocol papers over the drift, but it steers, never guarantees: it corrects in small nudges when it can and jumps the clock outright when it must — backward, sometimes, with no announcement. So 12:00:00 on machine A and 12:00:00 on machine B are two different claims about two different universes, made by witnesses who have never met.
Into this gap walks the most common sentence in system design: “just order the writes by timestamp.” It sounds like physics. It is last-write-wins, a common timestamp-based merge policy — and it works exactly as well as its clocks deserve. The Lag needed an ordering for its rungs; The Cache needed timestamps for its TTLs. This essay opens the box those essays only borrowed from: where does order come from, when no two clocks agree? Three exhibits: a crime, a fix, and an invoice.
The earlier edit wins.
A shared document. Mira edits it; two seconds later Jae improves her wording. Their edits sync to a server that resolves conflicts by LWW — keep the bigger timestamp. The trap is already set: Mira’s clock runs a few seconds fast. Set the skews below and RUN THE CRIME.
LWW doesn’t compare edits. It compares testimonies — and one of the witnesses is drunk.THE CRIME SCENE, SUMMARIZED
What makes this the nastiest class of bug in replication is the silence. Every component did its job: the clocks ticked, the sync shipped, the server kept the “latest” write. The timestamp was testimony from a witness with a different watch, and the verdict was rendered without cross-examination. NTP would eventually have corrected Mira’s clock — but the merge happens once, and the correction comes too late for the edit that no longer exists.
A clock that needs no time.
Lamport’s 1978 move is to abandon the question “what time is it?” and ask the only question that matters: “what caused what?” The Lamport clock is a counter with two rules — bump on every event; on receiving a message, take the max of yours and what arrived, then bump. The guarantee is one-directional: if A caused B, then A’s counter is smaller. Causality can never run backward again.
The toy below is live. Create events, pass messages, and click any two events — the oracle tells you whether one caused the other, or whether they are concurrent strangers that the counter ranks only because a total order demands it.
What you bought: an order that respects causality everywhere, on every machine, with no physics and no sync — the foundation every replication state machine, including Raft’s terms from the CAP essay, actually stands on. What you didn’t: concurrency detection (a small counter can rank concurrent events but can’t flag them — that needs vector clocks; see the field guide), and any connection to real deadlines. “Commit before Friday” is a physics question. For that, you pay.
Atomic clocks, and the humility to wait.
Spanner’s answer is neither to trust wall clocks nor to abandon them, but to bound them: GPS receivers and atomic clocks in every datacenter hold the uncertainty to milliseconds. The API stops pretending time is an instant — TrueTime returns an interval. Then the commit rule: pick τ past the far edge of every participant’s band, and wait until τ is certainly in the past before acknowledging — commit-wait. Drag the uncertainty, write at both machines, and commit.
Spanner’s trick is not better clocks. It is bounded clocks — plus the humility to wait out the bound.AFTER CORBETT ET AL., OSDI 2012
Read the skip-the-wait ghost carefully and you’ll recognize the criminal from FIG. 01 — it is the same lie wearing a suit: a stamp that outranks reality. The wait is what makes the stamp trustworthy, and its price is honest: about 2ε per commit, the PACELC invoice signed in milliseconds. You met this bargain twice before — Kafka’s scoped exactly-once (The Ack) is real only inside one system’s log, and Spanner’s external consistency is real only inside its bounded clocks. Guarantees have boundaries, and the boundaries are where the invoices are paid.
MODEL NOTES — times compressed and modeled: ε is your dial; Spanner’s real ε is ~1–7 ms · τ = the commit-time upper bound (modeled as true now + ε) · ACK waits until the current lower bound passes τ, about 2ε in this symmetric model.
The field guide.
- Wall clock
now()from the OS. Skews, steps backward, and — alone — can order nothing across machines. Fine for stamps humans read; treacherous for anything that decides.- Monotonic clock
- A clock that only moves forward: perfect for measuring durations on one machine, meaningless for comparing across machines. If you’re timing something, this is the one.
- Drift
- Quartz crystals disagree by seconds per day. NTP steers them back — in nudges, or in a backward jump that can scramble anything built on “timestamps always increase.”
- LWW
- Last-write-wins: keep the greatest timestamp, discard the rest. Sound only if the stamps are — and they never fully are. See FIG. 01 for the funeral.
- Lamport clock
- A counter with two rules. Guarantees causality is never inverted; costs nothing; needs no infrastructure. The load-bearing wall of every replication protocol.
- Happens-before
- The relation Lamport actually defined: a → b if same process and a came first, or a is a message’s send and b its receive, or transitively. The order that survives unfaithful wires.
- Vector clock
- A counter per process, compared componentwise. Unlike Lamport’s, it detects concurrency: neither before the other → report it. Riak shipped this as “siblings” — your shopping cart, honestly admitting the conflict instead of quietly eating it.
- TrueTime
- Time as a bounded interval, from GPS and atomic clocks. The honest API: not “now” but “now, plus-or-minus ε, and ε is small and known.”
- Commit-wait
- Wait until τ is surely in the past before acking. Buys external consistency; bills ~2ε. The reason Spanner can promise “any later transaction sees this one.”
- Hybrid logical clocks (HLC)
- Lamport’s rules grafted onto a wall-clock nanosecond field — causally sound and human-readable. What CockroachDB and YugabyteDB ship for exactly this reason.
WHY THE ONE-WAY STREET IS STILL ENOUGH
Lamport gives a→b ⟹ C(a)<C(b), but C(a)<C(b) proves nothing — concurrent events are ranked arbitrarily. For ordering a replication log, arbitrary-but-consistent is exactly what you want: every replica picks the same winner, and causality is never violated.
Detecting conflict is a different job. Vector clocks give each process its own counter and compare componentwise: V(a) ≤ V(b) with one strict entry means a→b; incomparable vectors mean concurrent — report it, ship siblings, let the application decide.
The ladder in one line: Lamport orders, vectors detect, TrueTime anchors to physics. Each rung buys the guarantee the previous one lacked — and prices it.
Three scenarios.
From these figures, you can separate wall time from causal order and explain why a bounded-clock commit must wait. Try changing one assumption and check whether your explanation still holds.
The sealed sheetThree questions are sealed inside this sheet. Nobody is asked to open it — the counters don't need your time.Break the seal
Mira edits at true 10:00:00 on a machine running +4 s fast; Jae edits at true 10:00:02 on an honest clock. An LWW store receives both. Which edit survives?