A field manual for the day one truth became two
The Lock.
Inside one database, COMMIT is a promise kept in nanoseconds. Split the booking across two services and the word transaction quietly stops spanning your system — and every fix for that buys something else. You will break the promise yourself, kill the coordinator in the one instant where it matters, and watch the world freeze while a number climbs. Then you will learn why the fix is a story you tell backward.
one decision, two databases, and the padlock in between.
BEGINOne word, doing its job.
Inside one database, a booking looks like this: BEGIN; mark seat 14C taken; insert the booking row; COMMIT. Two writes, one certainty — both happened, or neither did. That certainty is the A in ACID, and you have never once thought about it while writing the code, because it is simply true. It is true because both writes live inside one lock and one log that the database owns.
Then the company grows. A bookings service, a pricing service, separate databases, an event bus so the rest of the company can react. And the word transaction — without anyone announcing it — quietly stops spanning your system. This essay is about what replaces it, what it costs, and which parts of the bill nobody warns you about.
Two writes, one truth.
The booking flow every team writes first: save the row, publish the event. The code reads like one operation. It is two — separated by a window you did not write and cannot see. Below, that window is held open for a second and a half so you can find it. Book something, then kill the worker while it publishes.
What you just made is called the dual-write problem, and its deepest property is that both systems look completely healthy. The database is up. The bus is up. The worker restarted and carries on. Nothing crashes; something is simply, permanently wrong — the database says one thing about the world and every downstream subscriber believes another. No code review catches it, because the gap is not between two lines of code. It is between two systems.
Two-phase commit exists to make that window not exist — by refusing to let either write become true until both systems have promised, in writing, that they could make it true.
The machine that asks permission.
Phase one: the coordinator asks every participant to prepare: lock the row, write the intent to your own durable log, and answer yes only if you can guarantee both commit and abort later. Phase two: when every vote is yes, the coordinator writes one line to its own log — the decision — and only then tells anyone. Watch the padlocks: between the first lock and the last commit, every locked row is unusable to everyone else.
Now the experiment this whole page exists for. STEP through the protocol one event at a time — and KILL THE COORDINATOR somewhere along the way. There is exactly one fatal stretch. Find it.
WHY A PREPARED PARTICIPANT MAY NOT DECIDE ALONE
“Prepared” is a double promise: I can commit, and I can abort — you choose. After the coordinator vanishes, the participant does not know which was chosen, and no timeout can tell it.
If it guesses “abort” while the decision was “commit,” the cluster has made a split decision — half the world booked, half not. That is the one outcome worse than waiting, so waiting is mandatory.
Before prepare, the calculus is reversed: no promise has been made, so unilateral abort is safe. That asymmetry is why a kill before the first lock heals itself — and a kill after it never can.
MODEL NOTES — one locked row per participant · deterministic message order · no concurrent bookings (their math lives in FIG. 03) · participant timeouts compressed to seconds.
You pay even on a good day.
When nothing fails, 2PC is fast — flat in N, often faster than a saga. That is not the scandal. The invoice has lines that never appear on the happy path: every participant’s lock spans the whole commit; the blast radius of one failure is all N participants; and the odds that everyone can still say yes shrink geometrically. Drag the sliders and watch which costs grow — and which stubbornly don’t.
Notice what did not grow: 2PC’s duration. Its real costs are lock span and blast radius, not latency — a distinction interviews love and triangle diagrams hide. And one cost is absolute: the freeze you produced in FIG. 02 does not degrade gracefully. It is binary. Either the log comes back, or the world waits.
The long way back.
A saga replaces the protocol with a process: each step is an ordinary local transaction, and each step ships with a compensation — the pre-written answer to “how do we un-live this?” Run forward until something breaks; then run the compensations backward, last-first. Pick where it breaks below.
| DR/CR | event | customer |
|---|
Compensation is not undo. The flight was never yours; the fee was always real.THE SAGA’S FINE PRINT — GARCIA-MOLINA & SALEM, 1987
Run it with CARD selected: the charge lands, the confirmation never comes, and the refund arrives as a new line, not an erasure. While the saga ran, every intermediate state was public — the card really was charged while the car really was missing. What you buy: no distributed locks held across services; an orchestrated saga still has coordinator state, and a failure can trigger compensations in earlier services. What you pay: you write your own undo — and it must be idempotent, because with retries, every step will eventually run twice.
The field guide.
- Atomicity
- The A in ACID: all of it or none of it — guaranteed by one database’s lock and log, on one machine, in nanoseconds.
- Two-phase commit
- Prepare, then commit. One coordinator, N written promises, and one line of log that decides everything.
- In-doubt
- A participant that prepared and is waiting for a decision. It can neither commit nor abort. It can only wait.
- Saga
- A sequence of local transactions with a compensation for each — atomicity rebuilt as a process instead of a protocol.
- Idempotency
- Running the same step twice has the effect of once. Not optional: retries are a certainty, and every saga step eventually runs twice.
- Transactional outbox
- Write your row and the event in one local transaction; a relay publishes later, with retries. The database stays the single truth.
Three scenarios.
From these figures, you can find the blocking point in two-phase commit and account for compensation in a saga. Try changing one assumption and check whether your explanation still holds.
The sealed sheetThree questions are sealed inside this sheet. Nobody is asked to open it — the booking survives either way.Break the seal
In FIG. 02, the coordinator dies after both databases voted YES but before any COMMIT was delivered. What happens?