A field manual for the day the network splits

Two out
of three.

In 2000, Eric Brewer proposed that a distributed database must trade away one of three promises: consistency, availability, or the ability to survive a network partition. Two years later the guess became a theorem. This page makes you feel it — you will cut the cables yourself.

three replicas, gossiping peacefully. for now.

BEGIN
№ 01THE HAPPY PATH

A system that works.

Hold something small in your head: a ticket shop with five tickets left, and three datacenters — Frankfurt, Virginia, Singapore — that must agree on how many remain. Every replica keeps a full copy of the count. Every customer asks the nearest one.

Buy a ticket below and watch it travel: the sale is accepted at one site, then copied to the others. Every replica shows the same number. Every customer gets an answer.

FIG. 01 — ONE STOCK, THREE REPLICAS

With the cables intact, updates can propagate and replicas eventually agree. The blinking lamp marks an update still in flight; this figure has not promised linearizable reads during that interval. Every node is answering requests here. The stronger consistency promise, and what it costs under a split, comes next.

Remember this moment. It is the part of the “two out of three” posters that everyone forgets: on a good day, you keep everything. The theorem has nothing to say about good days.

№ 02THE DIAGRAM LIES

The lines are not always there.

That tidy figure makes one quiet assumption: that the cables hold. They don’t. Backhoes find fiber-optic cables with remarkable enthusiasm. Switches install firmware at 3 a.m. and never come back. Google engineers have reported sharks biting undersea fiber. And even with every cable intact, a network can delay a packet so long that slow and dead become indistinguishable — in an asynchronous network, no timeout can tell them apart.

In a network where messages can be lost or delayed forever, no system can promise both linearizable consistency and perfect availability.THE CAP THEOREM — AFTER GILBERT & LYNCH, SIGACT NEWS, 2002
WHY IT’S A THEOREM — GILBERT & LYNCH IN FOUR SENTENCES

Consider two executions that look identical to Singapore during an indefinite partition. In one, Frankfurt never sold the last ticket; in the other, it sold that ticket just before the split. No message from Frankfurt reaches Singapore, so Singapore cannot tell which history is real.

If Singapore must answer a sale request in both executions, it must give the same answer in both. Selling violates the one-ticket rule in the second; refusing a valid sale violates the ticket-shop specification in the first. Waiting for information forever preserves consistency but sacrifices availability at Singapore.

This argument needs a partition that can delay messages indefinitely. One dropped message may be recovered by retrying. If communication eventually resumes, the system can recover; it still cannot guarantee both promises throughout the partition.

This is why partition tolerance is not a feature you build. It is the world you live in. You don’t get to declare it away:

FIG. 02 — THE WEATHER, ARRIVINGONE (1) NETWORK, ALLEGEDLY RELIABLE

A sale is mid-flight between Frankfurt and Singapore. Nothing could possibly interrupt it.

Since we are about to trade these away, let us be precise about what they are:

Consistency
Linearizability: every read reflects the most recent write — as if there were exactly one copy, touched in one agreed order.
Availability
Every request to a living node returns a real, non-error answer. No refusals, no hanging timeouts.
Partition tolerance
The system survives messages being lost or delayed without limit between its nodes. On any real network, this one is not negotiable.
№ 03CUT IT YOURSELF

Now it’s your network.

Click a cable to cut it — and learn the first real lesson: with three sites and three links, one cut changes almost nothing. Frankfurt still reaches Singapore through Virginia. Redundancy works; nobody even notices. To truly split a system, you must isolate it — cut a second line.

When you do, the theorem stops being a diagram and becomes a decision made under pressure about real customers. A sale arrives on the island side, which holds no quorum. Refuse it — or serve it and let the histories drift? Flip the policy switch and run the same partition both ways.

FIG. 03 — THE SANDBOXCLICK A CABLE TO SEVER IT · CLICK AGAIN TO REPAIR

MODEL NOTES — each replica displays a local count. AP merges per-origin sales after a heal and can oversell. CP synchronizes the majority’s check-and-decrement before acknowledging a sale; the packet animation is a teaching aid, not a real consensus protocol. Minority writes are refused, and this toy does not model strongly consistent reads.

№ 04THE HEALING

What “eventually” actually does.

Eventual consistency sounds gentle — as if the data were merely late, like a train. Watch what really happens when you repair the cable: the islands — a state engineers call split-brain — compare their histories and merge them, sale by sale, counting every confirmed order. If both sides sold the same scarce ticket while apart, the merge does not smooth it over. It reveals it: the shop is oversold, and some specific customer will meet the CAP theorem in person — at the door, without a ticket.

A CP system takes the opposite bargain. While split, it refuses every sale it cannot vouch for. Nobody is oversold; somebody is simply told no. Neither choice is kind. The theorem doesn’t prefer either — it only removes the third option, the one where every customer goes home happy.

Without a partition, CAP itself does not force a choice. Replication still has latency and failure costs, and a split makes the consistency and availability conflict unavoidable.THE SHAPE OF BREWER’S RETROSPECTIVE, IEEE COMPUTER, 2012
№ 05THE HIDDEN PRICE

You pay even on a good day.

Here is what the triangle hides: even with every cable intact, consistency and availability are not simultaneously free. If every replica must agree the instant a write is acknowledged, the write must first reach them — at the speed of light, through a trench, across an ocean. If you don’t wait, you answer instantly from a local copy that is, briefly, wrong.

Daniel Abadi named this PACELC: if Partition, choose Availability or Consistency; Else, choose Latency or Consistency. Feel the cost:

FIG. 04 — PACELC, THE HIDDEN INVOICELINK ~145 ms EACH WAY · TIMES MODELED
IF PARTITIONAkeep answering — histories may driftORCrefuse rather than lieELSE · HEALTHYLanswer in ~2 ms — maybe staleORCwait ~290 ms — always current
No operations yet — buy a ticket and watch the invoice.
0145290320 ms

Spanner and ZooKeeper live on the left of that bar — they wait. Cassandra and DynamoDB live on the right — they answer. Nobody lives at the corner where a write is both instant and already everywhere; physics has opinions.

№ 06A FIELD GUIDE

Where real systems sit.

Roughly — and each of these will happily argue with you about its dot at great length. What matters is the shape of the space. Notice that the far corner, the one marked CA, is not actually in the room.

FIG. 05 — THE SPECTRUMCLICK A DOT · POSITIONS APPROXIMATE
Select a system to read how it behaves when the network splits.

The only way to hold consistency and availability unconditionally is to hold no network at all — one machine, no replicas. The moment a second copy exists across a wire, partitions enter the story, and CAP starts writing your error messages for you.

№ 07WANNA TRY OUT WHAT YOU HAVE LEARNED?

Three scenarios.

From these figures, you can trace a partitioned sale and distinguish availability from a single coherent stock count. Try changing one assumption and check whether your explanation still holds.

The sealed sheetThree questions are sealed inside this sheet. Nobody is asked to open it — the machines already ran for you.Break the seal
question 1 of 3

Your three-node CP cluster loses two of its three links — one node is now an island. A client sends that island a write. What happens?