A field manual for the day three nines became 99.7
The Nines.
Every box in your architecture has an availability number, and every number looks fine. The user does not meet your boxes — they meet the pipeline, and pipelines do not average. They multiply. You will add a component and watch reliability fall, buy a redundant twin and watch shared fate eat it, then spend a real error budget until feature launches freeze. This is the arithmetic under every SLA you have ever signed.
four healthy boxes. one number nobody put on the slide.
BEGIN99.9% is 8.8 hours a year.
Nobody feels a nine. Convert it and it gets heavy: 99.9% is 8.8 hours of downtime a year — 44 minutes every month, ten minutes every week. 99.99% is 52 minutes a year. 99.999% is five. Each extra nine is not one percent better; it is ten times better, which is why they are bought with architecture rather than effort, and why the fifth nine costs more than the first four together.
The trap is structural, and this page exists to make your hands feel it: the promise is written per box — “our database is 99.9%” — but the user meets the whole pipeline. A request crosses the load balancer, the web tier, the cache, the database. For the user’s request to succeed, every one of them must be up at that instant. All must hold: the probabilities multiply. Four boxes at 99.9% do not give you 99.9%. They give you 99.6. And every additional box on the critical path — every microservice, every sidecar, every “just one more hop” — multiplies the number down.
Availability is not a property of your boxes. It is a property of their conjunction.THE SENTENCE THIS ENTIRE ESSAY DRAWS
Add a box, lose a nine.
Below is a pipeline: four boxes, each individually respectable. The big number is the only one users ever meet — the product. Drag any dial. Then press ADD A COMPONENT, because growth means adding boxes, and you should watch what growth does to the promise.
Two lessons live in that figure. First, the backfire: every ADD drops the number. The new box is individually fine — 99.9%, perfectly shippable — and the pipeline gets worse, because a series request now needs one more yes. Capacity was added; availability was subtracted. Nobody notices in the planning meeting because each box’s SLA is green. Second, the spending rule: the amber-outlined box is your worst link, and it is where every upgrade dollar belongs. Improve your best box from 99.99 to 99.999 and the chain barely shrugs. Bring the worst from 99.9 to 99.99 and the whole product improves. The multiplication makes reliability a weakest-link economy: the chain’s downtime is dominated by its least available member, and every other dial drowns in it.
MODEL NOTES — boxes assumed independent (a fiction FIG. 02 corrects) · promised line fixed at 99.9% · conversions from the exact product: year = 525,600 min, month = 43,800.
Redundancy — and what eats it.
The escape from multiplication is addition: run two copies and the user needs only one to be alive. Two 99.9% boxes in parallel are 99.9999% — three extra nines from one doubling. That trade is the entire return on replication, and the whole reason the series exists. But the parallel formula carries a fine print so dangerous it deserves its own dial: shared fate. The math assumes your twins fail like strangers. Twins that share a machine, a rack, an availability zone, or a config deploy are not strangers — they are the same witness twice, and their failure is one event, not a coincidence of two.
Drag SHARED FATE and watch the twin’s gift evaporate: by the time the boxes share an availability zone, your “99.9999%” pair delivers barely one better nine; on the same machine, it delivers nothing — a backup of your own outages. This is the senior-signal question in every design review: not “is it redundant?” but “what does this redundancy share?” And it is why the real availability of a system is set less by its boxes than by its correlations.
Worse: the multiplication never asked your diagram’s opinion. The boxes that quietly sit in every chain — DNS, the identity service, the load balancer itself, service discovery, the deploy pipeline — multiply too, and none of them appears on the slide. The unlabeled box is still in the math. It is usually the weakest one, per the spending rule.
THE ARITHMETIC OF TRUST — SERIES, PARALLEL, AND THE LOAD-BEARING ASSUMPTION
Series: a request succeeds only if every component is up, so A = a₁ × a₂ × … × aₙ. Products of numbers just below one shrink — and every added factor, however close to one it looks, pulls the whole product toward zero. That is the chain, and the reason “each box is fine” is not an argument.
Parallel: a request succeeds if at least one copy is up, so A = 1 − (1−a)². Two 99.9% boxes: 1 − 0.001² = 99.9999%. Redundancy multiplies the independent failure probabilities, then takes the complement — and that independence is the expensive part.
The assumption: both formulas assume independence — that no common cause can reach two boxes at once. The world is generous with common causes: one config pipeline, one availability zone, one schema migration, one human. With correlation ρ, joint failure drifts from p² toward p — and the twin’s promised nines evaporate. Redundancy without independence is not engineering. It is accounting.
Downtime is currency. Spend it on purpose.
Here is the reframe that separates operators from aspirants: your delivered downtime is not a scandal — it is a error budget, pre-funded by your SLO. At a 99.9% SLO you hold 43.8 minutes of allowed failure in an average month — and you are supposed to spend some of it, on risky deploys and experiments, because a team that never spends its budget is a team that never ships. The rule is not “never fail.” The rule is: when the budget is gone, features stop and reliability work starts. That contract, invented by Google’s SREs, is what keeps “move fast” and “don’t melt” from being rival religions. Below: your pipeline’s month. Spend it.
Feel the exchange rates the figure teaches: at a 99.9% SLO target, one 35-minute incident spends about 80% of the monthly budget. At a 99.75% target, that same 35-minute incident spends about a third of the budget. The same outage is a scandal or a rounding error depending entirely on the promise you wrote down — which means the SLO is a resource-allocation decision, not a vanity number. And when the budget hits zero, notice what the freeze actually says: not “stop failing” — you can’t command that — but “stop spending until you’ve refilled.” There are only two refills: fail less often, or recover faster.
The second refill is the underrated one. 8.8 hours a year can be one outage in March or twenty-six twenty-minute blips and eight more minutes — identical nines, unrecognizably different companies. At a fixed failure frequency, halving mean time to recovery halves downtime; availability rises, but does not double. Recovery work is often cheaper than preventing every failure: canaries, rollbacks, and the discipline to page early are nines you can buy this sprint.
MODEL NOTES — month = an average 365-day year / 12 = 43,800 min · budget = allowance from the SLO target at the slider · spends modeled: deploy 12, canary 2, incident 35 · freeze at 100%.
The field guide.
- SLI / SLO / SLA
- The indicator (measured, e.g. “99.85% of requests OK”), the objective (the internal target), the agreement (the contract with penalties). An internal availability SLO is set above the contractual SLA to leave a safety buffer — the gap is your working room.
- Error budget
- 1 − SLO, per window. The monthly allowance of failure that buys deploys, experiments, and honest postmortems. Emptied → freeze features, fix reliability.
- Nine
- Each is 10× the last. 99.9% = 8.8 h/yr; 99.99% = 52 min; 99.999% = 5 min. Buy them with architecture, not heroics — the fifth nine costs more than the first four.
- MTTR / MTBF
- Mean time to recover / between failures. The two levers of every nine: fail less, or heal faster. Halving MTTR equals doubling reliability — and is usually the cheaper purchase.
- Single point of failure
- Any box whose death the product cannot survive. The dangerous ones are unlabeled: DNS, auth, the load balancer, service discovery, the deploy pipeline. Not on the diagram — fully in the math.
- Correlated failure
- One cause, many victims: same AZ, same config push, same migration. The assumption that quietly powers the parallel formula — and the first thing real outages delete.
- Fault domain / blast radius
- The set of things one failure can take down. Cells and shards exist for scale, but equally for this: cap the blast radius so the chain multiplies in pieces, not wholesale.
- Graceful degradation
- When the database dies, serve cached pages and disable checkout. A degraded experience reads the same on the dashboard and feels entirely different to users — nines hiding in plain sight.
- Failover
- Switching to the spare. The least-exercised code path in any system — an untested failover is a hypothesis. Chaos days exist to turn hypotheses into facts.
Three scenarios.
From these figures, you can compute an error budget and explain which shared dependency dominates availability. Try changing one assumption and check whether your explanation still holds.
The sealed sheetThree questions are sealed inside this sheet. Nobody is asked to open it — the budget can spare this.Break the seal
A request path crosses three services, each promising 99.9%. The team's design doc claims end-to-end availability of 99.9% “since each service meets its SLA.” What does the user actually get?