A field manual for the day the cache answered
The Cache.
“There are only two hard things in Computer Science: cache invalidation and naming things” — and, the joke adds, off-by-one errors. A cache is a bet that the future resembles the past, and every part of this page is about collecting on it. You will watch the bet pay, discover what a TTL actually buys, compound a page into the ground, and then melt an origin with one expired key — before fixing it three different ways.
most questions come back. the red one went all the way to the truth.
BEGINA copy of the past, with a lease.
Here is the whole technology in one sentence: keep an answer close to the questioner, and hope the question comes back. A cache is not a copy of the database and not a smaller database — it is a bet on locality, placed with other people’s patience. The pattern is called cache-aside, and it is three lines of code that every system on earth has written.
Below, the bet is running: every client wants one thing — the price of seat 14C. Green reads are answered from the cache in ~2 ms. Red reads miss, walk to the origin, pay ~45 ms, and come back with a copy. Watch the two meters: hit rate and average read. Then press WRITE TO ORIGIN — the world changes price, and the cache keeps telling the old one.
Now the arithmetic the meters are quietly doing. A read through the cache costs T = Tc + (1−H)Tm — every request pays the cache hop, and misses also pay the origin. To beat the origin alone, H > Tc/Tm. With a 2 ms cache and a 50 ms origin, that means more than 4% hits; against a 5 ms origin it needs more than 40%. A cache is not automatically faster. Below break-even it is a tax you added in front of your database.
And notice what the cache never does: refuse. Every read gets an answer, fast or slow. The danger was never speed.
The question was never “how fresh.”
You pressed the surge button and the lamp went amber: the world said one price, your users saw another — for up to your entire TTL. Here is the reframe that separates people who operate caches from people who configure them: a TTL is not a freshness guarantee. It is a staleness budget. The question is never “how do we keep the cache fresh” — you can’t, not exactly, not distributed. The question is: how wrong are we allowed to be, and for how long? That is an SLA on age, and you are allowed to choose it. Price stability for six seconds is a feature. Six hours of stale seat inventory is The Lock’s dual-write lamp, wearing a wig.
You can drive the age to zero — evict on every write. But now the writer must know about every reader’s cache, and the invalidation itself must be reliable. You have already built that machinery once: it was called the outbox. Which is why the quote survives:
There are only two hard things in Computer Science: cache invalidation and naming things.ATTRIBUTED TO PHIL KARLTON — AND THE OFF-BY-ONE IS IN THIS PAGE SOMEWHERE
Your hit rate is a lie of aggregation.
A page is not one read. A dashboard is twenty widgets; a product page is thirty pieces of data; each cached independently, each with its own healthy hit rate. The page is fully cached only if every piece hits: HN. Twenty widgets at 95% each — a config any team would celebrate — gives the page 36%. Most page loads touch the origin for something. Nobody’s dashboard shows this number, because every individual metric is green.
This is why page-level caching, fragment caching, and CDNs exist — not to make reads faster, but to collapse N independent bets into one. And it is why “our cache has a 95% hit rate” is, in an interview, a sentence that should be finished with “…per key, so the pages our users actually load see…”.
One key, five thousand questioners, zero seconds.
Locality means popularity concentrates. A handful of keys take most of the traffic — they sit warm in the cache, and everything above was fine. Until the lease lapses. The moment your hottest key expires, every client in the same second becomes a miss, and every miss queues at an origin that was serving zero of them a moment before. One missing key is one query. Ten thousand simultaneous misses of the same key is a melting building.
The sim below runs forever: the hot key expires every 8 seconds. First watch one full cycle untouched. Notice the shape — the stampede does not end when the cache re-warms. It ends when the queue drains, long after, because the herd was already in flight. Then try the fixes.
THE STAMPEDE FAMILY — FOUR WAYS TO SPREAD THE THUNDER
Request coalescing (single-flight). The first miss fetches; the rest subscribe to its result. Five thousand origin reads become one. Memcache at Facebook ships “leases” — a single-flight with teeth: the loser of a race is told to wait, not to fetch.
Stale-while-revalidate. Serve the expired copy instantly while one refresh runs in the background. You traded one second of age for one hundred percent of the wait.
TTL jitter. Expiry at TTL + rand(0…600 s), per key — so ten thousand keys never strike twelve together. The synchronized clock is the herd’s weapon; desynchronize it.
Probabilistic early refresh (XFetch). As expiry approaches, an exponential probability based on remaining lifetime and recent fetch duration sometimes starts a refresh early. Expensive refreshes are more likely to begin early; cheap ones can wait longer. The stampede becomes a gentle drizzle.
MODEL NOTES — one hot key; origin serves 40 reads/s at ~1 s each; clients time out at 2 s and retry once (jitter optional); dots are sampled (~8 requests each); fluid queue, times compressed and modeled.
Every miss is someone else’s hit.
So far, one cache. Real systems are stratified — five or six copies of the same bet stacked between a user’s finger and your disk. Each layer has its own hit rate, its own staleness budget, and its own idea of what a miss costs. Click through them: the useful question was never “should we cache?” It is where should the miss land, and who is allowed to be how stale there.
Read the stack bottom-up and something odd appears: the database was a cache all along. A “query” that touches RAM instead of disk is a cache hit; the buffer pool is just the oldest, most successful cache you never had to configure. Caching isn’t a layer you add. It’s a property of every layer you have.
The field guide.
- Cache-aside
- The app checks the cache, falls through to the origin on a miss, and fills the cache itself. The default pattern — and the one that stampedes, because nothing coordinates the misses.
- Read-through
- On a miss, the cache loads from the origin before answering. This describes how reads are filled; it says nothing about write consistency.
- Write-through
- Writes pass through the cache to the origin before acknowledgment. The extra write latency narrows the stale window, but failure handling still matters.
- Write-behind
- Writes land in cache and flush later. You bought latency and sold durability: a crash can lose writes the user saw confirmed.
- TTL
- A staleness budget, not a freshness guarantee. Choosing it means answering: how wrong may we be, for how long?
- Single-flight
- One request per key refreshes; the rest wait on its result. The stampede’s cure, and why the herd in FIG. 03 is optional.
- Stale-while-revalidate
- Serve the expired copy, refresh in the background. Age in exchange for a queue that never forms.
- LRU / LFU
- Eviction policy: least-recently or least-frequently used. Locality’s cleanup crew — it decides which bets the cache keeps room for.
- Hot key
- One key taking most of the traffic. Hashing spread keys across servers; it cannot spread one key — you watched that in The Ring. Replicate it or local-cache it.
Three scenarios.
From these figures, you can estimate when a cache helps and name the race that an asynchronous eviction leaves open. Try changing one assumption and check whether your explanation still holds.
The sealed sheetThree questions are sealed inside this sheet. Nobody is asked to open it — the origin never notices.Break the seal
A dashboard assembles 20 widgets, each individually cached with a 95% hit rate. Roughly how often does at least one widget miss and hit the origin?