RATE LIMITINGDOWNTIME PRESSAN INTERACTIVE ESSAY

A field manual for the day the door held

The Bucket.

Every catastrophe in this series ended at some door: the retry storm arrived somewhere, the stampede queued somewhere, the ten-million-write fan-out knocked somewhere. This essay is about that somewhere — the one component whose entire job is refusing, on schedule, with a number attached. You will exploit a window’s blind spot, spend a bucket burst by burst, and watch a fleet of honest servers multiply your promise eightfold.

two get through. one gets a number and an apology. that is the whole product.

BEGIN
№ 01THE DOOR

The sentence to dismantle.

“We’ll add rate limiting if traffic becomes a problem.” It sounds like pragmatism. It is a category error: the limiter is not an optimization you bolt on when traffic gets bad — it is the precondition for “bad” being survivable. Every storm this series staged met a door. The Ack showed how clients retry when a reply vanishes; without limits, those loops pound this door. The Cache’s stampede is what a door looks like when it opens for everyone at once. The Feed’s queues bloom where no gate sheds load. The door is not infrastructure; it is the immune system.

And every limiter is an answer to a single question: what does “too much” mean, and over what span of time? The algorithms are just different calendars. The fixed window counts per page of the calendar — simple, and blind at the page turn. The token bucket hands out a budget that refills continuously — a loan you spend. The leaky bucket is a metronome: outflow exactly, always, whatever arrives. Choose the calendar, and you have chosen which bursts are forgiven and which are crimes.

The kindest thing an overloaded system can say is a fast, honest no — with a number attached.THE WHOLE JOB DESCRIPTION

One warning before the figures: everything here is the server side. The client side of this contract — back off, with jitter, when the no arrives — was The Ack’s essay. The two halves only work together; a limiter without polite clients is a filter for self-inflicted DDoS.

№ 02THE WINDOW

The calendar with a blind spot.

The naive limiter counts requests per fixed window — “100 per minute” — and resets the counter when the minute rolls over. Watch the counter fill, watch the 429s begin, and then press BURST AT THE BOUNDARY to commit the oldest crime in the book: half the burst lands on the dying window’s unused budget, half on the newborn’s fresh one. Each window is legally innocent. The client got twice the promise in a blink.

FIG. 01 — THE FIXED WINDOWWATCH THE COUNTER · THEN COMMIT THE BOUNDARY CRIME

The boundary crime scales with the window: a “per-day” quota can allow 2× at midnight. One fix is a sliding-window log of recent timestamps: exact, but memory hungry. A cheaper counter estimates the slide by weighting the previous fixed window by the fraction still overlapping the last minute, then adding this window’s count. It smooths the page turn, but its error depends on when requests arrived inside that previous window; there is no universal “few percent” bound. The deeper alternative changes shape: stop counting pages, start lending budget. That is the bucket.

MODEL NOTES — window drawn as 4 s, labeled “10 s window · time compressed” · the decision happens at arrival · the burst button stages both windows empty first, so the exploit is shown under its fair-weather conditions — which is exactly when it is exploited in the wild (idle quotas, midnight crons).

№ 03THE BUCKET

A loan, not a calendar.

The token bucket holds up to b tokens and refills at r per second. Each accepted request spends one. The two knobs set the sustained rate and the burst allowance. Switch to the leaky meter to see a stricter variant: with enough offered traffic it admits no more than r on average; idle time produces no output.

FIG. 02 — THE BUCKET · BURSTS ARE LOANSPICK A TRAFFIC SHAPE · THEN FLIP THE METER

Run the experiments the figure is built for. BURSTY + token bucket (b=15): each cloudburst largely passes — the bucket had been quietly refilling between rains — and the long-run rate lands on r. Now BURSTY + leaky: the same rain, and the door admits a metronome. Neither is wrong; they are different promises. And FLOOD + anything: the door holds at r forever, which is the entire point — a limiter is not fair, it is survivable. When you hear an API describe its limits as “rate and burst,” you now know its two knobs by their real names: r and b.

MODEL NOTES — decisions happen at arrival and refill is continuous · leaky mode is a reject-excess meter with capacity one · a queueing variant delays excess requests instead · dots sample about six arrivals each · rates are compressed. NGINX limit_req has separate burst, delay, and nodelay settings; this figure is one simplified configuration.

№ 04THE FLEET

Eight honest servers, one dishonest promise.

Your API runs on N stateless servers behind a load balancer. Each one dutifully limits each key to L requests per second — locally. Now the arithmetic you already distrust: the key’s true allowance is N × L, because the balancer spreads its traffic and no counter is shared. You have met this shape twice — The Nines multiplied availability down a chain; this multiplies your promise across a fleet. The cure costs a round trip.

FIG. 03 — THE FLEET · LOCAL HONESTY, GLOBAL LIESADD SERVERS · THEN CENTRALIZE THE COUNTER
—

Drag SERVERS and watch the delivered number climb while every server stays individually correct — the same backfire as ADD A COMPONENT in The Nines, with the sign flipped: there, adding boxes subtracted availability; here, adding boxes adds permission. Tick CENTRALIZED COUNTER and the fleet collapses into one honest number — paid for in a round trip to Redis per check, which is The Cache’s invoice reissued, and why production systems shard counters, sync them asynchronously, or accept a small window of overshoot. There is also the cheap wrong-feeling right answer: set each local limit to promise ÷ N. It honors the SLA — and starves any server that happens to be busy while its siblings idle, because local quotas know nothing about where traffic actually lands.

And when the door does say no, the rejection is a contract, not a shrug: status 429, a Retry-After: 12 header so the client can wait like a grown-up, an X-RateLimit-Remaining: 0 so it can stop before the wall. A polite client reads those and backs off with jitter — The Ack’s lesson, one layer up. A rude one retries in a tight loop and converts your safety mechanism into the outage. The door is half the system. The etiquette is the other half.

WHY THE LOCAL MULTIPLIER IS EXACTLY N — AND WHEN IT ISN’T

With a good load balancer, a heavy client’s requests are spread roughly evenly across N servers. Each server admits up to L locally, so the fleet admits up to N×L before any single door shuts.

The multiplier shrinks only when traffic skews — a client pinned to few connections, or balancer affinity — which is why some teams get away with local limits for years, until a client with a big connection pool arrives and finds the true ceiling.

The shared counter restores honesty by making the limit a fact about the key rather than the server — the same move as moving truth into one durable log in The Lock, or one outbox row in The Cache. One place where the number lives; everyone asks it.

MODEL NOTES — assumes even load balancing and no affinity · centralized counter cost modeled as one synchronous round trip per request · production softeners (sharded counters, async sync, small overshoot windows) noted in prose, not drawn.

№ 05FIELD GUIDE

The field guide.

Fixed window
Count per calendar page, reset at the turn. One integer of state — and blind at the boundary, twice the promise in a blink (FIG. 01).
Sliding window log
Keep each request’s timestamp; a request passes if fewer than N sit within the last T. Exact, boundary-proof — and O(N) memory per key, which is the bill.
Sliding window counter
Approximate the moving window: multiply the previous window’s count by the fraction of it still overlapping, then add the current count. It uses two counters per key, but can be inaccurate for uneven arrivals.
Token bucket
Budget b, refill r/s, spend one per request. Bursts forgiven to b; long run capped at r. The default because its two knobs are the two questions clients actually ask.
Leaky bucket (meter)
Under sustained demand, admit up to rate r and reject excess. An idle meter emits nothing. A queueing variant can delay accepted excess instead.
429 + Retry-After
The polite no: status, a wait hint, a remaining-allowance header. A limiter without them is a door that slams.
Distributed limiter
One shared counter (Redis + a Lua script for atomicity), or local counters synced async at the price of brief overshoot. Local-only limits multiply by the fleet — FIG. 03.
Load shedding vs throttling
Throttling refuses by identity (this key is over budget); shedding refuses by pressure (the system is drowning — The Cache’s stampede answer). A production door does both and labels which.
Reset stampede
A fixed window’s flip is a synchronized refill — every blocked client’s next attempt lands together. A mini-stampede (The Cache) on a timer; jitter the window edges or slide them.
№ 06WANNA TRY OUT WHAT YOU HAVE LEARNED?

Three scenarios.

From these figures, you can choose a limiter from the traffic shape and explain what its counter approximates. Try changing one assumption and check whether your explanation still holds.

The sealed sheetThree questions are sealed inside this sheet. Nobody is asked to open it — the door holds either way.Break the seal
question 1 of 3

The limit is 100 requests per calendar minute. A script fires 100 requests at 23:59:59.90 and 100 more at 00:00:00.10. How many are admitted?