In this chapter
We'll meet the cache stampede — a real, dramatic failure mode where a popular entry's synchronized expiry sends every request that would have been a cheap hit to the real origin all at once — and cover the three real mitigations (locking, jittered TTLs, refresh-ahead) that prevent it.
The Problem in Real Life
A flash sale goes live. Traffic spikes exactly as planned — until, three minutes in, the database falls over completely. Mike is baffled. "We just fixed caching. How did adding more caching make this worse?"
Sarah checks the timing. The database crash landed exactly 60 seconds after the sale page's cache entry — a nice, round TTL — expired. "It's not that caching failed," she says. "It's that it failed all at once."
We just fixed our caching. How did the database still fall over?
Mike
One Request Rebuilding the Cache vs. Thousands, All at Once
Many requests, one synchronized miss
When a popular entry expires, every request behind it can miss the cache at the exact same instant, all hitting the origin together.
Locking, jitter, and refresh-ahead
Three real mitigations, each working by breaking the synchronization that turns one expiry into a coordinated overload.
Cache Stampede
A cache stampede (also called a thundering herd) is a real, specific failure mode: a popular cache entry expires, and every request that would have been served from it — sometimes thousands, arriving within the same real second — instead misses the cache simultaneously and hits the real, slower origin at once, often enough to genuinely overwhelm it.
- Exactly how it happens. A cache entry with real, heavy traffic behind it — GreenMart's own flash-sale page — expires via its TTL. The very next request after that moment finds a cache miss, and, following the normal, reasonable logic, goes to the real origin to recompute or refetch the answer. Completely unremarkable, on its own. The real problem: if traffic to that entry is high enough, dozens or thousands of other requests arrive in that same real gap before the first request's fresh answer makes it back into the cache — and every single one of them, finding the exact same miss, independently does the exact same expensive real work, all at once.
- Why this is genuinely worse than steady load. GreenMart's real database was built to handle the sale's actual traffic just fine, spread out normally across the cache's real hit rate. A stampede doesn't increase real total traffic — it removes the cache's own real protection for one brief, synchronized instant, letting a spike of requests that would normally each cost nothing (a cache hit) all suddenly cost the real, full, expensive price at the exact same moment.
- Mitigation 1 — a lock, so only one request rebuilds it. The first request to hit a real miss acquires a real, temporary lock signaling "I'm already rebuilding this entry." Every other request arriving during that same window sees the lock and either waits for the real rebuild to finish (and then gets the freshly-cached result) or is served a real, slightly-stale copy in the meantime — either way, the origin only gets hit once, not thousands of times.
- Mitigation 2 — jittered TTLs, so entries don't expire in lockstep. Instead of every related cache entry sharing the exact same fixed TTL — the real root cause of GreenMart's own synchronized crash — a small, random amount of extra time (jitter) gets added to each entry's real expiration individually, spreading real expirations out over a few seconds or minutes instead of all landing in the same instant.
- Mitigation 3 — refresh ahead of expiry, not after. Rather than waiting for an entry to actually expire before recomputing it, a background real process refreshes popular entries proactively, shortly before their real TTL runs out — so by the time real traffic would have caused a miss, a fresh copy is already quietly in place, and no request ever actually experiences the gap at all.
GreenMart now has the real, exact name for what took the database down — not "caching stopped working," but a stampede: caching's own real protection vanishing for one synchronized instant, precisely because every related entry was set up to expire at exactly the same time.
Key Takeaway
A cache stampede happens when a popular entry's expiry removes a cache's real protection for one synchronized instant — every real request that would have been a cheap hit becomes an expensive miss at the exact same moment, and real fixes (locking, jitter, refresh-ahead) all work by breaking that synchronization.
Why This Matters
Any high-traffic moment GreenMart plans for — a flash sale, a major restock announcement, a viral product — is exactly when a cache stampede is most likely, and most damaging. Understanding the real mechanism turns "the database mysteriously fell over during our biggest sale" into a specific, fixable, well-understood real failure mode.
GreenMart now has the real vocabulary and real mitigations for cache stampedes — locking, jitter, and refresh-ahead. The Act's final chapter turns every mechanism covered so far into one real, practical framework for deciding when, what, and how to cache.
