Cache Invalidation

2.Cache Invalidation

M

In this chapter

We'll open up cache invalidation — the real, famously hard problem of knowing when a cached copy has gone stale — and see the two real strategies, time-based expiry and event-based signaling, and exactly which one GreenMart's own stock page was missing.

7–9 min

The Problem in Real Life

Sarah finds the proxy's own config: pages are cached for 15 minutes, no matter what. "So even without the stock count changing," Mike asks, "how would the cache ever know to update itself sooner, if it needed to?"

Sarah nods. "That's the actual, real problem underneath all of this. A cache doesn't know anything on its own. Something has to tell it."

M

How does a cache know when the real data behind it has changed?

Mike

A Cache That Waits vs. A Cache That's Told

Time-based — expires on a timer, no cooperation needed

Simple and requires nothing from the data source, but accepts a real, honest window of possible staleness.

Event-based — told the instant data changes

Far more accurate, but requires real, deliberate coordination between whatever writes the data and whatever caches it.

Cache Invalidation

Cache invalidation is the real, genuinely hard problem of deciding when a cached copy should stop being trusted. A cache, left alone, has no idea the real data behind it has changed — every invalidation strategy is really a different answer to "how does it find out?"

  • A real, honestly hard problem — not a slogan. Cache invalidation is famously named as one of the genuinely hard problems in computer science, and the reason is concrete, not clever wordplay: it requires correctly predicting, ahead of time, exactly when a piece of cached data will stop matching reality, for every single way that data could possibly change — including changes the caching layer itself might never be told about.
  • Strategy one — time-based (expire it eventually). The cache simply assumes its own copy is only good for a set period, and throws it away (or checks in on it) once that period passes, regardless of whether the real data actually changed. This is simple to implement and requires zero cooperation from whatever produces the real data — but it accepts a real window where the cache can be confidently, silently wrong, for however long that period lasts.
  • Strategy two — event-based (get told directly). The system that actually changes the real data takes real responsibility for telling the cache: the moment GreenMart's own stock count updates, that same write also triggers a real, deliberate signal — invalidate this cached entry, or update it directly. This can be far more accurate than time-based expiry, since the cache learns about a change the instant it happens, but it requires real, deliberate coordination between whatever writes the data and whatever caches it — coordination that's easy to forget in exactly the kind of ad-hoc, un-designed cache GreenMart's own incident revealed.
  • GreenMart's own incident, named precisely. The proxy's 15-minute, purely time-based expiry is real strategy one, alone, with nothing backing it up — no event-based signal existed anywhere to tell it the stock count had actually changed sooner. The stock page wasn't wrong because caching is inherently unreliable; it was wrong because only the weaker of the two real strategies was in place, for data that genuinely needed the stronger one.

This is the real fork every caching decision eventually reaches: accept a bounded, honest window of possible staleness (time-based), or build real, deliberate coordination so the cache finds out the moment truth changes (event-based). The next two chapters go deep on the mechanics of the first path; a later chapter in this Act covers exactly how the second path gets built.

Key Takeaway

A cache never knows on its own that it's gone stale — it either assumes a fixed shelf life and expires on a timer, or it's told directly, the instant the real data changes. Every invalidation strategy is really just a choice between those two real answers.

Why This Matters

Every stale-data incident GreenMart will ever debug — a price that's wrong for a few minutes, a review count that lags, a stock number that's confidently false — traces back to exactly this one real question: which of the two invalidation strategies was actually in place, and was it the right one for how often, and how badly, that specific data can afford to be wrong?

GreenMart now has the real vocabulary for why the stock page went wrong: time-based expiry alone, with no event-based signal for data that genuinely needed one. The next chapter goes deep on the time-based side of this — TTL — and the genuinely separate real question of what happens when a cache simply runs out of room.

Next