In this chapter
We'll meet Redis's everyday operating problems — hot keys, memory limits, and eviction policies — and put persistence, replication, and failure together into one real picture of what a restart actually recovers.
The Problem in Real Life
Persistence and replication cover the dramatic failures — a restart, a dead machine. Mike's last question is about the quieter kind: nothing crashes, and it's still slow, or still full.
Those are Redis's own everyday operating problems, distinct from any crash.
And if nothing crashes at all, but it still slows down or fills up — what then?
Mike
What Actually Goes Wrong Day to Day
Hot keys and eviction policies
One key can bottleneck a whole instance; a full, memory-bounded Redis has to evict something under a chosen policy.
A real failure scenario, put together
What actually gets recovered on a restart depends entirely on the persistence and replication choices made in advance.
Scaling, Hot Keys & Failure Scenarios
Two separate everyday problems hide inside Mike's question — one about traffic, one about memory — plus the honest, final answer to what a real restart actually recovers, now that every piece is on the table.
- On traffic — one key, overloaded. A hot key is Redis's own specific version of a hot shard: one key getting hit far more than any other. Because access to a single key is inherently handled one at a time, a hot key can bottleneck a whole Redis instance no matter how well everything else is distributed across a Cluster.
- On memory — running out of room. Memory management matters because Redis is memory-bounded in a way a disk-based database usually isn't — once it's full, something has to give. An eviction policy decides what: GreenMart can choose, for example, to evict whichever key was least recently used first, rather than refuse new writes outright.
Put together, with persistence and replication both on the table, a real Redis failure scenario looks like this: the machine restarts, and whatever's recovered depends entirely on what GreenMart chose earlier. No persistence at all — everything's gone, and for a pure cache in front of a real database, that's often a perfectly acceptable, deliberate choice, not a disaster. RDB only — recovers to the last snapshot, losing anything written after it. AOF enabled — recovers close to everything, at the cost of the write overhead paid the whole time it was running. And through all of it, replication and Sentinel decide whether GreenMart even notices the restart at all, or a replica has already taken over.
None of this is optional risk GreenMart stumbled into. It's a real, deliberate set of trade-offs — exactly the same kind of trade-off this whole Act opened with: memory over disk, for speed, with the specifics of what that costs finally spelled out in full.
Key Takeaway
A hot key and a full memory aren't crashes — they're the everyday cost of running Redis at real scale, and eviction policies and traffic-aware key design are how GreenMart keeps them from becoming one.
Why This Matters
This closes the loop the very first chapter of this Act opened: Redis is fast because it lives in memory, and memory has real, honest costs that only show up when something goes wrong — or when nothing crashes, but the traffic itself becomes the problem. Every production Redis deployment makes these exact choices — this chapter is what makes it an informed decision instead of an unpleasant surprise during an actual outage.
GreenMart now understands Redis's speed and its real limits, in the same breath. The checkpoint ahead asks Sarah — or you — to design GreenMart's actual flash-sale cache from scratch, with every one of this Act's trade-offs made on purpose.
