Replication & High Availability

7.Enough Copies to Survive Anything

M

In this chapter

We'll meet Redis's answers to losing a whole machine — replication, Sentinel's automatic failover, and Cluster's built-in sharding.

9–11 min

The Problem in Real Life

Persistence answers what survives a restart of the same machine. Mike's next question is bigger: what if the machine itself is gone — not restarted, just gone?

RDB and AOF don't help there. Both live on the same disk, on the same machine, that just disappeared.

M

Okay, but what if the whole server dies mid-sale? Not restarts — dies.

Mike

Surviving a Machine, Not Just a Restart

Replication and Sentinel

Copies for durability, plus automatic failover so a dead primary doesn't mean a dead cache.

Redis Cluster is sharding, not failover

Cluster splits the keyspace across nodes for capacity — a separate job from Sentinel's availability role.

Replication & High Availability

Persistence protects against losing data. It does nothing if the machine holding that data is simply unreachable. That needs copies, living somewhere else.

  • Replication works the same way it has everywhere else in this course: one primary, one or more replicas continuously copying from it. If the primary goes down, a replica already has (nearly) everything it had, on a different machine.
  • Redis Sentinel watches a small primary-replica setup and handles automatic failover — promoting a replica to primary if the current primary goes down, and redirecting clients to it — the same job an election played for MongoDB's replica sets, Redis's own name for it.
  • Redis Cluster is the bigger-scale answer to a different problem: not just surviving one machine's death, but holding more data than one machine can. Cluster is Redis's own built-in sharding — splitting the keyspace itself across multiple nodes, each owning a slice of it, so no single machine has to hold all of it, the same idea as MongoDB's sharding, applied to Redis's own architecture.

Replication plus Sentinel answers "what if one machine dies" without needing to touch the data model at all. Cluster answers a separate question — "what if the data no longer fits on one machine" — and both can run together.

Key Takeaway

Replication and Sentinel buy availability; Cluster buys capacity. They solve different problems, and a real production Redis deployment often needs both, not one instead of the other.

Why This Matters

A flash sale that goes down because one Redis machine failed, with no replica ready to take over, is a self-inflicted outage — the exact kind this Act exists to prevent. Sentinel and Cluster are how GreenMart turns "we hope the server stays up" into "we've already planned for it not to."

GreenMart now has real durability (persistence) and real availability (replication and Sentinel) covered, plus a path to scale beyond one machine (Cluster). One category of problem is still open: what goes wrong day to day, even when nothing has crashed.

Next