In this chapter
We'll meet failover and fault tolerance, tied back to Act 5's hinted handoff and Act 4's Global Tables as GreenMart's own prior real answers, multi-region design and disaster recovery as the standing, planned-in-advance version of this whole Act's toolkit, and RTO/RPO as the real, checkable numbers that prove a recovery plan actually works — closing with a direct, honest roundup of the real distributed failure scenarios this course has now named one by one.
The Problem in Real Life
Mike asks the last real question this Act needs answered: not what went wrong, not even the fix for that one incident — what does GreenMart actually do, on purpose, in advance, so the next regional outage is a non-event instead of a war room?
Sarah realizes this Act's whole toolkit — quorum, consensus, consistent hashing — was never really about one incident. It's the real, standing design GreenMart needs running underneath every region, all the time, not just during a crisis.
The next outage shouldn't be a war room. It should be Tuesday.
Sarah
Reacting to an Outage vs. Being Built for One
Failover shifts responsibility automatically
The same real idea behind Cassandra's hinted handoff and DynamoDB's Global Tables — now named as the general pattern.
Fault tolerance means working while broken
Not "eventually recovers" — a system that keeps functioning correctly while parts of it are genuinely down.
Multi-region design happens in advance
Replication strategy, quorum, and consensus chosen deliberately before an outage — not improvised during one.
RTO and RPO make a plan checkable
Real, named numbers per kind of data — how long down, how much lost — proving a plan actually works, not just exists.
Failover, Disaster Recovery & Multi-Region Design
Three real, separate questions, the same pattern this course keeps returning to. What actually happens the moment a region genuinely fails? What does GreenMart's multi-region design need to look like, standing, so that failure is survivable by default? And — practically — how does GreenMart even know, in advance, whether its own plan would actually work?
- On what happens when a region fails. Failover is the real, automatic process of shifting traffic and responsibility away from a failed region onto the ones still healthy — this course has already met real, specific versions of this idea: Act 5's Cassandra hinted handoff (a healthy replica temporarily absorbing a down node's writes) and Act 4's DynamoDB Global Tables (multi-region replication designed specifically so one region's failure doesn't take the whole system down) were both real, working answers to pieces of this exact problem, before this Act ever named the general pattern. Fault tolerance is the broader property failover is built to provide: a system continuing to function correctly, for real, even while some of its parts have genuinely failed — not "eventually recovers," but "keeps working while it's actually broken."
- On designing for it in advance. Multi-region databases, done deliberately, mean choosing real replication strategy, real quorum requirements, and real consensus mechanisms before an outage, not improvising them during one — precisely the standing design this Act's own incident revealed GreenMart never actually had. Disaster recovery is the honest, complete plan for the worst real case: not just one link breaking, but an entire region genuinely disappearing — real backups, a real, tested plan for which other region takes over, and a real, known answer for how much data (if any) could genuinely be lost in the worst case, decided in advance, not discovered live.
- On actually knowing the plan would work. A disaster-recovery plan nobody has tested is really just a hope with extra paperwork. Two real, standard metrics make a plan honestly checkable instead of aspirational: RTO (Recovery Time Objective — how long the system is allowed to actually be down before failover completes) and RPO (Recovery Point Objective — how much recent data, at most, the business accepts might be lost in the worst case, measured as a real time window). Naming both numbers, on purpose, for GreenMart's own scarce-inventory data specifically — and then genuinely testing whether the real system actually meets them — is what turns "we have a disaster recovery plan" from a sentence into something actually true.
| Data | RTO (max downtime accepted) | RPO (max data loss accepted) | Why |
|---|---|---|---|
| Scarce inventory (gift box, 3 units) | Under 1 minute | Zero — no lost sale record | A real, immediate business cost if wrong |
| Ordinary catalog browsing | Several minutes | A few minutes of staleness | Inconvenient, not costly, if briefly behind |
| Order history archive | Up to an hour | Up to an hour | Read rarely, urgency is genuinely lower |
The real point: RTO and RPO aren't one number for the whole company — like the quorum chapter's own lesson, they're chosen deliberately, per kind of data, based on what being wrong or being down actually costs.
Distributed database failure scenarios, named directly and honestly, one more time: a network partition (chapter 1's own incident), a coordinator failure mid-transaction (the previous chapter), a minority side incorrectly believing it's still in charge (split-brain), an entire region going dark (this chapter). None of these are edge cases at real, multi-region scale — they're the normal, expected shape of what eventually happens, which is exactly why every mechanism this Act has taught (quorum, consensus, consistent hashing, two-phase commit) exists to be designed in before one of these happens, not improvised during it.
Key Takeaway
Distributed databases aren't difficult because storing data is difficult. They're difficult because machines fail independently while the application still expects one coherent system — and the real answer was never to prevent failure entirely, it's to design, in advance, for a system that keeps behaving correctly while parts of it are genuinely broken, with real, named numbers (RTO, RPO) proving the plan actually works, the same way Act 5's hinted handoff and Act 4's Global Tables already did, one specific piece at a time.
Why This Matters
This closes the loop this Act's very first chapter opened with: GreenMart's own incident wasn't a freak event, it was the predictable result of running multi-region without a standing design for exactly this — and this chapter is the honest, complete, and checkable version of what that standing design actually requires.
GreenMart now has a real, complete, and checkable picture of what it takes to survive the next regional outage by design, not by luck. The checkpoint ahead asks Sarah — or you — to actually fix GreenMart's own split-brain incident, with every one of this Act's lessons made on purpose.
