Distributed Systems Checkpoint

Fix GreenMart's Split-Brain Incident

Reasoning Checkpoint

A design challenge, worked through in writing — no auto-grading, just a real attempt.

25–30 min

The Challenge

Twelve chapters of real distributed-systems reasoning, all converging on one real challenge: fix GreenMart's actual split-brain incident, precisely — not with a general appreciation that "distributed systems are hard," but with the real vocabulary and real mechanisms this Act built, one at a time.

You have the actual incident in front of you: two regions, a genuine network partition, both regions independently reporting 3 units of stock and both selling against it, 8 real orders landing against 3 real units once reunited. Write out Sarah's actual fix: what the incident precisely was, in this Act's own real terms, and what concrete, standing design would prevent it from happening again — plus what should happen to the conflicting data that already exists.

What Your Fix Needs

  • Name the incident precisely using this Act's own vocabulary — which CAP-theorem trade-off each region made during the partition, and why that alone explains the diverging stock counts.
  • Using the real quorum formula (R + W > N), explain what GreenMart's effective read/write quorum actually was during the incident, and propose a specific, real quorum configuration that would have prevented two regions from both confirming a sale without checking each other.
  • Explain, precisely, why this incident is genuinely split-brain and not just "eventual consistency working as intended" — and propose the specific consensus mechanism (majority-quorum leader election) that would have prevented it.
  • Propose one concrete piece of standing multi-region design — tied explicitly to either Cassandra's hinted handoff or DynamoDB's Global Tables, both already real, working answers this course has built — that GreenMart should have had in place before the sale, not improvised during it.
  • Separately from prevention: explain what a conflict-free approach (a CRDT, like the PN-Counter worked example) would and would not have changed about this specific incident, and name a real RTO/RPO pair GreenMart should set for this specific, scarce-inventory data going forward.
Stuck? A Few Hints
  • Reread the CAP-theorem chapter's own precise framing — the real question is which trade-off each region made once the partition happened, not whether "CAP was violated."
  • The quorum-mathematics chapter's own worked table is the exact tool for the second requirement — work out what R and W effectively were during the incident before proposing new ones.
  • The split-brain chapter draws a real, specific line between split-brain and ordinary eventual consistency — the difference is whether anything ever required agreement before acting as authoritative, not just whether views diverged. That same chapter's own PN-Counter table is the right reference for the fifth requirement.

Before You Move On

Notice what this checkpoint never asked: a promise that distributed systems can be made perfectly simple. That was never the point — the point was replacing "we don't know why this happened" with a real, precise diagnosis and a real, specific fix, using this Act's own vocabulary throughout. The final Act of this course asks a different, calmer question: given everything covered across all thirteen Acts, how does GreenMart actually argue for the right database, not just pick one that works.

Next