Replication & Partitioning

4.Copies Everywhere

M

In this chapter

We'll separate replication (multiple copies, for durability) from partitioning (split data, for scale) — and see why a real system usually needs both.

8–10 min

The Problem in Real Life

It's a normal Tuesday afternoon — sellers actively editing listings, a few new signups coming in — when GreenMart's hosting provider sends an automated alert: a hardware fault on the machine running the database, forcing an emergency restart. For four minutes, nothing can reach the database at all. No listings load. No seller can save an edit.

It comes back on its own. Nothing is actually lost this time. But Mike can't stop thinking about the four minutes where it just wasn't there.

M

What happens to every seller's listings if that machine just... doesn't come back?

Mike

One Copy vs. Several

One copy is one failure away from gone

A single machine holding the only copy of live seller data is a real business risk, not a hypothetical one.

Replication = more copies of the same data

The same listings, kept in sync across multiple machines — so losing one machine doesn't mean losing the data.

Partitioning = splitting data into pieces

The same idea GreenMart already used to split Orders by city — spreading one dataset across machines instead of copying it whole.

Real systems usually need both

Each partition of the data still needs its own replicas — splitting for scale doesn't remove the risk of losing a single copy.

Replication & Partitioning

GreenMart has actually solved a version of this problem before. Once Orders outgrew what one table could comfortably hold, Sarah split it across separate tables by city — horizontal partitioning, often called sharding. That idea carries straight over to NoSQL databases, and for the same reason: one dataset, spread across multiple machines, so no single machine has to hold all of it. In most NoSQL databases, this isn't something Sarah sets up by hand with separate tables anymore — the database itself spreads data across machines automatically, based on a partition key Sarah chooses.

But partitioning was never built to solve today's problem. Splitting the listings data across machines doesn't help if any one of those machines can still just disappear, taking its slice with it. What today's four minutes actually exposed is a different gap entirely: there was exactly one copy of the listings data, on exactly one machine, anywhere.

Table — Three Copies of the Same Listings Data
CopyRoleIf this one machine fails
Primary (Machine A)Accepts every writeA replica takes over as the new primary
Replica (Machine B)Kept in sync with the primary, can serve readsNo data lost — two other copies still exist
Replica (Machine C)Kept in sync with the primary, can serve readsNo data lost — two other copies still exist

The same listing exists on all three machines at once. Losing any one of them still leaves two intact, complete copies behind.

This is replication: keeping multiple, kept-in-sync copies of the same data spread across different machines, specifically so that one machine failing doesn't mean the data is gone. A common shape is one primary — the one machine that accepts writes — plus one or more replicas, which stay synced with the primary and can often serve read traffic too. If the primary goes down, a replica is ready to take over.

Replication and partitioning solve genuinely different problems, and it's worth keeping them distinct. Partitioning splits a dataset into different pieces spread across machines — it solves outgrowing what one machine can hold. Replication keeps identical copies of the same data on different machines — it solves losing that data if one machine fails. A real, large-scale NoSQL system almost always needs both at once: each partition of the data still needs its own replicas, or that one partition is right back to being a single point of failure, just a smaller one.

Key Takeaway

Partitioning protects against outgrowing one machine. Replication protects against losing one. A real distributed system almost always needs both, for two entirely different reasons.

Why This Matters

The moment GreenMart's data exists as more than one copy, a new question shows up that a single copy never had to answer: what happens when the copies don't perfectly agree with each other, even for a moment? That question — not a hypothetical, but something every real distributed database has to make a real decision about — is exactly where the next chapter starts.

GreenMart's listings will live as three synced copies, not one. Mike feels better about that — right up until Sarah points out the harder question replication quietly creates: what if two of those copies briefly disagree?

Next