Backup & Recovery

2.The Day the Server Died

M

In this chapter

GreenMart's server dies overnight, and the only question that matters is whether anyone ever actually tested a restore — we look at recovery point/time objectives, the gap every backup schedule leaves open, and why SQLite's real backup features (VACUUM INTO, WAL mode) can't be honestly demoed in a browser sandbox.

10–12 min

The Problem in Real Life

It's 6 AM when Mike's phone rings. The server hosting GreenMart's database has crashed overnight — a hardware failure, nothing anyone did wrong. By the time Sarah gets a new server running an hour later, one question matters more than anything else: is the data still there?

Mike asks the question he's never had to ask before. "We do... back this up somewhere, right?" Sarah realizes she genuinely isn't sure. Nobody ever tested it.

M

We do back this up somewhere, right?

Mike

Hoping the Data Survives vs. Knowing It Will

A backup schedule always leaves a gap

Anything that changed since the last backup is at risk if a failure happens before the next one runs.

An untested backup is an assumption, not a plan

The only way to actually know a restore works is to have already done it, before an emergency forces the question.

RPO and RTO measure two different things

Recovery Point Objective is how much data you can afford to lose; Recovery Time Objective is how long you can afford to be down.

SQLite has real backup tools, honestly out of reach here

VACUUM INTO and WAL mode are genuine SQLite features that this browser Playground can't meaningfully demonstrate.

What Actually Makes a Backup Strategy Real

A backup is a saved copy of the database, taken at some point in time, that can be restored if the original is lost or corrupted. The gap between when a backup was last taken and when a failure happens is real, unavoidable, and has a name: the recovery point objective (RPO) — the maximum amount of data a business is willing to lose. Backing up once a day means an RPO of up to 24 hours of lost data if disaster strikes right before the next backup.

How long it takes to actually get back up and running afterward is a separate number: the recovery time objective (RTO). A backup that exists but takes three days to restore, or that nobody has ever actually tested restoring, isn't a real safety net — it's an assumption.

The Gap a Backup Schedule Leaves Open

Backup Taken

11:00 PM — a full, known-good copy exists

hours pass

Normal Operation

Orders, updates, new customers — none of this is in the backup yet

unplanned failure

Server Crashes

6:00 AM — everything since 11:00 PM is at risk

recovery

Restore From Backup

Recovers everything up to 11:00 PM — the RPO gap (7 hours) is genuinely gone

SQLite does have real backup mechanisms worth knowing by name: the VACUUM INTO 'filename' SQL statement writes a complete, consistent copy of the database to a new file, and WAL mode (write-ahead logging) keeps a log of recent changes that helps the database recover cleanly if it's interrupted mid-write. Both are genuine, useful SQLite features — but neither can be meaningfully demonstrated in this browser-based Playground, since there's no real, restorable file system here to actually prove a backup and restore against.

The lesson that actually matters isn't a specific command — it's the discipline behind Mike's 6 AM question. A backup you've never tried to restore is a belief, not a plan. The only way to know a backup actually works is to have already restored from it, on purpose, before the day you desperately need to.

Key Takeaway

A backup schedule doesn't eliminate data loss — it bounds it. The real measure of a backup strategy isn't whether backups exist, but whether anyone has actually verified a restore works, before a real crash forces the question.

Why This Matters

An untested backup is one of the most common false senses of security in real operations — teams that assume backups work, because a backup job runs on a schedule, often discover the restore process is broken only during an actual outage, when it's far too late to fix calmly.

GreenMart survives its first real outage — barely, and only because a backup happened to exist. The next chapter looks at what it actually takes to run across multiple cities without every city sharing a single point of failure.

Next