Storage Reliability & Recovery

5.Storage Reliability & Recovery

M

In this chapter

We'll unify crash consistency (journaling, Act 3), durability (physically-separated copies, Act 5), and recoverability (versioning, Act 5) into one real reliability framework, and run GreenMart's own honest audit — finding that some of its storage layers have explicitly quantified reliability, while others have only ever assumed it.

7–9 min

The Problem in Real Life

Mike asks a real, final, uncomfortable question before the capstone. "Across everything we've built — the file system, the caches, the object storage — if something actually goes wrong, right now, do we genuinely know what survives?"

Sarah pauses. "Honestly? Let's actually check, layer by layer, instead of assuming."

M

If something actually goes wrong right now, do we genuinely know what survives and what doesn't?

Mike

Assuming Reliability vs. Actually Naming It, Layer by Layer

Reliability is layered, not one property

Crash consistency, durability, and recoverability are genuinely separate guarantees — each covered by a specific mechanism already met.

Unnamed reliability is only assumed

A guarantee that isn't explicitly stated for a given storage layer is a real, honest risk, not a confirmed protection.

Storage Reliability & Recovery

Storage reliability isn't one real property a system either has or lacks — it's a real, layered set of guarantees this whole course has already built, mechanism by mechanism, and a genuinely reliable architecture means naming, deliberately, exactly what survives at every one of those layers.

  • Crash consistency — surviving a mid-write failure. Act 3's own journaling chapter already covered this: a file system's own real structures stay consistent even if power cuts mid-write, by writing intent down first. Any real storage layer GreenMart builds on top of a file system inherits this real guarantee — but only for the file system's own structures, not automatically for GreenMart's own application-level data unless it makes a similar real, deliberate choice.
  • Durability — surviving a hardware failure. Act 5's own real durability chapter covered this precisely: multiple, physically-separated real copies, checked continuously for silent corruption, are what let GreenMart state a genuine, quantifiable survival probability for its own object storage. GreenMart's relational database and Cassandra cluster need this same real question asked explicitly — how many real copies exist, and how physically separated are they, really?
  • Recoverability — surviving a real, human mistake. Act 5's own versioning chapter already covered this: keeping prior real versions turns an accidental overwrite from permanent loss into a recoverable event. This same real principle applies everywhere data can be overwritten or deleted — a database backup strategy is really the exact same real idea, applied to a different real storage type.
  • GreenMart's own, real, honest audit. Object storage: durability is real and quantified (Act 5). File system: crash consistency is real, via journaling (Act 3). The relational database and Cassandra: real replication exists, but GreenMart has never actually stated its own durability number the way object storage's own "11 nines" was named precisely. This chapter's own honest point: reliability that isn't named explicitly, layer by layer, is reliability GreenMart is only assuming, not actually verifying.

GreenMart now has the real, complete reliability vocabulary this whole course has quietly built — crash consistency, durability, and recoverability — and the honest discipline of naming each one explicitly, for every real storage layer, rather than assuming it by default.

Key Takeaway

Storage reliability is a real, layered set of guarantees — crash consistency, durability, and recoverability — each already built by a specific mechanism this course covered, and a genuinely reliable architecture means naming, explicitly, exactly which of these guarantees each storage layer actually provides, rather than assuming them.

Why This Matters

Before GreenMart commits its next real expansion to any specific storage architecture, explicitly naming each layer's real reliability guarantees — rather than assuming them — is what turns a genuine incident from a surprise into a real, already-understood, already-accepted risk.

GreenMart now has the real, complete reliability framework spanning every storage layer this course has covered. The Act's final reading chapter turns everything covered so far into one, real, complete process: designing a full storage architecture.

Next