In this chapter
We'll turn all three of this Act's real layers — storage engine, serialization format, and compression — into one real, ordered decision process, and close with GreenMart's own event log fully, deliberately architected: LSM-tree-based, in a binary format, compressed with a strong algorithm, each layer matched to its own real question.
The Problem in Real Life
Sarah stands in front of a whiteboard covering everything this Act built — storage engines, serialization, compression. "We started this Act with one confusing question," she says. "'Why does one database choke on writes the other handles easily?' We're closing it with a real, repeatable process for every storage decision like that one, going forward."
Mike looks at the board. "So walk me through it — start to finish, for something new."
Walk me through it, start to finish — how do we actually decide, next time?
Mike
Three Real, Separate Layers vs. One Real Decision, Made Together
Three real, stackable layers
Storage engine, serialization format, and compression are genuinely separate decisions that compound rather than compete.
Each layer, matched to its own real question
Read/write shape decides the engine; who reads raw bytes decides the format; write/read frequency decides compression.
Choosing a Storage Format
Every mechanism this Act covered sits at one of three genuinely separate, real, stackable layers — storage engine, serialization format, compression — and a real, complete storage decision means choosing deliberately at each one, matched to the actual real shape of the data involved.
- Layer one — storage engine, matched to read/write shape. (Chapters 1-4.) Read-heavy or steady-write data favors a B-tree-based system; write-heavy, high-throughput, append-style data favors an LSM-tree-based one. GreenMart's own real, correct answer for the event log: LSM-tree-based (Cassandra) — the exact fix this Act's own opening incident needed.
- Layer two — serialization format, matched to who actually reads the raw bytes. (Chapters 5-6.) Data inspected directly by a person favors JSON's own real readability; high-volume data read only by internal systems that already agree on its shape favors a genuinely more compact binary format. GreenMart's own real answer: a binary format, for exactly that reason.
- Layer three — compression, matched to write/read frequency. (Chapter 7.) Data written far more often than it's read favors a stronger, slower real algorithm; data read constantly favors a faster, lighter one. GreenMart's own real answer: a stronger algorithm, since the event log is written constantly but read only occasionally.
- These three real layers stack — they don't compete. GreenMart's own event log ended up LSM-tree-based, in a binary format, compressed with a strong algorithm — three genuinely separate, real decisions, each made independently at its own layer, each compounding with the others rather than replacing them.
| Layer | Real Question | GreenMart's Answer |
|---|---|---|
| Storage engine | Read-heavy/steady writes, or write-heavy/bursty? | LSM tree (Cassandra) — write-heavy, append-style |
| Serialization format | Read by a person, or only by internal systems? | Binary format — high-volume, internally consumed |
| Compression | Written often and read rarely, or read constantly? | Strong, slower algorithm — written constantly, read rarely |
Three genuinely separate, real, stackable decisions — the same real discipline applied at each layer, not one universal answer copied across all three.
GreenMart closes this Act having done, for real, what the opening chapter only diagnosed: turned "one database chokes, the other doesn't" into three specific, deliberate, layered real decisions — engine, format, compression — each matched to the actual, real shape of the data, not chosen by habit or convenience.
Key Takeaway
"How should this data actually be stored" was never one question — it's three real, separate, stackable layers (storage engine, serialization format, compression), each with its own real decision, and this Act just handed GreenMart all three, proven against a real incident from start to finish.
Why This Matters
Every future data-heavy real feature GreenMart builds — a new analytics pipeline, a new high-volume log, a new archive — now has a real, repeatable, three-layer decision process behind it, instead of defaulting to whatever storage engine and format happen to already be in use.
GreenMart closes Act 6 with the write-heavy-workload mystery fully, physically explained, and a real framework for every storage-layout decision still to come. Act 7 turns to the real, final synthesis of this whole course: designing GreenMart's own complete storage architecture.
