Compression & Storage

7.Compression & Storage

M

In this chapter

We'll meet compression as a genuinely separate, stackable real lever from serialization format — shrinking any bytes by finding real redundancy, at a real, deliberate trade-off between compression ratio and CPU cost — and land on the right real choice for GreenMart's own event log: a stronger, slower algorithm, since it's written constantly but read rarely.

7–9 min

The Problem in Real Life

The binary format helped — real, meaningfully smaller events. But GreenMart's finance team asks a real, obvious next question. "Can we shrink it further? Every byte we're not storing is real money we're not spending."

Sarah nods. "There's a real, separate lever left — one that works on any bytes at all, whatever format they're already in."

M

We already made the format smaller. Is there anything else we can actually do?

Mike

A Smaller Format vs. The Same Bytes, Compressed

A real, separate, stackable lever

Compression works on any bytes regardless of format — it's a genuinely different decision from choosing JSON vs. binary.

Ratio vs. CPU cost, paid on both ends

A stronger algorithm shrinks data further but costs more CPU at both compression and decompression time.

Compression & Storage

Compression is a real, genuinely separate lever from serialization format — it works on any real sequence of bytes, finding and removing real, statistical redundancy within them, regardless of whether those bytes are JSON text or an already-compact binary format.

  • Why compression and format choice are real, separate, stackable decisions. The last chapter picked a real format — how the data is structured into bytes. Compression is a genuinely different, later real step — taking whatever bytes result and shrinking them further by finding real, repeated patterns. GreenMart can compress JSON, compress a binary format, or compress neither — the two decisions don't conflict, and are often applied together for real, compounding savings.
  • The real, fundamental trade-off: ratio vs. CPU cost. A compression algorithm that searches harder for real redundancy produces a genuinely smaller result (a better real compression ratio) but takes real, measurable CPU time to do that searching — and that same real CPU cost is paid again, in reverse, every time the data is decompressed for reading. A faster, lighter algorithm compresses and decompresses quickly but leaves real, avoidable size on the table; a slower, more thorough one shrinks further but costs more real CPU on both ends.
  • Real algorithm families, not one universal answer. Some real algorithms (like gzip) favor a strong real compression ratio at a real, higher CPU cost; others (like Snappy or LZ4) favor real speed, accepting a somewhat weaker ratio in exchange for compressing and decompressing extremely fast — a genuinely better fit when data is read constantly and CPU time matters more than every last real byte saved.
  • GreenMart's own, real, correct choice. The event log is written constantly and read relatively rarely (mostly for occasional debugging or analysis) — a real workload shape that favors a stronger, slower real algorithm, since the CPU cost is paid mostly at write time, once, in exchange for real, meaningful storage savings that persist for as long as the data is kept. A real, different workload — data read constantly, on every request — would instead favor a fast, lighter algorithm, the same real "know your access pattern first" discipline Act 5's own lifecycle-management chapter applied to storage-class choices.
Table — Compression Algorithm Trade-offs — The Real Choice
PriorityReal Trade-offFits Best When
Strong compression ratio (e.g. gzip-family)Higher CPU cost, smaller real resultWritten once, read rarely — cost paid mostly at write time
Fast compression/decompression (e.g. Snappy, LZ4-family)Weaker ratio, lower CPU costRead constantly — CPU cost paid repeatedly, so speed matters more

GreenMart's own event log — written constantly, read rarely — favors a stronger, slower algorithm; a frequently-read dataset would favor the opposite choice.

GreenMart now has a real, second, genuinely separate lever on the exact same cost problem — stacked on top of the binary format decision from the last chapter, not competing with it, for real, compounding storage savings.

Key Takeaway

Compression is a genuinely separate, stackable real lever from serialization format — it shrinks whatever bytes already exist by finding real redundancy, and the right real algorithm choice depends on the same real trade-off every storage decision in this course keeps returning to: how often the data is written versus how often it's actually read.

Why This Matters

As GreenMart's own event log and other growing datasets accumulate real, ongoing storage cost, compression is a real, direct, deliberate lever GreenMart can apply independently of format — genuine, compounding savings for a real, deliberate CPU trade-off, not a free, magic reduction.

GreenMart now has both real, stackable levers on storage cost — format choice and compression — and the real workload-shape discipline for choosing each deliberately. The Act's final chapter turns every mechanism covered so far into one real, practical framework for choosing a storage format.

Next