B-Tree vs. LSM Tree

4.B-Tree vs. LSM Tree

M

In this chapter

We'll put B-trees and LSM trees directly side by side — two real, important families of storage structures making genuinely different bets about when to pay the cost of organizing data — and close with the precise, physical explanation for GreenMart's own opening incident: its write-heavy event log needed the LSM-tree-based system, not the B-tree-based one.

8–10 min

The Problem in Real Life

Sarah finally draws the two real trees side by side. "Now you can actually see it," she says. "Not 'SQL vs. NoSQL.' Not 'relational vs. Cassandra.' Just two real, different physical bets about where the real cost of storing data should land."

Mike studies both. "So which one is actually better?"

M

Now that we've seen both — which one is actually the better storage engine?

Mike

Two Real Bets, Not a Better and a Worse Option

Two real bets, not a hierarchy

B-trees pay organizational cost at write time for fast reads; LSM trees defer it for fast writes — neither is universally better.

Match the engine to the real workload shape

Read-heavy or steady writes favor B-trees; write-heavy, high-throughput, append-style workloads favor LSM trees.

B-Tree vs. LSM Tree

B-trees and LSM trees are two important families of storage and indexing structures used by many database systems — genuinely different real bets about when to pay the cost of keeping data organized, not a hierarchy with one correct winner.

  • The real bet each one makes. A B-tree pays real organizational cost at write time — keeping data thoroughly sorted and structured on disk continuously — in exchange for genuinely fast, direct reads. An LSM tree defers that real organizational cost to a background compaction process, accepting writes as fast as physically possible, at the real, honest cost of potentially checking multiple places on a read.
  • Where each real bet actually pays off. A B-tree genuinely excels at workloads dominated by reads, or by writes spread out over real time rather than arriving in a sudden, heavy burst — GreenMart's own relational database, handling real customer lookups and moderate order-writing traffic, fits this well under ordinary conditions. An LSM tree genuinely excels at write-heavy, high-throughput workloads — exactly the real event-logging pattern that overwhelmed the relational database but barely registered on Cassandra.
  • GreenMart's own real, correct fix, named precisely. The new event log shouldn't have gone to the relational database at all — its real write pattern (constant, high-volume, append-style) is precisely the shape an LSM-tree-based system is built for, which is exactly why moving it to the existing Cassandra cluster (already familiar, now understood at the real physical layer for the first time) solved the problem completely, without either database needing to change or scale up.
  • A real, honest caution: this isn't the only real axis, and these aren't the only two families. Real-world storage engines sometimes use hybrid or genuinely different structures (some newer designs blend both real ideas), and the read-vs-write trade-off is one real, important axis among several a real storage-engine decision actually involves — not the single, complete answer to every case.
Table — B-Tree vs. LSM Tree — The Real Decision
Real Workload ShapeBetter Real Fit
Read-heavy, or moderate, steady write volumeB-tree
Write-heavy, high-throughput, append-styleLSM tree
Needs the most current data instantly reflected on readB-tree (no compaction lag)
Needs to absorb bursts of writes without degradingLSM tree

This is the same real decision behind GreenMart's own fix — the event log's write-heavy, append-style shape is precisely what an LSM-tree-based system (Cassandra) is built for.

GreenMart closes this real comparison with the precise, physical explanation for its own opening incident — not a vague "NoSQL scales better" but a specific, correct, named reason: two real storage-engine families, each making a genuinely different bet, and GreenMart's own event log simply needed the one built for its actual real shape.

Key Takeaway

B-trees and LSM trees are two real, important families of storage structures, each making a genuinely different bet about when to pay the cost of organizing data — and the right real choice always depends on matching a specific workload's actual shape to the engine actually built for it, not picking a universal "better" option.

Why This Matters

Every future write-heavy or read-heavy real feature GreenMart builds — a new logging system, a new lookup-heavy dashboard — now has a real, physical vocabulary behind the decision of which existing system to build it on, instead of an intuition-based guess.

GreenMart now has the real, complete physical explanation for its own opening incident, and a real framework for every future storage-engine decision. The next chapter shifts from how data is organized on disk to a related but distinct real concern: how data actually gets converted into bytes in the first place — serialization.

Next