In this chapter
We'll see why NoSQL design starts from the actual questions an app will ask, not from the entities alone — and why deliberately duplicating data is sometimes the right call.
The Problem in Real Life
Sarah starts designing the new listing record the way GreenMart's database has always been designed: find the entities, keep each one in its own place, don't repeat data. A Listing record. A separate Seller record, holding the seller's name, rating, and contact details. A listing just points at its seller by ID — the same instinct that made Products, Customers, and Sales three separate tables in the first place.
Then she actually looks at how the storefront will use it. Every single listing page — and there will be thousands of page views a day — needs the seller's name and rating shown right there, immediately. With Seller kept separate, that's two lookups every time: fetch the listing, then fetch its seller. Sarah checks how often a seller's profile actually changes. A few times a year, if that.
I split this the SQL way. But nobody reads a listing that way.
Sarah
Entities First vs. Questions First
Start from the question, not the entity
Not "what tables do I need" but "what will the app actually ask, and how often?"
Denormalization is a decision, not a mistake
Deliberately duplicating data — like a seller's name inside every one of their listings — trades a rarer, costlier write for a much more frequent, much cheaper read.
Joins get expensive
Most NoSQL databases don't do cheap joins across records the way SQL does — a design that quietly depends on one gets slow, page after page.
Read-heavy vs. write-heavy
A listing gets viewed thousands of times for every one time it's edited — design for the ratio that's actually true, not a general rule.
Access-Pattern-Driven Design
GreenMart's relational database has always started from the data: what are the real-world things (GreenMart's entities), and how do they relate? Normalize first, keep every fact in exactly one place, ask questions later. That approach works because a relational database makes joining those separate pieces back together, at query time, genuinely cheap.
Most NoSQL databases don't make that promise. Joining across separate records — especially once data is spread across many machines, which several later Acts get into — is often slow, sometimes not supported at all. So the design question flips. Instead of "what are the entities, then what can I ask," NoSQL design starts with: "what will the application actually ask, and how often — then what shape makes that fast?" This is called access-pattern-driven design.
- "Show me this one listing" — the single most common read on the whole storefront, happening thousands of times a day
- "Show me every listing from this seller" — a seller's own dashboard, much rarer
- "Update a seller's name or rating" — rare, maybe a few times a year per seller
| Approach | What a listing page has to do | Lookups per page view |
|---|---|---|
| Normalized — seller kept separate | Fetch the listing, then fetch its seller record to get the name and rating | 2 |
| Denormalized — seller's name & rating copied into the listing | Fetch the listing — the seller's name and rating are already sitting right there | 1 |
Same information either way. The only difference is where the work happens: every read, or once at write time.
Once Sarah sees the numbers that way, the fix is obvious: copy the seller's name and rating directly into every one of their listing documents. This is called denormalization — deliberately duplicating data instead of keeping one single source of truth in one place. It sounds like exactly the mistake Sarah spent a whole Act learning to avoid. Here, it's the opposite: a deliberate trade, made on purpose, because the numbers say it's the right one.
The trade is real, not free. If a seller changes their display name, every listing document holding a copy now needs updating too, not just one Seller record. But that update happens a handful of times a year, against a read that happens thousands of times a day. Optimizing for the thing that happens thousands of times more often than the other is exactly what read-heavy vs. write-heavy thinking means — look at which one is actually true for this specific piece of data, and design for that, not for a general rule that data should never be duplicated.
This one idea — start from the question, not the entity — is the thread running through almost every NoSQL Act still ahead. Document databases call it embedding vs. referencing. DynamoDB calls it single-table design. Cassandra calls it query-driven modelling. Different names, same underlying move: shape the data around how it's actually going to be read, and accept some duplication as the cost of making that read cheap.
Key Takeaway
In a relational database, you design around what the data IS. In most NoSQL databases, you design around what you'll actually ASK — and that often means deliberately duplicating data.
Why This Matters
This isn't a one-time trick for the seller-listing problem. It's the design principle every NoSQL Act from here builds on, under a different name each time. Understanding it now, with one concrete, low-stakes example, means the same idea won't have to be re-taught from scratch when it resurfaces as "single-table design" three Acts from now.
Sarah's listing documents are shaped around how they'll be read. But a single copy of a document, sitting on a single machine, is one failure away from disappearing entirely — and GreenMart can't risk that with real seller data. The next chapter is where copies of the data start to matter.
