Access-Pattern-Driven Design

3.Design for the Question, Not the Table

M

In this chapter

We'll see why NoSQL design starts from the actual questions an app will ask, not from the entities alone — and why deliberately duplicating data is sometimes the right call.

8–10 min

The Problem in Real Life

Sarah starts designing the new listing record the way GreenMart's database has always been designed: find the entities, keep each one in its own place, don't repeat data. A Listing record. A separate Seller record, holding the seller's name, rating, and contact details. A listing just points at its seller by ID — the same instinct that made Products, Customers, and Sales three separate tables in the first place.

Then she actually looks at how the storefront will use it. Every single listing page — and there will be thousands of page views a day — needs the seller's name and rating shown right there, immediately. With Seller kept separate, that's two lookups every time: fetch the listing, then fetch its seller. Sarah checks how often a seller's profile actually changes. A few times a year, if that.

S

I split this the SQL way. But nobody reads a listing that way.

Sarah

Entities First vs. Questions First

Start from the question, not the entity

Not "what tables do I need" but "what will the app actually ask, and how often?"

Denormalization is a decision, not a mistake

Deliberately duplicating data — like a seller's name inside every one of their listings — trades a rarer, costlier write for a much more frequent, much cheaper read.

Joins get expensive

Most NoSQL databases don't do cheap joins across records the way SQL does — a design that quietly depends on one gets slow, page after page.

Read-heavy vs. write-heavy

A listing gets viewed thousands of times for every one time it's edited — design for the ratio that's actually true, not a general rule.

Access-Pattern-Driven Design

GreenMart's relational database has always started from the data: what are the real-world things (GreenMart's entities), and how do they relate? Normalize first, keep every fact in exactly one place, ask questions later. That approach works because a relational database makes joining those separate pieces back together, at query time, genuinely cheap.

Most NoSQL databases don't make that promise. Joining across separate records — especially once data is spread across many machines, which several later Acts get into — is often slow, sometimes not supported at all. So the design question flips. Instead of "what are the entities, then what can I ask," NoSQL design starts with: "what will the application actually ask, and how often — then what shape makes that fast?" This is called access-pattern-driven design.

  • "Show me this one listing" — the single most common read on the whole storefront, happening thousands of times a day
  • "Show me every listing from this seller" — a seller's own dashboard, much rarer
  • "Update a seller's name or rating" — rare, maybe a few times a year per seller
Table — Listing Page Load — Normalized vs. Denormalized
ApproachWhat a listing page has to doLookups per page view
Normalized — seller kept separateFetch the listing, then fetch its seller record to get the name and rating2
Denormalized — seller's name & rating copied into the listingFetch the listing — the seller's name and rating are already sitting right there1

Same information either way. The only difference is where the work happens: every read, or once at write time.

Once Sarah sees the numbers that way, the fix is obvious: copy the seller's name and rating directly into every one of their listing documents. This is called denormalization — deliberately duplicating data instead of keeping one single source of truth in one place. It sounds like exactly the mistake Sarah spent a whole Act learning to avoid. Here, it's the opposite: a deliberate trade, made on purpose, because the numbers say it's the right one.

The trade is real, not free. If a seller changes their display name, every listing document holding a copy now needs updating too, not just one Seller record. But that update happens a handful of times a year, against a read that happens thousands of times a day. Optimizing for the thing that happens thousands of times more often than the other is exactly what read-heavy vs. write-heavy thinking means — look at which one is actually true for this specific piece of data, and design for that, not for a general rule that data should never be duplicated.

This one idea — start from the question, not the entity — is the thread running through almost every NoSQL Act still ahead. Document databases call it embedding vs. referencing. DynamoDB calls it single-table design. Cassandra calls it query-driven modelling. Different names, same underlying move: shape the data around how it's actually going to be read, and accept some duplication as the cost of making that read cheap.

Key Takeaway

In a relational database, you design around what the data IS. In most NoSQL databases, you design around what you'll actually ASK — and that often means deliberately duplicating data.

Why This Matters

This isn't a one-time trick for the seller-listing problem. It's the design principle every NoSQL Act from here builds on, under a different name each time. Understanding it now, with one concrete, low-stakes example, means the same idea won't have to be re-taught from scratch when it resurfaces as "single-table design" three Acts from now.

Sarah's listing documents are shaped around how they'll be read. But a single copy of a document, sitting on a single machine, is one failure away from disappearing entirely — and GreenMart can't risk that with real seller data. The next chapter is where copies of the data start to matter.

Next