In this chapter
We'll see why cramming every fact into one table causes real bugs — and split GreenMart's data into proper entities: Customer, Product, and Sale.
The Problem in Real Life
Sarah starts building GreenMart's first real table the obvious way: one wide table, every fact she can think of, all in one place — Customer Name, Phone, Product, Price, Quantity, Date. It works fine, for about ten rows.
Then Mike mentions that Alex moved apartments and has a new phone number. Sarah goes to update it, and finds Alex's old number sitting in three different rows — one for every single thing Alex has ever bought.
| Customer | Phone | Product | Price | Date |
|---|---|---|---|---|
| Alex | 555-0198 | Apples | $4 | Aug 2 |
| Alex | 555-0198 | Bread | $3 | Aug 9 |
| Alex | 555-0198 | Milk | $3 | Aug 16 |
Alex's phone number is typed three separate times — once per purchase. Update it in one row, forget the other two, and GreenMart now has two different "correct" phone numbers for the same person, and no way to tell which one is real.
Why am I storing Alex's phone number three times?
Sarah
One Big Table vs. Real Entities
Facts got duplicated
Alex's phone number gets typed in again every single time she buys something.
Updates don't stick everywhere
Change the phone number in one row, and the other two still show the old one.
Two different things, one row
Facts about a customer and facts about a sale are jammed together, so neither is fully correct on its own.
Simple questions get harder
"How many customers do we actually have?" now means untangling duplicate rows to count unique people.
What Are Entities and Attributes, Really?
That one wide table feels natural, because it matches how the notebook worked — one line, every fact about that moment, all crammed together. But look closely at what's actually happening on each row: some of those facts are about Alex, the person. Other facts are about one specific purchase Alex happened to make. Those are two completely different kinds of things, sharing one row by accident.
Every time Alex buys something new, her name and phone number get typed in all over again — not because anyone made a mistake, but because the table itself has no idea these are two different kinds of things. Nothing in a wide table like this distinguishes "a fact about Alex" from "a fact about this one sale." It just sees columns.
This is exactly the judgment call database design is really about — and it comes with its own vocabulary, because it's important enough to name precisely:
- An entity is one distinct kind of thing a business needs to keep track of — for GreenMart, that's a Customer, a Product, and a Sale.
- An attribute is a single fact about one entity — a Customer's name and phone are attributes of the Customer entity. A Sale's date and amount are attributes of the Sale entity.
- Each entity gets its own table. Deciding which facts belong to which entity, before a single column gets created, is most of what database design actually is.
| Customer | Phone |
|---|---|
| Alex | 555-0198 |
Alex's phone number now lives in exactly one place. Update it here, once, and it's correct everywhere it matters — because nowhere else stores it.
| Product | Price | In Stock |
|---|---|---|
| Apples | $4 | 40 |
| Bread | $3 | 15 |
| Milk | $3 | 22 |
Product info — price, how many are left — is its own entity too, with its own table, exactly like Customer.
| Customer | Product | Price | Date |
|---|---|---|---|
| Alex | Apples | $4 | Aug 2 |
| Alex | Bread | $3 | Aug 9 |
| Alex | Milk | $3 | Aug 16 |
This table just says which customer bought which product, by name for now. That "by name" is doing more work than it should — two customers can easily share a name, and this exact problem is where the next chapter picks up.
Splitting one big table into three smaller ones can feel like more work, not less — more tables, more things to keep track of. But look at what it actually buys GreenMart: Alex's phone number lives in exactly one row, in one table. Change it there, once, and every part of the business that ever needs it reads the same correct value automatically, because there's only ever one copy to be wrong.
This is the real skill this chapter is training — not "how do I create a table," but "which facts actually belong together, and which don't." Every table this course builds from here on is really just this same question, asked about a different part of GreenMart.
Key Takeaway
Group facts by what they're actually about — one entity, one table — and duplication, along with the bugs it causes, disappears on its own.
Why This Matters
This chapter's real lesson — figure out which facts belong to which thing, before building anything — is the seed of nearly everything ahead. The ER diagrams in chapter 8 are a formal drawing of exactly this decision. The normalization chapter much later in the course (Act 4) is this same instinct, applied rigorously to a schema that's already been through real use. Even the next chapter's problem — telling two customers named Alex apart — only exists because this chapter correctly gave Customer its own table in the first place.
Sarah has split GreenMart's mess into three clean entities: Customer, Product, Sale. But the Sales table currently points at a customer by typing out their name — and names aren't as unique as they feel. The next chapter is about exactly what happens when two customers share one.
