Embedding vs. Referencing

2.Inside the Box or Down the Hall?

M

In this chapter

We'll meet MongoDB's central design question — should this data live inside the document, or somewhere else — through a real embedded-array failure and its fix.

9–11 min

The Problem in Real Life

Sarah adds the seller's name and rating directly into every listing document — the same call she already made once, for the same reason: GreenMart's storefront shows it on every single page view, and a seller barely ever touches their own profile. In MongoDB's own words, she's embedding it.

The fast charger listing takes off a few months later, after a well-known local reviewer mentions it. Reviews pour in. Sarah had embedded those too — each new review just pushed onto an array field inside the same listing document, the same instinct that worked fine for the seller's name. One evening, a new review fails to save.

S

It didn't error on the review. It errored on the whole document.

Sarah

Embed vs. Reference

Embedding works — until it doesn't

The seller's name embeds fine forever. An unbounded array of reviews embeds fine, right up until it doesn't.

A real, hard limit: 16MB per document

Almost nothing gets close to it — except data that grows without a natural stopping point.

References aren't enforced

Unlike a SQL foreign key, MongoDB never checks that a reference still points at something real — that's on the application.

The question that never gets easier

"Should this data live inside this document, or somewhere else?" — asked fresh, every time, not answered once for the whole database.

Embedding vs. Referencing

Every MongoDB document has a hard limit: 16 megabytes. Almost no document ever gets close to it — a listing with a name, a price, a few specs is a few hundred bytes at most. But an array that grows without any natural stopping point, one push at a time, forever, doesn't respect that limit just because nobody happened to be watching. A wildly popular listing with thousands of reviews, each one appended to the same document, is exactly the kind of growth that eventually gets there.

  • Embed when the data is read together constantly, stays small and naturally bounded, and doesn't need to be queried or updated on its own — like a seller's name and rating riding along inside every listing.
  • Reference when the data can grow without a natural limit, is shared across many parent documents, or genuinely needs to be queried or updated independently — like a listing's reviews.

A review now lives in its own document, in its own reviews collection, holding a plain listingId field that points back to the listing it belongs to. That field is what MongoDB calls a reference — nothing more than an ordinary value, not a special type.

Reviews, Moved to Their Own Collection
db.reviews.insertOne({
_id: "rev-8841",
listingId: "lst-fast-charger-20w",
rating: 5,
text: "Charges my phone faster than the one it replaced.",
reviewer: "priya92"
})

Worth being honest about: unlike a foreign key in a relational database, MongoDB never automatically checks that listingId still points at a real listing. Keeping references valid is the application's job, not the database's.

This is the Act's real, recurring question, and it never gets easier to answer by instinct alone: should this data live inside this document, or somewhere else? Embedding and referencing aren't a beginner move versus an advanced one — they're two honest answers to two genuinely different situations, and the fast charger's reviews just showed what happens when the wrong one gets picked without checking whether the data's actual growth pattern was ever bounded in the first place.

Notice what didn't change: the seller's name and rating are still embedded, and that's still the right call. Nothing about the reviews problem makes embedding wrong in general — it makes embedding wrong specifically for data that can grow without limit. The two decisions live side by side in the same listing document, each one made for its own reason.

Key Takeaway

Should this data live inside this document, or somewhere else? Embed when it's small, bounded, and always read together. Reference the moment it can grow without limit.

Why This Matters

This exact question — inside the document, or somewhere else — comes back in every chapter for the rest of this Act. Querying across a reference (next chapter), aggregating data that's split across collections, and even how a document gets sharded later all trace back to this one design decision, made document by document.

GreenMart's listings now hold what belongs with them and reference what doesn't. The next question is more practical: now that the shape is right, how does Sarah actually read and write this data day to day?

Next