Indexes & Query Performance

7.Finding a Product Without Scanning Everything

M

In this chapter

We'll meet MongoDB indexes for real — compound and multikey indexes, index selection, and explain()'s IXSCAN vs. COLLSCAN — the same indexing idea, MongoDB's own syntax.

9–11 min

The Problem in Real Life

GreenMart's catalog isn't small anymore. Thousands of listings now, growing every week as more sellers join. A query that used to return instantly — "find this one listing by its id" — starts taking a visible moment. Nothing broke. The database is just doing exactly what it always did: checking every single document, one at a time, until it finds a match or runs out of documents to check.

Sarah already knows the fix for this shape of problem — an index. The idea carries straight over. The syntax doesn't.

S

Same idea as before. Different everything else.

Sarah

Checking Every Document vs. Checking an Index

A collection scan checks everything

Without an index, MongoDB walks every document in the collection, one at a time, to answer a query.

Compound indexes cover more than one field

One index, several fields, in a specific order — helping queries that filter or sort on that same combination.

Multikey indexes handle arrays automatically

Indexing an array field indexes every element individually, with no extra setup required.

explain() shows what actually happened

IXSCAN means an index was used. COLLSCAN means every document was checked — worth confirming, not assuming.

Indexes & Query Performance

An index in MongoDB does the same job an index does anywhere else: a separate, ordered structure built on one or more fields, letting the database jump straight to matching documents instead of checking every single one. Without an index, a query checking a field forces a collection scan — MongoDB's own term for exactly what it sounds like, walking every document in the collection in turn.

A compound index covers more than one field at once, in a specific order. This one helps any query that filters or sorts by category, or by category and price together — like the "electronics under $50, cheapest first" query from a couple chapters back — but it won't help a query that only filters on price alone, since price isn't the leading field.

A Compound Index — More Than One Field
db.listings.createIndex({ category: 1, price: 1 })

In this simulator, createIndex() only records that the index exists — it's syntax-only, and doesn't actually change how fast a query runs here. Real MongoDB's query planner genuinely uses an index like this to skip a full collection scan.

Indexing a field whose value is an array — like a listing's tags — automatically creates a multikey index: MongoDB indexes every element of the array individually, so a query for one tag inside a long list still gets the benefit of the index, without GreenMart having to do anything different.

A Multikey Index — Indexing an Array Field
db.listings.createIndex({ tags: 1 })

Having an index doesn't guarantee MongoDB actually uses it. Index selection is the query planner's own job — given several indexes and one query, it picks the one it estimates will do the least work, and it's allowed to decide none of them are worth using at all for a given query. explain() is how Sarah checks its actual decision instead of guessing: it reports whether a query ran as an IXSCAN (an index was used) or a COLLSCAN (every document was checked), plus how many documents each stage actually touched.

Query performance isn't really about adding indexes everywhere — every index also costs something on every write, since MongoDB has to keep it updated whenever a matching document changes. The real skill is matching indexes to GreenMart's actual, real access patterns — the same access-pattern-driven thinking from the very start of this course, now applied to a very concrete decision: which fields get indexed, and in what order.

Key Takeaway

An index doesn't speed up every query just by existing — it speeds up the specific queries whose filters and sorts actually match it. explain() is how you check that, instead of assuming it.

Why This Matters

Every chapter left in this Act — replication, sharding — is really about the same underlying question at a bigger scale: how does GreenMart's data stay fast and reliable as it keeps growing? Indexes solve it for a single query on a single machine. The next two chapters solve it across multiple machines.

GreenMart's catalog can now be found quickly, not just correctly. But a fast single copy of the data is still just one copy — and GreenMart already learned, back in Act 1, exactly what that risks.

Next