In this chapter
We'll build a real aggregation pipeline — $match, $group, $sort, $limit — turning raw listings into an actual seller dashboard.
The Problem in Real Life
Sarah needs a real seller dashboard now: how many electronics listings does each seller have, and what's their average price? find() can hand back every matching document — but nobody wants to scroll through hundreds of listings and do the counting and averaging by hand.
Sarah needs the database to actually combine documents into an answer, not just filter them.
I don't want the listings. I want what they add up to.
Sarah
A Single Query vs. A Pipeline of Stages
A pipeline, not a single query
Each stage takes the previous stage's output and transforms it further — filter, then group, then sort, then limit.
$group turns documents into summaries
Many documents sharing an _id expression collapse into one summary document per group.
Accumulators compute across a group
$sum, $avg, $min, $max, $count — real numbers computed across every document in a group, not just one.
The output shape can be brand new
A pipeline's result doesn't have to look anything like the documents it started from.
The Aggregation Pipeline
This is what MongoDB's aggregation pipeline is for. Unlike find(), which filters and shapes documents but always hands back documents, an aggregation runs a sequence of stages — each one taking the previous stage's output and transforming it further — until what's left isn't matching documents anymore. It's an actual computed answer.
Before any code, watch what actually happens to the data, stage by stage, using four real listings.
| Seller | Category | Price |
|---|---|---|
| sel-01 | electronics | $8.99 |
| sel-01 | electronics | $24.99 |
| sel-02 | electronics | $12.00 |
| sel-01 | produce | $1.99 |
| Seller | Category | Price |
|---|---|---|
| sel-01 | electronics | $8.99 |
| sel-01 | electronics | $24.99 |
| sel-02 | electronics | $12.00 |
Predict before scrolling: sel-01 has two electronics listings here ($8.99 and $24.99), sel-02 has one ($12.00). If the next stage groups by seller and computes a count and an average price, what would it produce for each one?
| _id (seller) | listingCount | averagePrice |
|---|---|---|
| sel-01 | 2 | $16.99 |
| sel-02 | 1 | $12.00 |
Three documents became two — one summary per seller. sel-01's two prices, $8.99 and $24.99, collapsed into a single averagePrice of $16.99. This is the actual, computed result $group produces, not an estimate.
A pipeline is an array of stages, run in order. $match usually goes first — it filters down to the documents worth doing any real work on, exactly like a find() filter. This is exactly what produced the second table above.
db.listings.aggregate([{ $match: { category: "electronics" } }])
$group collapses many documents sharing the same _id expression — here, $sellerId — into one summary document per group, exactly the transformation the third table above walked through. $sum: 1 counts how many documents landed in each group; $avg: "$price" averages a field across that same group. MongoDB also supports $min, $max, and $count as the same kind of accumulator.
db.listings.aggregate([{ $match: { category: "electronics" } },{ $group: {_id: "$sellerId",listingCount: { $sum: 1 },averagePrice: { $avg: "$price" }} }])
The same $sort and $limit from a find() query work here too — as stages, not as chained methods. Added onto the pipeline above, this becomes "the top 3 sellers by electronics listing count," a real, specific business question, answered in one call.
db.listings.aggregate([{ $match: { category: "electronics" } },{ $group: {_id: "$sellerId",listingCount: { $sum: 1 },averagePrice: { $avg: "$price" }} },{ $sort: { listingCount: -1 } },{ $limit: 3 }])
$count (how many documents landed in the group), $min, and $max round out $sum/$avg as the full set of accumulators. $project, used as a pipeline stage rather than find()'s second argument, trims the final output down to just the fields a report actually needs — here, dropping cheapest/priciest from what actually gets returned, without ever removing them from the underlying documents.
db.listings.aggregate([{ $group: {_id: "$category",totalListings: { $count: {} },cheapest: { $min: "$price" },priciest: { $max: "$price" },averagePrice: { $avg: "$price" }} },{ $project: { _id: 1, totalListings: 1, averagePrice: 1 } }])
Notice what changed across those three stages: the input was individual listing documents, but the output was seller summaries — a shape that never existed in the raw data at all. That's the actual point of an aggregation pipeline. A $project stage (reshaping which fields survive into the next stage, the same idea as find()'s projection but usable mid-pipeline) rounds out the toolkit for building exactly the shape a dashboard or report actually needs.
This only works because GreenMart's listings already carry what a summary like this needs — a sellerId on every listing, price as an actual number, not a formatted string. Aggregation can transform data, but it can't invent a relationship that was never captured on the documents in the first place.
Key Takeaway
An aggregation isn't one query — it's a pipeline of stages, each transforming the previous stage's output, until documents become the actual computed answer.
Why This Matters
Every report, dashboard, or summary GreenMart ever needs from MongoDB runs through this same pipeline shape — filter, group, shape, sort, limit, in whatever combination the actual question needs. It's also the direct MongoDB answer to something SQL handled with GROUP BY and HAVING — a genuinely different syntax, but the same underlying job: turn many rows (or documents) into fewer, more meaningful ones.
GreenMart's catalog can now answer real questions, not just return matching listings. But every one of these queries has been running against however many documents GreenMart happens to have today — and that number is about to stop being small.
