In this chapter
We'll meet scale, data volume, data velocity, data variety, and geographic distribution — four separate real dimensions of size, applied to GreenMart's own next expansion.
The Problem in Real Life
The promises are named. Mike's next question is about size — not today's size, but the size this has to survive.
Sarah has answered a version of this question in nearly every earlier Act. This time, she answers it for the whole expansion at once, before a single line of it exists.
How big does this actually need to get, and how many countries does it need to work in?
Mike
How Big, How Fast, How Far
Scale is four separate questions, not one
How much data, how fast it arrives, how varied its shape is, and how far it's spread — each stresses a database differently.
Geographic distribution has its own real cost
One country and three continents are genuinely different problems, already answered concretely by Act 4's Global Tables and Act 12's replication.
Scale, Volume & Geographic Requirements
Three real, separate size questions, plus one about shape variety — each with its own real earlier Act already teaching the honest cost of getting it wrong.
- On sheer size. Scale and data volume ask, honestly: how many total records, and how much total traffic, is this actually expected to reach — not today, but at the size GreenMart is planning for. Act 4's DynamoDB and Act 5's Cassandra were both built specifically for scale most single-machine databases never need to reach.
- On speed of arrival. Data velocity asks how fast new data actually arrives — a handful of orders a minute is a very different problem from thousands of fleet-location pings a second, Act 8's own time-series territory, where high-frequency writes were the entire opening problem.
- On shape diversity. Data variety asks whether the incoming data is uniform (every order looks like every other order) or genuinely varied (some orders have three items, some have thirty, some come with special handling notes and some don't) — the same flexible-schema question Act 2's MongoDB opened this whole course with.
- On where it all has to work. Geographic distribution asks how spread out GreenMart's actual users and warehouses are — one country is a very different problem from three continents, and Act 4's Global Tables and Act 12's replication topologies both exist specifically to answer it.
For GreenMart's fleet-and-order requirement: real scale (millions of orders across the expansion, not thousands), high velocity on fleet-location updates specifically, moderate variety (orders vary somewhat, but stay within a predictable shape), and genuine geographic distribution across every country GreenMart operates in. Four more honest facts, on top of the ones already gathered — and still no database named.
Key Takeaway
Scale isn't one number — it's four separate questions (how much, how fast, how varied, how far), and a database that's genuinely strong on one of them can still be the wrong choice if the real workload's actual pressure is on a different one.
Why This Matters
Choosing a database that scales beautifully in the one dimension a workload doesn't actually stress, while quietly struggling in the dimension it does, is one of the most common real-world mismatches this kind of premature choice produces. Naming all four dimensions honestly is what prevents it.
GreenMart now has a third real category of facts: how big, how fast-arriving, how varied, and how geographically spread the fleet-and-order requirement actually is. The next chapter turns to a different kind of question entirely — who actually has to run this, day to day, and what it costs.
