Cardinality & Observability at Scale

5.Too Many Things to Watch

M

In this chapter

We'll meet cardinality and high-cardinality problems — a real, verified demonstration of a tag that fragments a metric into millions of nearly-empty series — the same hot-key/hot-partition lesson this course keeps re-teaching, and see where this whole toolkit shows up in the real world: monitoring, IoT, financial data, and observability.

9–11 min

The Problem in Real Life

Sarah adds one more tag to the server-monitoring metric, thinking it'll help: requestId, a unique id per request, so any single slow request can be looked up directly. Within a day, the metric has millions of distinct tag combinations — one for essentially every request that ever happened.

Dashboards that used to load instantly start crawling. Queries that used to summarize "today's average response time" in one number now have millions of nearly-empty series to wade through instead.

S

I added one tag. Why did everything just get slower?

Sarah

A Few Meaningful Series vs. Millions of Nearly-Empty Ones

Cardinality counts distinct tag combinations

A tag whose values repeat across many points (truck id) stays healthy — one that's nearly unique per point doesn't.

High cardinality fragments a metric

Verified: filtering by a unique-per-point tag always returns exactly one point — nothing left to meaningfully group or aggregate.

The same lesson, a new shape

Redis's hot keys, DynamoDB's/Cassandra's hot partitions, now a metric fragmented into millions of nearly-empty series.

One real toolkit, many real domains

Monitoring, observability, IoT, and financial data all share the exact same underlying shape this Act has built.

Cardinality & Observability at Scale

Cardinality is the number of distinct tag-value combinations a metric actually has. truck.speed tagged only by truck (a few dozen real trucks) has low, healthy cardinality — a few dozen genuinely meaningful series, each with plenty of points, each worth aggregating and charting. Add a tag like requestId, unique to every single request, and cardinality explodes: essentially one "series" per point, ever.

Filtering by one specific requestId value returns exactly one point — always, by definition, since no two requests share that tag. A tag this unique can never usefully group or aggregate anything; every "series" it creates has, at most, a single point in it.

A High-Cardinality Tag, Demonstrated
INSERT server.latency 120 requestId=req-8f3a2 at +0m
INSERT server.latency 135 requestId=req-9c1b7 at +1m
INSERT server.latency 118 requestId=req-2e4d9 at +2m
QUERY server.latency WHERE requestId=req-8f3a2
QUERY server.latency

Compare this to truck=T1 from earlier chapters — a tag shared by many real points over time is exactly what makes grouping and rollups meaningful in the first place.

This is a high-cardinality problem, and it's the same underlying shape as a lesson this course keeps re-teaching in new places: Redis's hot keys, DynamoDB's and Cassandra's hot partitions, and now a metric with too many distinct tag combinations — a design choice that looked harmless in isolation ("just add one more tag") becoming a real, systemic cost once enough data exists. The fix is the same kind of discipline as every earlier version of this lesson: keep tags to values that repeat meaningfully across many points (truck id, region, server name), and keep genuinely unique identifiers (a specific request id, a specific order id) out of tags entirely — logged or stored elsewhere, not used to fragment a time-series metric into millions of one-point series.

Key Takeaway

Cardinality is healthy when a tag's values repeat across many real points, and dangerous the moment a tag becomes nearly unique per point — the same "one value gets disproportionately, structurally overloaded" shape as a hot key or hot partition, here showing up as millions of nearly-empty series instead of one overloaded one.

Why This Matters

This is the one mistake in this Act that's genuinely easy to make by accident — adding "just one more tag" feels harmless until cardinality has already exploded. Recognizing it here means GreenMart designs tags deliberately from the start, the same discipline every earlier Act's own version of this lesson has already built.

GreenMart now knows why a tag needs to repeat across real points to be useful, and can name the failure mode when it doesn't. What this whole toolkit is actually used for, beyond GreenMart's own dashboards, is worth naming directly before this Act closes.

Next