In this chapter
We'll see a real rollup correctly update once a late-arriving, out-of-order point lands inside an already-reported time window, and meet continuous aggregation as the real production technique for keeping rollups correct as data keeps arriving.
The Problem in Real Life
A sensor on one of GreenMart's delivery trucks loses signal for a few minutes, then reconnects and sends its queued readings all at once. One of them is timestamped before readings the dashboard has already displayed and already rolled up into an hourly average.
Sarah's first instinct is that this reading is just... late. Does it even count anymore?
The reading's real. It's just showing up after everything that came after it.
Sarah
A Reading That Arrives in Order vs. One That Doesn't
Late data is normal, not an error
Network delays, offline buffering, batch catch-ups — real, common reasons a point's timestamp is earlier than data already stored.
A rollup can be recomputed correctly
Verified: an already-reported window's average genuinely updates once a late point lands inside it.
Continuous aggregation keeps rollups current
A real production technique for updating rollups incrementally as data arrives, rather than only computing them once.
Getting this wrong is silently wrong
A rollup computed before a late point arrives, and never revisited, stays incorrect forever — not just momentarily behind.
Out-of-Order Data & Continuous Aggregation
Late-arriving data — and the more general case, out-of-order data — is exactly this: a point whose timestamp is earlier than points the database already has, but which is only being written now. Network delays, a sensor buffering readings while offline, a batch job catching up — all real, common reasons this happens constantly in production, not an edge case worth ignoring.
The first rollup query reports the 0-10m window's average as 20.00 (one point). Then a reading timestamped +5m — inside that same already-reported window — arrives late, after the +10m reading that came "before" it in write order. Run the identical rollup query again, and the 0-10m window's average correctly updates to 21.00, now genuinely including the late point.
INSERT sensor.temp 20 at +0mINSERT sensor.temp 25 at +10mQUERY sensor.temp FROM +0m TO +20m GROUP BY 10m AGG(avg)INSERT sensor.temp 22 at +5mQUERY sensor.temp FROM +0m TO +20m GROUP BY 10m AGG(avg)
This is real, verified behavior — the late point wasn't lost, and the rollup wasn't stuck at its first answer. Run this yourself and watch the average actually change between the two identical queries.
This is what continuous aggregation is really solving for in a real production system: keeping a rollup correct as data continues arriving, not just correct once at the moment it was first computed. A real system typically does this incrementally — updating an already-computed rollup as new points land, rather than recalculating the whole window from scratch every time — a genuine performance technique this course's simulator doesn't literally implement (it recomputes fresh on every query, since that's simple and fast enough at this scale). The correctness behavior demonstrated above is real and honest; the incremental-update mechanism itself is a production detail worth knowing exists, not something claimed here as physically modeled.
The deeper point worth sitting with: a time-series database that couldn't handle a late point gracefully would be quietly wrong, not just inconvenient — every rollup computed before that point arrived would stay wrong forever. Treating out-of-order data as a normal, expected case (not a special error condition) is part of what "time is the organizing dimension" actually requires in practice.
Key Takeaway
A late-arriving point isn't an error to reject — it's a normal, expected case a real time-series database has to fold correctly into rollups that may have already been computed. Getting this wrong doesn't just miss one reading; it leaves every summary built before that point arrived silently incorrect.
Why This Matters
GreenMart's fleet sensors, and every real-world source of time-series data, will eventually deliver something late — a network hiccup, a queued batch, a device reconnecting. A dashboard that can't handle that gracefully quietly reports wrong numbers exactly when GreenMart is trusting it most.
GreenMart now knows a late reading doesn't get lost or ignored — a rollup correctly includes it once it arrives. None of this Act has addressed what happens once there are too many different tag combinations to track sanely, though — exactly where the next chapter goes.
