#01The 120ms API That Still Feels Slow
Picture the on-call engineer's favorite kind of ticket: a customer says search "feels broken," support escalates it, and the first thing anyone does is pull up the API dashboard. The /search endpoint's p99 latency is 118ms. The p50 is 40ms. Error rate is flat at zero. Every graph a backend team has ever been taught to care about is green. By every number the API team owns, search is fast — arguably faster than it's ever been, after a quarter spent trimming a slow database query down to almost nothing.
Make the scenario concrete: an e-commerce storefront, a shopper typing "waterproof hiking boots" into the search bar on their phone during a commute. The /search API genuinely does its job in 118ms. But that same search box also fires an autocomplete call, a personalization call, and an inventory-check call — three more round trips nobody thought to mention, because nobody thinks of them as "the search API." Add a page still finishing its own initial load, competing for the same connections and the same JavaScript thread, and four seconds stops being mysterious. It was never one slow thing. It was five moderately fast things nobody had ever measured together.
And yet the customer is right. Open the same search box, type the same query, and the results don't appear for over four seconds. Nothing in that four seconds is a lie the API dashboard is telling — the API genuinely did respond in 118ms. It's also almost irrelevant to what the person sitting in front of the screen actually experienced, because 118ms is the answer to a much narrower question than the one they'd ask: not "was the server fast," but "how long from the moment I hit enter until I could read a result." Those are different questions with different answers, and the entire rest of this piece is about the distance between them.
This isn't a story about a badly built API. It's a story about a well-built API sitting inside a system nobody measured end to end — a far more common, far less obvious failure than a slow query. A slow endpoint shows up immediately, in the one dashboard everyone already watches. A slow system with a fast API in the middle of it can hide for months, because every individual metric keeps saying everything is fine.
The next reasonable step the on-call engineer takes is usually to reproduce it themselves — and here the mystery gets stranger, not clearer. On the office network, on a fast laptop, search feels instant. The four-second complaint only reproduces on a customer's actual phone, on an actual mobile connection, loading an actual page that also has to fetch a stylesheet, three fonts, an analytics script, and a chat widget before the search results the API already delivered can actually appear on screen. Nothing about the API changed between those two attempts. Almost everything else about the journey did.
That gap between "reproduces instantly at my desk" and "takes four seconds on a real device in the field" is not a fluke of one customer's bad wifi. It's the predictable result of measuring a system from the one vantage point — a fast machine, a fast connection, a warm cache — that happens to be the least representative of how most real users actually experience it. A dashboard built entirely from server-side numbers can't see that gap at all, because the server-side number is identical in both cases. Only something watching the whole journey, on real devices, would ever notice the difference.
A fast API answering the wrong question is indistinguishable, from a dashboard, from a fast API answering the right one — until a real user tells you otherwise.
#02Your API Isn't the Whole Request
The 118ms in that dashboard almost always measures one specific slice of time: from the moment a request lands on the application server to the moment that server hands a response back to whatever's in front of it. That's a perfectly reasonable thing to measure — it's the part of the system the backend team actually built and can actually change. It is not, however, the part of the round trip the user is waiting on, which starts well before the server ever sees a byte and doesn't end when the server sends one back.
Before "the API" gets involved at all, a browser has to resolve a domain name via DNSDomain Name System — the distributed lookup service that translates human-readable domain names into the IP addresses computers actually route traffic to., open a TCPTransmission Control Protocol — adds reliable, ordered delivery and congestion control on top of the internet's inherently unreliable, unordered raw packet delivery. connection to whatever address that resolves to, and — for anything running over HTTPS, which is effectively everything now — negotiate a TLSTransport Layer Security — the cryptographic protocol that encrypts, authenticates, and verifies the integrity of data in transit between a client and a server. handshake on top of that connection before a single application byte can move in either direction. Each of those is its own round trip: one for DNS, one for the TCP three-way handshake (a SYN, a SYN-ACK, and an ACK before either side has sent real data), and typically one more for TLS 1.3's handshake. On a cold connection, that's two to three full round trips before the request that actually matters can even be sent — and on a mobile network, where one round trip alone can run into the hundreds of milliseconds depending on signal quality, that setup cost can rival or exceed the API's own processing time.
Connection reuse is what keeps this from happening on every single request: a browser that keeps a connection alive across multiple calls to the same host, via HTTP keep-alive or HTTP/2's multiplexed streams, pays that DNS-plus-TCP-plus-TLS cost once, not per request. Only the first request to a new host pays it in full — which is exactly why a performance test run entirely on warm, already-open connections will systematically understate how slow a real user's first request of a session actually is.
Underneath all of that sits a limit no amount of engineering removes: the physical distance between a user and the server, and how long it takes a signal to cross it. A user in Mumbai talking to a server in Virginia can already spend a significant chunk of round-trip time on nothing but that distance, before any handshake or application logic gets a chance to run — real transit time, not anything a backend team wrote. This is the actual reason a CDNContent Delivery Network — a geographically distributed set of edge servers that cache content close to visitors, shrinking the network distance a request has to travel. (content delivery network) exists: not to make a server faster, but to move a copy of what it would have answered physically closer to the user, shrinking that unavoidable round trip from "halfway around the world" to "the nearest edge location" — a fix for geography that no amount of query optimization was ever going to provide, because query optimization only starts after the request has already finished traveling.
Then the request actually reaches the server, and the 118ms clock starts — but even that number is doing more than "run one query." A typical API handler authenticates the caller, maybe checks a rate limit, executes one or more database calls, possibly serializes a response through several layers of framework middleware, and only then hands bytes back to the network. Each of those is real, billable time inside the number the dashboard reports as "the API" — which means even the part everyone agrees to measure is already an aggregate, not a single atomic operation.
And the response leaving the server isn't the end of the user's wait either. Those bytes have to travel back across the same network, arrive at the browser, get parsed, and get turned into something a human can actually look at — a step this piece comes back to properly in a few chapters, because it turns out to be far from free. "The API responded" and "the user has something usable on screen" are two different moments in time, separated by work that no API latency metric was ever built to see.
Line up every one of those hops and the API dashboard instruments exactly one segment in the middle of the list — everything before and after it is real time on the user's clock that dashboard was simply never pointed at.
"API latency" was always measuring one segment of a much longer trip — the mistake is treating that segment as if it were the whole journey.
#03Latency Is a Journey, Not a Number
Once you accept that a request is a chain of steps rather than a single event, a single latency number stops being able to tell the whole story on its own — because a single number, by definition, can't distinguish "one moderately slow thing happened" from "five totally normal things happened back to back." A p99 of 400ms could mean one occasionally-slow database call. It could just as easily mean five services, each individually well within its own SLA, chained together in a way nobody's dashboard represents as a single unit.
This gets sharper once you stop thinking in averages and start thinking in probabilities. If a request depends on five sequential steps and each one independently has a 2% chance of landing in its own slow tail, the request's chance of hitting at least one slow step is meaningfully higher than 2% — each step is 98% likely to be fast, so all five being fast at once is only about 90% likely (0.98 raised to the fifth power), leaving roughly a one-in-ten chance something along the way is slow, even though every step looks 98% healthy on its own dashboard. This is what distributed-systems engineers call tail latency amplification, and it only gets worse as an architecture grows: the same math pushes past a one-in-five chance at ten dependencies, and below 70% at twenty — with not one individual service having gotten any less reliable. A system doesn't need one consistently slow component to feel unreliable; it just needs enough independently-healthy components chained on the same request.
That reframes what "measuring latency" should even mean. A metric that reports the median or the 99th percentile of one component, in isolation, is answering "how does this one piece behave on its own." It is structurally unable to answer "how does the chain this piece sits inside behave as a whole," because that question depends on how the pieces are arranged relative to each other — which ones happen one after another, and which ones could, in principle, happen at the same time. That distinction is the entire subject of the next chapter, and it's the single biggest lever most systems have never pulled.
A latency number describes one link in a chain — it was never going to describe the chain, because the chain is a property of the arrangement, not any single link.
#04The Hidden Enemy: Waterfalls
Here is the arrangement that quietly does the most damage, and it's one every engineer has drawn without necessarily naming it: a strictly sequential chain, where each step can't start until the one before it finishes.
If A takes 80ms, B takes 120ms, C takes 60ms, and D takes 100ms, the total time a user waits is the sum of all four — 360ms — even though no single step took longer than 120ms. This is a waterfall, and it's exactly the shape you'll see if you open a browser's network panel on almost any unoptimized page: one request finishing before the next one is even allowed to begin, cascading downward, each step adding its own full duration on top of everything that came before it.
#05Breaking the Waterfall: Running Steps in Parallel
Compare that to the same four pieces of work arranged so that the independent ones — the steps that don't actually depend on each other's output — run at the same time instead of one after another:
Here, A still has to finish before B, C, and E can start, and D still has to wait for all three of them to complete before it can run — but B, C, and E themselves run concurrently, not in sequence. The total time is no longer the sum of every step; it's A, plus whichever of B, C, or E takes the longest, plus D. If those three parallel branches take 120ms, 60ms, and 90ms respectively, the group only costs 120ms total, not 270ms — the two faster branches finish for free, hidden inside the slowest one's duration.
That "A, plus the slowest of the parallel group, plus D" path has a name: the critical path — the specific chain of dependent steps that actually determines total duration, regardless of how much other work is happening alongside it. Speeding up anything not on that path — making C twice as fast when B was already the slowest of the three — does essentially nothing to the number the user experiences, because C was never what they were waiting on in the first place. This is the single most common blind spot in performance work: teams optimize the component they can most easily measure, not the one actually sitting on the critical path, and then find the user-facing number barely moved.
This isn't a hypothetical diagram — it's the exact picture Chrome DevTools' Network panel draws for every real page load, and it's usually the fastest way to see whether a system is genuinely parallelized or only pretending to be. A waterfall view lays every request out as a horizontal bar, positioned by when it started and how long it took; a page built on sequential dependencies shows a visible staircase, each bar beginning only where the previous one ends, while a well-parallelized page shows a cluster of bars starting at roughly the same moment and finishing at different lengths. Anyone who has ever opened that panel on a page that "feels slow for no reason" has almost certainly seen a staircase where a cluster should have been — a page that could have loaded in the time of its single slowest request instead loading in the sum of six requests nobody realized were needlessly blocking each other.
There's a structural reason browsers used to make this worse than it needed to be, independent of anything an application team did wrong: HTTP/1.1 browsers historically capped concurrent connections per host at a small number — six was a common default — which meant a page with more render-blocking assets than that ceiling was forced into partial sequencing no matter how well the backend supported parallel requests. HTTP/2's multiplexing — many logical streams sharing one physical connection — relaxed most of that constraint, and modern browser and protocol behavior (HTTP/2 and HTTP/3 connection handling in particular) has moved further still, but the underlying point holds regardless of the exact numbers at any given moment: "make it parallel" isn't purely an application-level decision, because the protocol underneath the application has its own say in how much parallelism is even possible.
A sequential chain pays for every step in full; a parallel one only pays for its slowest branch — and knowing which shape your system actually has matters more than optimizing any single box inside it.
#06Fast Services Can Create a Slow System
Take that waterfall shape and put a real backend behind each letter, and the pattern that causes the most damage in production systems comes into focus: every individual service can be fast, genuinely fast, comfortably inside whatever SLA it was given — and the total experience can still be slow, because the services are chained on the critical path instead of running alongside each other.
A realistic version of this: a single logical request has to pass through an authentication check (50ms), then a user-profile lookup (80ms), then a permissions service (60ms), then a feature-flag evaluation (40ms), before the actual business logic (100ms) even starts. Every one of those five numbers would pass a code review. Every one of those five services would show a healthy dashboard, individually. Chained together on one request's critical path, they add up to 330ms before the work the user actually asked for has even begun — and that's before any of the five has an unlucky moment of its own.
This is the tax architectural decomposition doesn't advertise up front. Splitting a monolith into services that each do one thing well is a genuinely good idea for a long list of reasons — ownership boundaries, independent deployment, isolated failure domains. None of those reasons have anything to do with latency, and decomposition, left unmanaged, tends to make latency worse, precisely because it turns what used to be in-process function calls into network hops, and hops are exactly the kind of thing that stack into a waterfall if nobody deliberately parallelizes them. BizTechLab's own look at the return to modular monoliths covers the organizational side of that same trade-off — this chapter is the latency-shaped version of the identical lesson: the number of fast, well-run services in a critical path is not the same question as how fast the request feels.
The fix isn't "go back to a monolith," and it isn't "make each service faster" either — both miss the actual lever. The fix is looking at the shape of the dependency graph a request travels through and asking, honestly, which of those five services actually needed to happen before the next one could start, and which were only sequential because nobody had gotten around to parallelizing them. Auth genuinely has to happen before permissions can be checked. There's a real argument that the user-profile lookup and the feature-flag evaluation don't depend on each other at all, and could run concurrently — turning part of that 330ms critical path into something closer to 270ms without a single service getting any faster.
At the code level, this is often a small, almost boring-looking change — the difference between await-ing each call one after another and firing the independent ones together with something like Promise.all and waiting once for the group. It reads like a minor refactor. It's actually the exact mechanical move that turns a waterfall into the parallel-branch diagram from the previous chapter, and it's routinely skipped simply because sequential code is what most people reach for by default when nothing about the code review process forces the question "did any of these calls actually need to wait for each other."
There's a related failure worth naming because it wears the same clothes as this one: the N+1 pattern familiar from database work — one query to fetch a list, then one additional query per item in that list — has a direct analog at the service level. A request that fetches ten items and then calls a pricing service once per item, sequentially, pays for ten full round trips where a single batched call, or ten calls fired in parallel, would have paid for roughly one. The database version of this mistake gets caught in code review fairly often now, because most engineers have been burned by it once. The service-call version is newer, less discussed, and just as expensive — and it produces exactly the "every individual call was fast, the aggregate was not" symptom this chapter keeps circling back to.
There's also a genuinely organizational reason this pattern survives as long as it does. When auth, profile, permissions, feature-flags, and business logic are five separate services owned by five separate teams, "is this request fast" stops having one obvious owner — each team can honestly report their own p99 is within target, in the same incident review, without anyone in the room being wrong, and without anyone being positioned to add all five numbers together. The chain's total latency belongs to the request itself, not to any single team's dashboard, and that ownership gap is exactly where a slow system keeps hiding behind five green ones.
Five fast, healthy services chained in sequence add up to one slow request — the system's speed was never the sum of its parts' health, it was a property of how those parts were wired together.
#07Third-Party Dependencies
Every example so far has stayed inside infrastructure a team actually controls. Real systems rarely get that luxury — a checkout flow calls a payment processor, a signup form verifies an email through a deliverability service, a dashboard pulls a map tile or a currency rate from an external feed before it considers a request complete. Your own API might genuinely be a 100ms service — and if it's synchronously waiting on a payment gateway that takes 800ms on a bad day, the user's actual wait is closer to 900ms, and there is exactly nothing your own engineering team can do to that external 800ms through code review or a faster database.
This is where the waterfall-versus-parallel distinction from a few chapters back stops being an academic diagram and starts being a real architecture decision with real financial stakes. Does a checkout request call the payment processor, then the fraud-check service, then the inventory-reservation service, one after another, each waiting on the last? Or do the ones that don't genuinely depend on each other's output run concurrently, so the total wait is bounded by the slowest single call rather than the sum of all three? Most systems default to sequential simply because it's the easier code to write and reason about — call one thing, get its result, call the next — with nobody making a deliberate choice to pay that cost.
The second, quieter risk is that a third party's reliability becomes part of your own system's reliability the moment you call it synchronously and block on the answer. A payment provider having a slow afternoon doesn't just make their own dashboard look bad — it drags your own p99 down with it, on every request that touches them, and there's a real temptation to treat that as "not our problem" right up until it shows up in a customer complaint about your product being slow. A circuit breaker — a pattern that stops calling a dependency once it's clearly struggling, failing fast instead of waiting out its timeout on every request — doesn't make the third party faster, but it stops one failing dependency from silently degrading everything downstream of it while everyone assumes the problem must be local.
The practical discipline this chapter argues for is uncomfortably simple to state and genuinely easy to skip under deadline pressure: know, explicitly, which of your request's dependencies are third-party, know whether each one sits on the critical path or could be moved off it, and set a timeout on every single one of them rather than trusting their default. A request that "waits as long as the payment provider takes" has no upper bound at all; a request that waits at most 2 seconds before failing gracefully has a bound its own team actually chose.
Timeouts alone aren't the whole answer, and the naive fix — "if it times out, just retry" — can make things worse: a struggling dependency under load needs less traffic, not a synchronized wave of retries arriving the instant it shows its first sign of trouble. That's a retry storm — degradation triggers timeouts, timeouts trigger retries, the extra traffic pushes the dependency further into trouble, and the cycle feeds itself. The standard defense is retrying with exponential backoff plus a little random jitter, so retries spread out across clients instead of all landing at once.
It's also worth being honest about which third-party calls genuinely need to be synchronous at all. A payment confirmation has to complete before checkout can proceed — there's no way around waiting on it. An analytics event, a marketing pixel, a "send a welcome email" call very often don't need to block the response at all; they can be fired asynchronously, queued, or handled after the user-facing response has already gone out. Every synchronous third-party call on the critical path is a deliberate decision, whether or not anyone made it deliberately — and a surprising number of them, audited honestly, turn out to be sitting on that critical path purely out of habit, not necessity.
There's a monitoring gap specific to third parties, easy to miss until it costs someone a bad afternoon: a team's own SLO (service level objective) describes what that team committed to, and says nothing about what a payment processor or mapping API actually delivered this month. Treating a vendor's published SLA as equivalent to observed behavior is trusting a claim instead of a measurement — the two diverge, and a synthetic canary check calling that same endpoint independent of live traffic is the difference between assuming a dependency is fine and having your own continuous evidence that it is.
Your own code being fast was never the whole contract — the moment you call something you don't control and wait for its answer, its latency became your latency too.
#08The Browser Is Part of Your Performance Budget
Even a perfectly parallelized backend, with every third-party call bounded and every unnecessary sequential step removed, only gets a response as far as the network handing bytes to the browser. Everything that has to happen after that — before a human being can actually read, click, or act on anything — is real, measurable time that a server-side API dashboard has no visibility into at all, because by the time it starts, the server's job is already done.
The response body has to be parsed. In a modern JavaScript-heavy application, parsing is often the smaller cost — updating application state, triggering a re-render, and letting the framework reconcile whatever changed against what was already on screen frequently takes longer than the parse itself, especially on a mid-range mobile device rather than the developer's own laptop. If the page uses lazy-loaded chunks, this might be the exact moment a browser discovers it needs to fetch another JavaScript bundle before it can even render the component that's waiting on the API response — a second network round trip the original request never accounted for, chained sequentially after the first the same way any waterfall step is.
There's a specific, well-documented gap here worth naming directly: a page can be visually complete — every pixel painted, everything that should be on screen already there — while still being functionally unusable, because the JavaScript needed to respond to a click hasn't finished executing yet. This is sometimes described as the "uncanny valley" of web performance: it looks done, so a user reaches out to interact with it, and nothing happens, which reads as more broken than a page that's honestly still loading. The industry has settled on a real, standardized way to talk about this precise moment — Interaction to Next Paint (INP), part of Google's Core Web Vitals — specifically because "the page looks finished" and "the page will actually respond to you" turned out to need two different measurements, not one.
The specific mechanism most often responsible for that gap, in server-rendered JavaScript applications, is HydrationThe step where a JavaScript framework re-attaches event listeners and internal state to server-rendered HTML already sitting in the DOM, turning a static-looking page into a fully interactive one. — taking HTML that arrived already rendered (often for a faster first paint and better SEO) and attaching client-side JavaScript behavior to it after the fact, so buttons respond, forms submit, and state updates. Hydration has to walk the entire rendered tree, compare it against what the framework expects, and wire up event handlers throughout — real, often single-threaded work scaling with how much markup the page contains. A page can finish its network requests, finish painting, and still sit for a further several hundred milliseconds — longer on a low-end phone — before it genuinely responds to a tap, entirely invisible to any metric watching only the network.
Render-blocking resources compound this in the other direction, before hydration even gets a chance to start. A <script> tag without async or defer, or a stylesheet the browser has decided it must resolve before painting anything, can hold up the very first pixel appearing on screen, regardless of how quickly the API responded or how small the JavaScript bundle ultimately is. This is a second, distinct waterfall — HTML arrives, discovers it needs a stylesheet, fetches that, discovers it needs a font, fetches that too — layered directly on top of the network-level one this piece already described, and it's every bit as real to the person waiting for something to appear.
None of this is the API's fault, and none of it shows up in an API latency dashboard, because it happens entirely on hardware the backend team doesn't own, running code the backend team may not have written, after the response the backend team is measuring has already been delivered. That's exactly why it's so easy to miss: it's real time on the user's clock, sitting in the one part of the request's entire lifecycle that the team most likely to get paged for "slowness" has the least visibility into by default.
"The API responded" and "the user can do something with what came back" are two different timestamps — and the distance between them lives entirely outside the system that gets measured first.
#09Perceived Performance vs Actual Performance
Everything up to this chapter has been about objective time — milliseconds a stopwatch would agree on. Human perception of "fast," it turns out, doesn't purely track that number, and the gap between the two is large enough that it changes what's actually worth optimizing.
There's a long-standing rule of thumb from usability research, one that's held up remarkably well as interfaces have changed around it: a response within about a tenth of a second feels instantaneous, one within about a second keeps a user's sense of flowing with the system despite a noticeable delay, and past roughly ten seconds their attention has genuinely left the task, whether or not the system eventually finishes. Almost nothing in modern software lands under that first threshold — which means most of the interesting perceived-performance work happens in that middle zone, where the actual number matters less than whether the interface signals that anything is happening at all.
That's why a well-placed skeleton screen, a progress indicator, or a response that streams in incrementally can make an identical total latency feel meaningfully faster without changing a single millisecond of real work — a blank screen for three seconds followed by an instant result reads as slower than three seconds of visible progress toward the same result, because uncertainty is a real component of how "slow" something feels, not just duration. Optimistic UI pushes this further still: updating the interface as if an action already succeeded, then reconciling quietly with the server's actual response a moment later, can make an action feel instant regardless of what the network is actually doing underneath it.
Not every reassurance is equally effective. A spinner confirms the system is alive but gives no evidence of progress, which is exactly why one running past a few seconds starts to feel worse than no indicator at all; a determinate progress bar at least gives a basis for deciding whether to keep waiting. A skeleton screen — a placeholder shaped like the content that's coming — goes one step further by setting an expectation for what's loading, not just that something is. None of these are free to build, and none substitute for the actual milliseconds the rest of this piece has been chasing — but skipping them leaves a genuinely well-studied lever for perceived speed sitting unused.
None of it rescues a wait that genuinely needs to be five seconds — a slow system with excellent loading states is still a slow system. But it is the missing half of the story this piece has been building: a technically fast system can still feel slow with no signal that anything is happening, and a technically slower one can still feel faster if it manages the wait well. Both directions matter, and most engineering organizations only ever measure the first one.
The stopwatch and the user's own sense of "slow" are correlated, not identical — and the gap between them is exactly where good interface design earns its keep.
#010What Engineers Should Actually Measure
Put the last several chapters together and the failure at the center of this piece stops being a mystery: most organizations instrument the one segment of the journey that's easiest to measure — server-side API latency — and quietly treat it as a stand-in for the whole experience, because it's the number that lives closest to the team holding the pager.
The fix isn't to throw that number away; server-side latency is still genuinely useful for catching regressions in the code a team directly owns. The fix is refusing to let it be the only number, and being explicit about the question it can't answer on its own: not "how long did the server take," but "how long from the moment a user did something until they could do the next thing they came to do." That's a fundamentally different measurement, and it requires fundamentally different tooling to capture.
Real User Monitoring (RUM) — instrumentation running in actual users' actual browsers, on actual devices and networks, rather than a controlled test environment — comes closest to measuring that real question directly, because it captures the full chain this piece has walked through: DNS, connection setup, server processing, network transfer, parsing, rendering, and the point the page actually becomes interactive, all for the same real request on the same real hardware. Synthetic monitoring — scripted checks run on a schedule from a fixed location — still earns its place for catching regressions consistently, but it measures a controlled approximation, not the messy reality RUM captures across every device and connection quality actually in use.
The other tool worth naming directly is distributed tracing: instrumentation that follows one logical request across every service, database call, and third-party dependency it touches, and stitches all of that into a single timeline instead of leaving each hop's latency sitting in its own isolated dashboard. A trace is the difference between five teams independently confirming their own services are each healthy, and one person seeing exactly where a request spent its 330ms — which step sat on the critical path, which ran in parallel for free, and which single slow hop, out of a dozen otherwise-healthy ones, was the actual reason a request felt slow. Without it, the sequential-versus-parallel question from earlier chapters is close to unanswerable in a system with more than a couple of moving parts.
Concretely, a trace is built from spans — one per unit of work, each with its own start time and duration, tagged with an identifier tying it back to the request it belongs to — rendered as exactly the kind of waterfall this piece has already leaned on twice, except built from real production data. OpenTelemetryAn open-source observability framework providing APIs, SDKs, and tooling for collecting distributed traces, metrics, and logs — the standard instrumentation layer for debugging microservices in production., the vendor-neutral open standard most tracing tooling has converged on, exists so a span from one team's service, another team's service, and a wrapped third-party call can all land in the same trace regardless of language or framework. Most of the value comes from instrumenting the boundaries — every outgoing call, every query, every queue publish — tagged consistently enough that a request's full journey becomes one connected picture instead of a dozen logs someone has to correlate by hand at 2am.
RUM and tracing answer different halves of the same question, not competing versions of it. RUM tells you what a real user's browser experienced end to end, including everything downstream of the API. Tracing tells you why a request, once it reached the backend, took as long as it did, and which internal hop was actually responsible. A team with only RUM knows something is slow but not where; a team with only tracing knows where their own services spent time but not what the browser did with the response afterward. Getting this right means treating the two as complementary, not a choice between them.
Measuring the part of the journey a team owns is necessary but never sufficient — the metric that actually matters was always "time until the user could do what they came to do," and that metric doesn't live inside any one service's own dashboard.
#011The Bigger Lesson
Everything in this piece collapses into one idea, and it's a genuinely uncomfortable one for any team that's spent a quarter shaving milliseconds off a service that shows up first in the dashboard: optimizing the fastest, most visible, easiest-to-measure component of a system does not make the system faster, unless that specific component happens to sit on the critical path the user actually experiences.
This is the same shape of result Amdahl's Law describes in a different context — speeding up one part of a system yields a real improvement only in proportion to how much of the total time that part actually accounts for, and essentially no improvement at all if it wasn't the bottleneck in the first place. Cut a 40ms feature-flag check down to 10ms and you've genuinely improved something — but if it was running in parallel with a 120ms call the whole time, the user's total wait doesn't move by a single millisecond, because it was never waiting on the flag check specifically. Halve the time of a service that wasn't on the critical path, and the arithmetic doesn't care how much better that service's own dashboard looks afterward.
Put a number on it and the point sharpens further. A component accounting for 10% of a user's total wait, made infinitely fast, improves the total by at most 10% — twice as fast improves it by around 5%. The identical "twice as fast" effort spent on a component quietly accounting for 40% of the total is worth roughly four times as much. The question that determines whether a quarter of optimization work was worth doing was never "how much faster did we make this thing" — it was "how much of the user's total wait did this thing ever actually represent."
That's the trap: the API is usually the easiest thing to measure and the thing a backend team most directly owns, so it becomes the thing that gets optimized — whether or not it was ever the actual bottleneck. Meanwhile the DNS lookup, the sequential third-party call, the hydration cost, the render-blocking script — the pieces genuinely on the critical path — stay invisible, because nobody's dashboard was built to show them.
The reframe this whole piece has been building toward is simple to state, even though building the tooling and the habits to actually live by it is real, ongoing work: stop asking "is my API fast" as if it were the whole question, and start asking "where does this specific user's specific journey actually spend its time, and which piece of that journey, if it got faster, would the user actually notice." Sometimes the honest answer is still "the API" — plenty of systems genuinely are backend-bound, and this piece isn't an argument that server-side latency never matters. But increasingly, in systems built from enough services, enough third parties, and enough client-side complexity, the honest answer is somewhere the API dashboard was never pointed at all.
None of this requires throwing out existing dashboards and starting over — "replace everything" is exactly the kind of advice that gets a genuinely useful idea ignored under deadline pressure. It requires one deliberate addition most teams haven't made yet: a single view — RUM data, a distributed trace, or at minimum an honest whiteboard exercise walking one real user journey hop by hop — that shows the whole path a request takes, not just the one segment a single team happens to own. Everything else in this piece — the waterfalls, the tail-latency math, the hydration cost, the third-party timeout budget — is a specific thing to look for once that whole-path view exists. Before it exists, all of it is invisible by construction, no matter how carefully any one team measures their own piece.
Optimizing the fastest component of a system to be faster still was never the same project as making the system faster — and until those two get separated on purpose, most performance work will keep polishing the wrong number.
Your monitoring may tell you that your API is fast. Your users tell you whether your system is fast.

