Tech #013•27 min read•13 August 2026 , Thursday

Everything Is Green. So Why Is the System Down?

How slow dependencies, retries, queues, and cascading failures bring healthy-looking distributed systems to their knees.

Rajnish Kumar

Rajnish Kumar

Editor-in-Chief & Founder

Everything Is Green. So Why Is the System Down? — Tech dispatch hero image
Editor's Context

The article approaches the problem from the system's point of view rather than any single technology, connecting these failure modes to the resilience patterns used to contain them — circuit breakers, bulkheads, backpressure, retry budgets, graceful degradation, and observability. Real production incidents and established engineering practices are used to ground the discussion in how distributed systems actually behave under failure, rather than treating each pattern as an isolated best practice.

#01The 200 OK Trap

Your API is returning 200. CPU usage looks normal. Memory is fine. The database is reachable. And yet, users are staring at a loading spinner that never resolves.

This is the moment most on-call engineers know well: every dashboard is green, every individual check is passing, and the system is still down. It isn't a paradox — it's a category error. A health check answers one narrow question: is this process alive and able to respond to a request deliberately made trivial? It says nothing about whether a real request, chained through several other services, a database, and maybe a queue, actually completes within a time a human will wait for.

The instinctive response, in the first few minutes of an incident like this, is to keep checking the same things harder — refresh the CPU graph again, re-run the same curl against the same health endpoint, double-check that the database really is reachable. All of that is checking the wrong layer. The failure this piece is actually about doesn't live inside any single service's own metrics; it lives in the interaction between services, in the time a request spends waiting on something else, and no amount of staring harder at one service's own dashboard will make that interaction visible.

A 200 response from every service in the chain is not the same claim as "the request the user actually made succeeded" — and conflating the two is exactly what turns a partial failure into a surprise.

The gap between "every component reports healthy" and "the system is healthy" is where most real production incidents actually live. This piece is about that gap: why it exists, how one slow dependency turns into a company-wide outage, and what actually stops it — not "add more monitoring," but the specific mechanical defenses that keep a partial failure partial instead of letting it eat the whole system.

#02A System Is More Than Its Services

Individual services being healthy and the overall system being healthy are different claims, and the gap between them has a name older than itself. In 1994, Sun Microsystems Fellow L. Peter Deutsch compiled a list of assumptions engineers habitually make about networks that are almost always wrong — later extended by fellow Sun engineer James Gosling, who added an eighth: the assumption that a network's hardware and behavior stay uniform end to end.

Two of them map directly onto everything covered later in this piece. "Latency is zero" is the assumption underneath every synchronous call chain that doesn't budget for how long each hop actually takes to reach across a network and back — the same assumption a timeout that's too short, or missing entirely, is quietly making. "Bandwidth is infinite" is the assumption underneath every retry policy that doesn't consider what happens when every client acts on it simultaneously — the same assumption a retry storm exposes as false, expensively, during a real incident.

Every one of these fallacies describes an assumption quietly baked into "the API returned 200, so that part of the system is fine" — and every one of them is false on a real production network.

These are known today as the Fallacies of Distributed Computing, and they're worth reading in full, because modern monitoring tooling still quietly assumes most of them are true:

  • The network is reliable
  • Latency is zero
  • Bandwidth is infinite
  • The network is secure
  • Topology doesn't change
  • There is one administrator
  • Transport cost is zero
  • The network is homogeneous

#03The Dependency Chain

A request rarely terminates at the first service it touches. In a typical production system, a single HTTP request that hits the API triggers a chain: the API calls Service A, Service A calls Service B, and Service B calls some further dependency — a database, a third-party API, an internal data service — before any response can travel back up the chain to the user.

Put concrete names on those boxes and the shape becomes familiar to almost anyone who has worked on a real production system. An e-commerce checkout button calls an orders service; the orders service calls a payments service to actually charge the card; the payments service calls out to a third-party fraud-detection API before it will authorize the charge at all. Four hops, three different teams' code, at least one dependency the company doesn't even control — and the checkout button either works or it doesn't, from the user's point of view, entirely as one indivisible thing, regardless of how many separate services quietly cooperated to answer it.

Every one of those hops adds its own latency, its own failure modes, and its own opportunity for the request to get stuck waiting on something the checkout page's own code never directly calls. The user experiences none of that structure — they experience a spinner, or a result, with no visibility into which of the four hops, if any, was the slow one.

Concretely, a chain like this is common enough that most engineers have drawn some version of it on a whiteboard during an incident, every hop reporting healthy except the one at the very bottom:

Every service in that chain can be individually healthy while the chain as a whole is not — because "healthy," for a single service, almost always means "I can respond to my own health check," not "everything I depend on is responding fast enough for me to do my job."

HopComponentOwn health check
1Client—
2APIHealthy
3Service AHealthy
4Service BHealthy
5DependencySlow

#04How One Slow Dependency Becomes a System-Wide Problem

Looking at that chain, only one box is actually unhealthy — the dependency at the bottom, running slow instead of down. Everything above it is still passing its own health check. And yet the user's request fails end to end, because Service B is now waiting on a dependency that isn't responding, Service A is waiting on Service B, and the API is waiting on Service A.

A single slow dependency doesn't just slow down the requests that actually need it — it can slow down every request an upstream service is handling, if that service uses a bounded pool of threads or connections to talk outward. If Service B has, say, 50 worker threads and each one is now stuck waiting 30 seconds on the slow dependency instead of returning in 50 milliseconds, those 50 threads fill up with waiting requests almost immediately — and every other request Service B would otherwise have handled instantly now queues behind them too, including requests that had nothing to do with the slow dependency in the first place.

This is the actual mechanism behind "one slow dependency took down the whole system": the failure doesn't spread through some mysterious contagion, it spreads because a bounded resource — threads, connections, memory — gets consumed by requests stuck waiting on the one slow thing, leaving nothing left for requests that would otherwise have succeeded instantly.

This is precisely the reasoning Netflix documented on its own engineering blog in 2012 when open-sourcing Hystrix: isolate every downstream dependency behind its own dedicated thread pool, so a dependency running slow can only ever exhaust the threads allocated to it, not the threads every other, unrelated call depends on. That isolation technique — the bulkhead pattern — gets its own section later in this piece.

Thread exhaustion isn't the only resource this pattern applies to — it's simply the most common one, because most server frameworks default to a fixed-size worker pool. Google's own Site Reliability Engineering book lists this same class of problem generically as resource exhaustion, and includes memory (a service that starts garbage-collecting more aggressively under load, slowing down further and making the underlying problem worse — a feedback loop rather than a one-time hit) and file descriptors (every open socket or file consumes one, and a server that leaks connections instead of releasing them can run out of descriptors entirely, refusing new connections even though it has spare CPU and memory sitting idle) as the same mechanism wearing different clothes. In every case, the service isn't down — it's simply out of one specific, bounded thing it needs to keep doing its job, and nothing on a basic CPU/memory dashboard reflects that shortage directly.

#05Timeouts Don't Always Protect You

The obvious fix for "Service B is waiting forever on a slow dependency" is a timeout — stop waiting after some fixed duration and fail fast instead. This helps, but it's a much weaker protection than it sounds like, for a few concrete reasons.

First, a timeout only bounds how long one call waits — it does nothing about how many calls are queued up waiting simultaneously before that timeout fires. If a dependency slows down and every one of Service B's incoming requests sets a 10-second timeout, Service B can still have all 50 of its worker threads occupied for the full 10 seconds each, even though none of them individually waits forever. The system doesn't hang indefinitely; it just becomes unavailable in 10-second waves instead.

A timeout is necessary — a call with no timeout at all can hang a worker thread indefinitely — but a timeout alone only converts "hangs forever" into "fails predictably," it doesn't prevent the failure from happening, and it doesn't stop that failure from consuming resources while it waits.

Second, a timeout that's too generous relative to what a human will actually wait for doesn't protect the user experience at all — a request that "succeeds" after 28 seconds because the timeout was set to 30 has already lost the user, whether or not it counts as a failure on a dashboard. A timeout that's too aggressive relative to a dependency's own legitimate worst-case latency does the opposite: it turns ordinary slow-but-successful responses into outright failures, which — combined with a retry policy — is exactly how well-intentioned defenses make an incident worse instead of better, covered next.

Third, timeouts have to be set per hop, and they have to actually decrease as a request moves deeper into a dependency chain — a client-facing timeout of 5 seconds is meaningless if Service B is allowed to wait 10 seconds on its own dependency, because the client will already have given up, and possibly retried, before Service B even finishes failing.

The more rigorous version of this idea is deadline propagation: instead of every hop independently deciding its own timeout, the original deadline — "this whole request needs to finish within 5 seconds" — travels with the request itself, and every downstream service computes its own remaining budget from however much of that 5 seconds is left when the call reaches it. A request that's already used 4 of its 5 seconds by the time it reaches Service B should get roughly one second of budget there, not a fresh, independently-configured 10-second allowance that guarantees the caller has already given up long before Service B is done. Google's SRE book lists deadline propagation explicitly as one of its recommended defenses against cascading failure, for exactly this reason — without it, individually reasonable per-hop timeouts can still add up to a chain that takes far longer than any single caller was ever willing to wait.

#06Retries Can Multiply the Damage

Retrying a failed request is one of the most reflexive defenses in distributed systems, and one of the easiest to get backwards. A retry is a reasonable response to a single request that failed due to a transient blip — a dropped packet, a momentary hiccup in a load balancer. It is a dangerous response to a dependency that's failing because it's already overloaded, because a retry in that situation doesn't wait for capacity to free up — it adds one more request on top of a system that just finished telling you, seconds ago, that it couldn't handle the load it already had.

Multiply this by every client retrying at roughly the same interval, and the result is what's usually called a retry storm: a dependency struggling under N requests per second gets hit with some multiple of N, from the exact clients whose retries were meant to help it recover. AWS engineer Marc Brooker described the mechanics of this precisely in a March 2015 AWS Architecture Blog post: naive exponential backoff on its own doesn't actually fix this, because every client computing backoff from the same failure event tends to retry in synchronized clusters rather than spreading out — the average load looks lower, but the peak load, arriving in bursts, doesn't. His fix — adding randomized jitter to the backoff interval — decorrelates clients from each other, turning a series of synchronized spikes into something closer to a steady, absorbable rate.

A naive retry policy assumes the problem it's retrying against is transient and isolated — but the moment the problem is capacity, not chance, every retry is additional load pointed at the one place that has the least capacity left to absorb it.

Concretely, this compounds across a dependency chain: if the API retries a failing call to Service A up to 3 times, and Service A retries its own failing call to Service B up to 3 times, a single client-initiated request can turn into up to 9 actual calls hitting Service B — a 9x amplification the client never asked for and the operators of Service B never provisioned for, generated entirely by defenses each layer thought were reasonable in isolation.

Production-grade retry logic increasingly bounds this with what's usually called a retry budget: a hard cap on total retry volume, expressed as a percentage of original request volume — commonly seen in proxies like Envoy and in gRPC's own retry configuration, both of which support capping retries at something like an extra 10-20% of normal traffic rather than leaving retry volume unbounded. Once a service's retries hit that budget, additional retries simply stop firing and the request is allowed to fail outright — a deliberate choice that a bounded number of visible failures is a better outcome than an unbounded, self-inflicted spike in load layered on top of a dependency that's already struggling.

#07The Database Is Healthy — So Why Are Requests Timing Out?

This is one of the most disorienting incidents to debug live, because every metric a database administrator would normally check — CPU, memory, disk I/O, replication lag, even query duration — can look completely normal while the application is timing out on every database call. The reason is almost always the same: the bottleneck isn't the database's ability to execute queries, it's the application's ability to get a connection to hand a query to in the first place.

A database like PostgreSQL treats connections as a rationed resource, not a free one — max_connections caps how many can exist simultaneously regardless of how much spare CPU or memory the server has left. If every one of those connections is currently checked out — held by requests that are themselves waiting on something else slow, or simply outnumbered by demand — a new request doesn't fail at the database at all. It queues, invisibly, at the connection pool, for a resource the database itself never even sees it asking for.

Notice, too, how neatly this connects back to the bounded-resource mechanism from earlier in this piece: a slow external dependency exhausts an upstream service's threads the same way a slow or overloaded database exhausts its own connection pool. It's the identical failure shape recurring at a different layer of the same system — bounded capacity, consumed by requests waiting rather than working, with nothing on a conventional dashboard distinguishing "busy doing useful work" from "stuck waiting for a turn."

The query genuinely can be fast and the database genuinely can be healthy, while the request is still slow or timing out, because the time lost happened entirely in a queue the database-side dashboards were never built to see.

This exact failure mode — a healthy database serving degraded requests because connections, not query speed, are the actual scarce resource — is its own deep, separate topic on this site, covered end to end with real sizing formulas and pool configuration; the mechanics described there are identical to what's happening here, just examined from the database side of the boundary instead of the distributed-systems side.

#08Queue / Consumer Lag

Not every dependency chain is a direct, synchronous call. A large share of production systems decouple producers from consumers with a queue — a payment event gets published, and a separate consumer process picks it up, processes it, and moves on. Queues exist precisely to absorb bursts that a synchronous chain couldn't, and most of the time that works exactly as intended.

The failure mode specific to queues is consumer lag: the gap between when a message is published and when it's actually processed, growing because consumers are processing slower than producers are publishing. Every individual service around the queue can still be healthy — the producer is happily publishing, the consumer process hasn't crashed, the queue itself hasn't run out of disk — while the time from event to effect quietly stretches from milliseconds to minutes.

Consumer lag is a partial failure with almost no visible symptoms at the service level — nothing is down, nothing is erroring, the queue is simply getting further and further behind, and unless lag itself is a monitored number, that growing gap is invisible until someone downstream asks why an event from twenty minutes ago still hasn't happened.

This matters specifically because lag compounds the same way retries do: a consumer that falls behind under load often has to catch up while new load keeps arriving, and if the consumer's own downstream call (writing to a database, calling another service) is the thing running slow, the queue backs up for the same reason a synchronous chain backs up — a bounded resource, not enough of it, and no signal reaching back to the producer to slow down.

In systems built on Kafka specifically, consumer lag — the difference between the latest offset written to a partition and the offset a given consumer group has actually processed — is a first-class, directly measurable number for exactly this reason: it turns an otherwise invisible "how far behind are we really" question into a metric that can be alerted on independently of whether the consumer process itself looks healthy. A consumer group can be running, consuming zero errors, and passing every liveness check, while its lag climbs into the hundreds of thousands of messages — which is, functionally, the queueing equivalent of the connection-pool queue from the database section above: the process is alive, but the work isn't happening on any timeline a downstream system can rely on.

#09Cascading Failure

Everything described so far — a slow dependency, exhausted threads, a retry storm, a growing queue — describes a single point of strain. Cascading failure is what happens when that strain doesn't stay contained to where it started, and instead pushes enough extra load onto the next healthy component that it starts failing too.

Google's own Site Reliability Engineering book devotes a full chapter to this exact pattern, defining a cascading failure as one that grows over time through positive feedback: one replica fails or slows under load, which pushes its share of traffic onto the remaining replicas, which increases their load and therefore their own probability of failing, which pushes even more traffic onto whatever's left. Each step looks locally reasonable — of course a load balancer routes around a failing replica — and the aggregate effect is a system eating itself from the inside.

A cascading failure rarely starts as a big, dramatic event — it starts as one component quietly running a little slower, and turns into an outage only because everything around it responded to that slowness in a way that added more load rather than less.

Overload isn't the only trigger, either. The same SRE book chapter lists a deployment or configuration push landing on many tasks simultaneously, a process crashing on a specific "query of death" input and taking capacity down with it, and even routine planned maintenance reducing available capacity right as an unrelated demand spike hits, as equally real starting points for the same feedback loop — the common thread across all of them isn't what removed the initial capacity, it's that the system's response to reduced capacity was to redistribute the same total load across fewer places able to absorb it.

A cascading failure doesn't need a service to actually crash to spread — AWS's own public postmortem of its November 25, 2020 Kinesis outage in US-EAST-1 is a precise, real-world illustration of this: added server capacity pushed a per-server operating-system thread limit past its ceiling, which broke an internal caching step, which left front-end servers with unusable routing information. Every one of those front-end servers was still running. None of them crashed. They simply couldn't do their job — and because Kinesis itself was a dependency for other AWS services, that single thread-limit problem cascaded outward across the region for hours: Cognito authentication degraded for roughly nine hours because of a latent bug in how its own webservers buffered Kinesis writes, CloudWatch metrics and log ingestion failed for over seventeen hours, Lambda invocation error rates climbed for four hours as buffered CloudWatch data caused memory contention on its hosts, and EventBridge's delayed event processing disrupted ECS and EKS cluster scaling for eleven hours — every one of those downstream services running code that never broke, sitting behind one dependency that had quietly stopped being able to route requests.

Recovery itself illustrates the same lesson from the opposite direction: AWS couldn't simply restart the whole Kinesis fleet at once, because every server coming back online simultaneously would have recreated the exact same thread-exhaustion condition that caused the outage in the first place. The fix had to be a slow, rate-limited restart — a few hundred servers brought back per hour — which is, functionally, the exact same "controlled, bounded reintroduction of load" principle behind a circuit breaker's half-open state, applied at the scale of an entire cloud region instead of one service.

Google's own SRE Book puts the definition this piece has been working from in one direct sentence:

“A cascading failure is one that grows over time as a result of positive feedback.”

#010Partial Failure — The Real Enemy

Total failure is, in a strange way, the easy case. A server that's completely down is unambiguous: health checks fail, alerts fire, load balancers route around it, and everyone agrees on what happened. Partial failure has none of that clarity, and it's the far more common real-world condition — a service that's up but slow, a dependency that succeeds 80% of the time and times out the rest, a queue that's processing but falling behind.

Partial failure is dangerous specifically because it doesn't trip the binary alarms most systems are built around. A load balancer's health check might pass even while every real request through that instance is timing out, because the health check hits a lightweight /healthz endpoint that never touches the slow dependency at all. The instance looks fine to the infrastructure and broken to every actual user — two contradictory truths, both technically correct, depending entirely on which question you asked.

Distributed systems are far more often brought down by something that's degraded than by something that's dead — a service failing 30% of its requests for six hours does more real damage, and is far harder to detect and diagnose, than a service that goes fully down for five minutes and gets paged on immediately.

This is also why partial failure is disproportionately represented in real incident postmortems relative to total outages: total failure has obvious, well-rehearsed playbooks — restart it, fail over, page someone. Partial failure requires recognizing that the system is unhealthy in the first place, which, as the rest of this piece has argued, individual service health checks are structurally unable to tell you.

Partial failure also tends to be inconsistent in a way that actively resists diagnosis: a request that fails might succeed perfectly on retry two seconds later, because the specific worker thread, connection, or downstream replica it happened to land on was the strained one, and the next attempt landed somewhere healthier. This is exactly the texture engineers describe when they say an issue is "intermittent" or "we can't reproduce it" — not because the underlying cause is rare, but because a distributed system under partial failure genuinely behaves differently request to request, depending on which specific path through its own components a given request happened to take.

#011Why Monitoring Every Service Isn't Enough

The natural response to all of this is "so monitor every service more closely" — and it's not wrong, exactly, it's just insufficient on its own. Monitoring every individual component answers "is this component doing what it's supposed to," one component at a time. It doesn't answer the question that actually determines whether a user's request succeeds: does the combination of every component's current behavior produce an acceptable outcome, end to end.

A system can have every individual service reporting 99.9% availability and still deliver a materially worse experience to users than any single one of those numbers suggests, because a request that touches five services each with 99.9% availability doesn't inherit 99.9% — it inherits something closer to the product of all five, assuming independent failures: 0.999 raised to the fifth power comes out to roughly 99.5%, meaning a user's actual end-to-end failure rate is close to five times higher than any single service's own dashboard reports, before accounting for the fact that failures in a dependency chain are rarely independent at all — a slow database, as this piece has already covered, doesn't fail one service's requests independently, it fails every service downstream of it at the same time.

Per-service monitoring tells you which component is unhealthy after something has already gone wrong for the user — it was never designed to tell you whether the request itself, as experienced end to end, is healthy in the first place.

This is the gap distributed tracing and request-level SLOs exist to close — not by replacing per-service monitoring, but by adding a layer that actually measures the thing users experience: the full path a request took, and how long the whole chain actually took to answer it, not how each box along the way answered a health check that never left its own process. Tools built specifically for this — is the open-source standard most teams reach for today — attach a single trace ID to a request at the edge and propagate it through every downstream call, so a slow or failed request can be reconstructed as one connected timeline across service boundaries instead of a dozen disconnected per-service logs someone has to correlate by hand, under pressure, during an incident.

#012How Production Systems Defend Against This

Nothing in the sections above is fixable by watching more dashboards harder. What actually stops a partial failure from becoming a full outage is a specific, well-established set of mechanical defenses, each addressing a different point where a healthy-looking system can still fail its users — and each one, notably, older than the "microservices" era that made this failure mode so common. The rest of this piece works through them one at a time.

It's worth being explicit about what these defenses have in common before going through each one individually: none of them try to prevent a dependency from ever failing — that's not achievable in a system with enough moving parts. Every one of them instead accepts that failure will happen and focuses entirely on containing its blast radius once it does, which is a fundamentally more realistic engineering goal than chasing a failure rate of zero.

#013Circuit Breakers

A circuit breaker does for a failing dependency what its electrical namesake does for a failing circuit: it stops sending current — in this case, requests — the moment failures cross a threshold, rather than letting every new request individually discover the same failure the hard way. Michael Nygard named and popularized the pattern for software in his 2007 book Release It!, borrowing the term directly from the physical device that trips to protect a circuit from damage rather than let it keep drawing current into a fault — a deliberate choice of metaphor, since an electrical breaker's whole purpose is also to fail fast and contain damage, not to somehow prevent the fault itself from ever occurring.

This connects directly back to the timeout discussion earlier in this piece: a timeout stops one call from waiting forever, but it still pays the full cost of attempting that call — the connection, the wait, the eventual failure — every single time. A circuit breaker sits one layer above that, tracking the pattern of repeated timeouts and failures, and once that pattern crosses a threshold, stops even attempting the call at all, returning a failure immediately instead of paying for a timeout the system has every reason to expect will happen again. That's also why a circuit breaker and a timeout are complementary, not competing, defenses — the timeout bounds the cost of any one attempt, and the circuit breaker decides, based on recent history, whether attempting the call at all is still worth that cost.

A circuit breaker tracks recent failures against a dependency and moves through three states: closed (requests flow normally), open (requests fail immediately, without even attempting the call, once failures cross a threshold), and half-open (after a cooldown period, a small number of test requests are allowed through to check whether the dependency has recovered, before fully reopening the flow).

The point of a circuit breaker isn't to make the dependency healthy — it's to stop paying the full cost of discovering, on every single request, that it isn't, freeing up the resources those failed attempts would have consumed to do useful work instead.

Netflix's Hystrix library, open-sourced in November 2012, was for years the reference implementation of this pattern at scale before entering maintenance mode; Netflix's own recommendation for new projects today is resilience4j, which implements the same core states.

The half-open state is the part most naive implementations get wrong, worth calling out on its own: a circuit breaker that reopens fully the instant a cooldown timer expires just recreates the exact overload it was protecting against, in one immediate burst, the moment it lets traffic back through. Sending a small, deliberately limited number of probe requests first — and only fully reopening if those succeed — is what actually distinguishes "the dependency recovered" from "the dependency is still down and we just found that out again the expensive way."

#014Bulkheads

A bulkhead is named after a ship's watertight compartments — if the hull is breached, only the flooded compartment fills with water, and the rest of the ship stays afloat. Applied to software, a bulkhead means isolating the resources used to call one dependency — a connection pool, a thread pool — from the resources used to call every other dependency, so that one dependency exhausting its allocation can't also exhaust everyone else's.

Microsoft's own Azure Architecture Center describes the exact scenario this piece has already walked through as the pattern's motivating case: a consumer calling an unresponsive service can exhaust its connection pool, and once that pool is exhausted, the consumer's calls to other, completely unrelated, perfectly healthy services start failing too — not because those services are unhealthy, but because there's no connection left to reach them with.

Without a bulkhead, every dependency effectively shares a blast radius with every other dependency, because they're all drawing from the same finite pool of threads or connections — a bulkhead's entire job is to make sure that radius stops at the dependency that's actually failing.

This is the same isolation Hystrix implemented by default: a separate thread pool per downstream dependency, so a slow payments service and a healthy recommendations service never compete for the same handful of worker threads.

The trade-off worth naming honestly is that bulkheads cost something to run: splitting one shared pool of 100 connections into five dedicated pools of 20 each means no single dependency can ever borrow spare capacity from another, even during the long stretches when that other dependency is barely being used. A shared pool is more efficient on paper, right up until the moment one dependency needs all of it and takes the rest of the system down in the process — which is precisely the failure this piece has spent several sections describing. Isolation is a deliberate exchange of some average-case efficiency for a hard ceiling on worst-case blast radius, and for anything a business actually depends on, that's usually the correct trade to make.

That same documentation states the pattern's naming motivation directly, in its own words:

“This pattern is named after the sectioned partitions (bulkheads) of a ship's hull. If the hull of a ship is compromised, only the damaged section fills with water, which prevents the ship from sinking.”

#015Rate Limiting and Backpressure

Rate limiting and backpressure solve a related but distinct problem: instead of a service passively absorbing however much load arrives and hoping it copes, it explicitly caps how much work it will accept, and pushes back — visibly, on purpose — once it's at capacity.

Rate limiting caps the number of requests a service accepts over a given window, rejecting or delaying anything beyond that cap rather than trying to serve everything and degrading unpredictably instead. Common implementations — token bucket and its close relative leaky bucket are the two seen most often in practice — allow short, natural bursts of traffic while still enforcing a steady average rate over time, rather than either rejecting every burst outright or letting bursts through unbounded. Backpressure is the same principle applied continuously, most commonly in queues and streaming pipelines: a slower consumer explicitly signals a faster producer to stop, rather than letting unbounded work pile up in memory somewhere in between.

Rate limiting and backpressure both trade a small, controlled, visible amount of rejected or delayed work now for avoiding a much larger, uncontrolled, invisible failure later — an intentional shedding of load instead of an accidental one.

The underlying idea predates modern distributed systems by decades — it's the same mechanism behind TCP flow control, where a receiver advertises how much buffer space it has left and can signal a sender to pause entirely once that buffer fills, rather than silently dropping data or crashing. The Reactive Streams specification formalizes the same idea at the application level for the JVM: a subscriber explicitly requests how many items it's ready to receive next, keeping every queue in between bounded by design rather than by luck.

#016Retries Done Right — Backoff and Jitter

Retries aren't inherently the problem described earlier in this piece — a naive retry policy is. Done correctly, a retry is still a reasonable, even necessary, response to genuine transient failures; the fix isn't to remove retries, it's to make them behave less like a synchronized pile-on and more like a system that backs off proportionally to how much trouble it just caused.

Exponential backoff is the first half of the fix: each successive retry waits longer than the last, so a client that fails once waits briefly, and a client still failing after several attempts waits substantially longer, reducing how much pressure any single client keeps applying to a struggling dependency. On its own, though, this doesn't solve the retry-storm problem described earlier, because every client computing backoff from the same failure event still tends to retry in synchronized waves.

Exponential backoff spreads out how long any one client waits between retries; jitter is what spreads different clients apart from each other, and a retry policy needs both, because backoff alone still lets every client fail together and retry together.

Jitter — adding a randomized component to each backoff interval — is the second half. Concretely: a thousand clients all following the same pure exponential schedule — 1 second, then 2, then 4, then 8 — all retry at second 1, then all retry again at second 3, then all retry again at second 7, arriving in the exact same synchronized spikes the schedule was supposed to spread out. Picking each actual wait time as a random value up to that exponential ceiling, instead of exactly at it, spreads those same thousand clients across the whole interval instead of into one instant — the average amount of waiting barely changes, but the peak concurrent load hitting the dependency drops sharply, which is the number that actually determines whether the dependency recovers or falls further behind.

Marc Brooker's analysis for AWS found that jittered backoff strategies meaningfully outperform pure exponential backoff specifically because they decorrelate clients from one another this way, converting synchronized bursts into something closer to a smooth, absorbable arrival rate. A production-grade retry policy also needs a hard cap on total attempts and, ideally, a circuit breaker sitting behind it — a retry policy with no ceiling is just a slower-motion version of the same retry storm.

#017Graceful Degradation

Every defense covered so far is about preventing failure from spreading. Graceful degradation is about what happens when, despite all of that, some part of the system still can't do its full job — and the answer, deliberately, isn't "fail the entire request."

A recommendations service that's down shouldn't take an e-commerce checkout flow down with it — the product page can render without personalized recommendations, and the request that would have failed entirely instead succeeds with slightly less functionality. A search service under heavy load might return fewer results, or skip an expensive ranking step, rather than time out completely. Google's own SRE book lists this kind of load shedding — serving a degraded but still-useful response instead of failing outright under overload — as one of the standard mitigations for cascading failure specifically.

This is also where a well-placed timeout and graceful degradation work together rather than separately: the timeout is what turns "wait indefinitely for the recommendations service" into a bounded, predictable failure, and graceful degradation is what decides what happens after that bounded failure — render the page anyway, with a generic "customers also viewed" fallback instead of a personalized one, rather than propagating that one failed call into a failed page. Neither pattern does much good without the other: a timeout with no fallback behavior just fails faster, and a fallback with no timeout still leaves the request hanging until something else eventually gives up on its behalf.

Graceful degradation accepts that "some functionality, delivered" is a genuinely better outcome than "full functionality or nothing," and builds that choice into the system deliberately, instead of discovering it by accident when something breaks.

The discipline this requires is knowing, ahead of time, which parts of a request are essential and which are enhancements — a decision that has to be made deliberately, before an incident, because it's the wrong moment to figure out for the first time, live, whether the checkout flow can survive without the recommendations widget. Feature flags are the common mechanism for this in practice — the same toggle a team uses to roll out a new feature gradually can, under load, be flipped the other direction to turn a non-essential dependency off deliberately, buying the essential path more headroom before anyone has to page an engineer.

#018Finding Out Before an Incident Does — Chaos Engineering

Every defense described in this piece has a quiet, uncomfortable problem: none of them are proven to work until the exact moment they're actually needed, which is usually the worst possible moment to discover a circuit breaker was misconfigured, a timeout was set too generously, or a bulkhead's isolation had a gap nobody noticed in code review.

Netflix's answer to this, publicly documented as part of what it called its Simian Army, was to stop waiting for real incidents to test these defenses and start deliberately causing smaller, controlled versions of them instead — a tool nicknamed Chaos Monkey that randomly terminates production instances during business hours, on the theory that a service that can't survive losing one instance shouldn't find that out for the first time during a real outage at 3 AM. The broader discipline this grew into is now generally called chaos engineering: deliberately injecting failure — killing a process, adding artificial latency to a dependency, cutting off a service's access to part of its own infrastructure — into a system that's expected to survive it, and treating any surprise that results as a finding, not an accident.

The entire premise of chaos engineering is that "we have a circuit breaker" and "our circuit breaker actually works under real failure conditions" are two different claims, and the only way to know which one is true is to cause the failure on purpose, on a schedule the team controls, instead of waiting for production to cause it on a schedule nobody controls.

Practiced well, this doesn't mean randomly breaking production without warning — most teams run it as a structured "game day," a scheduled exercise where a specific failure is injected deliberately, engineers watch how the system and its on-call responders actually behave, and the gap between "what we assumed would happen" and "what actually happened" becomes the next sprint's list of fixes. It's the same logic as a fire drill: the value isn't in the fire, it's in finding out which exit doors were actually unlocked before there's a real fire to find out during.

#019The Right Question Isn't "Is the API Healthy?"

Every pattern in this piece answers a version of the same underlying question differently than a simple health check does. A health check asks: is this process alive and able to respond. Timeouts, circuit breakers, bulkheads, rate limiting, backpressure, well-behaved retries, and graceful degradation all instead ask some version of: given everything this process depends on right now, can it actually do its job within a time that matters, and if not, does it fail in a way that protects everything around it.

That's a fundamentally different question, and it's the one that actually determines whether users experience an outage — not whether any individual box in the architecture diagram reports green.

"Is the API healthy" is a question about one component in isolation; "is the system healthy" is a question about how every component behaves together under real, current conditions — and only one of those two questions is the one users actually experience the answer to.

In practice, during a live incident, this reframing changes the actual first question an on-call engineer should ask. Not "which service is down" — that question assumes total failure and sends the investigation looking for something that, per this entire piece, may not exist. The more useful first question is closer to "which hop in the chain is the request actually spending its time in right now," which is a saturation-and-latency question, not an up-or-down one, and it's exactly the question end-to-end tracing was built to answer directly instead of by elimination.

The table below summarizes what each defense in this piece actually protects against — worth treating as a quick reference, not a substitute for understanding why each one exists:

DefenseWhat it actually protects against
TimeoutsA single call hanging a resource (thread, connection) indefinitely
Circuit breakersRepeatedly paying the full cost of calling a dependency that's already failing
BulkheadsOne dependency's resource exhaustion spreading to unrelated, healthy dependencies
Rate limitingAccepting more work than the system can actually complete
BackpressureUnbounded queue growth when a consumer falls behind a producer
Backoff + jitterSynchronized retries turning into a self-inflicted traffic spike
Graceful degradationA partial failure becoming a total failure for the entire request

#020What Should You Actually Monitor?

Given all of this, the practical question is what to actually watch — and the honest answer is that per-service uptime, while worth keeping, is close to the least useful number on this list for catching the failure mode this entire piece has described.

Google's SRE book offers a useful, deliberately small starting framework here, generally referred to as the four golden signals: latency, traffic, errors, and saturation. The reasoning behind keeping the list to four is explicit in the book itself — if only four metrics could be measured for a user-facing system, these are the four with the highest signal, and most incident-relevant dashboards are built by adding to this set, not by replacing it. Applied to everything this piece has described, latency and errors alone would have caught almost none of it — a partially degraded system routinely keeps error rates low and average latency looking fine — which is exactly why saturation, the fourth signal, matters as much as it does here.

Request-level latency at the percentile level — p95 and p99, not average — matters far more than uptime, because averages hide exactly the kind of partial failure this piece has been describing: a service can average 200ms while its p99 sits at 8 seconds, meaning one in a hundred users is having the exact bad experience the average completely conceals. Saturation metrics — how full a connection pool is, how deep a queue's backlog is, how many retries are currently in flight — catch a system heading toward failure before it actually fails, which is the entire point; by the time error rate spikes, the resource exhaustion that caused it has usually been building for minutes. And end-to-end request tracing, following one real request across every service and dependency it touched, is the only view that can actually answer "did the user's request succeed, and where did the time go" — the question every per-service dashboard, by construction, cannot.

None of these numbers matter individually as much as the discipline of tracking them together, because the entire premise of this piece is that a system fails in the space between its components — and that space is exactly what per-service health checks were never built to see.

#021The Real Lesson

None of the patterns in this piece are new, and that's the point worth sitting with rather than glossing over. Timeouts, circuit breakers, and bulkheads were named and written down almost two decades ago, in a book about production outages that happened before "microservices" was a word anyone used. TCP flow control predates the web itself. The fallacies at the start of this piece are three decades old. What's changed isn't the failure mode — it's how often modern systems, with their long chains of networked dependencies, actually encounter it, and how expensive it's become to relearn each of these lessons from scratch, one incident at a time, instead of building the defenses in deliberately before they're needed.

A healthy service doesn't mean a healthy system.

In distributed systems, the failure often lives between the components — in latency, dependencies, retries, queues, timeouts, and assumptions that were never quite true to begin with. Every pattern in this piece exists because someone, somewhere, learned that lesson the expensive way: in production, during an incident, staring at a wall of green dashboards while users stared at a spinner that never resolved.

Found this useful? Share it

Have a technical response or architectural perspective to share with the engineering desk?

Submit Engineering Feedback