Availability, Reliability and Failure

5.The Only Bridge Into Town

A

In this chapter

We'll learn how systems stay up when parts fail — availability and the "nines", reliability, fault tolerance and redundancy (an airplane with two engines), and the single point of failure (one bridge into town) — then hunt every one of them in BlueTicket's diagram.

14–16 min

The Problem in Real Life

John puts his finger on the load balancer in Anna's Sale Day diagram. "One box. If it dies at 10:00 on Sale Day, what happens?" She follows the arrows. Every fan's request goes through it. "Then... nobody gets in. Not even to the four healthy app servers behind it."

He moves his finger to PostgreSQL. "And this?" She doesn't need to answer. "Four app servers make us fast," he says. "They don't make us safe. For that, we need a different way of looking at the diagram."

J

Look at every box and ask one question: if this disappears right now, what stops working?

John

"It Works" vs. "It Keeps Working When Something Breaks"

Everything fails eventually

Servers crash, disks die, networks drop. The question isn't if, but what happens when.

Some parts are single doors

If the only path goes through one box, that box can take everything down.

Uptime has a number

"Always up" isn't a plan. Availability is measured, promised and designed for.

High Availability, Reliability and Fault Tolerance

The airplane analogy: a passenger plane has two engines, even though it can fly on one. Hospitals have backup generators. Cars carry a spare tyre. Nobody expects these parts to fail today — but everyone knows something will fail eventually, so the important parts have a spare. Systems that must stay up are designed the same way.

  • Availability — how much of the time it's up: availability is the share of time a system is working and reachable, usually written as a percentage with "nines": 99% ("two nines") sounds great, but means about 3.7 days of downtime a year. 99.9% ("three nines") is about 8.8 hours a year. 99.99% is about 53 minutes a year. Each extra nine is much harder and more expensive. High availability (HA) means designing a system to reach a high number like these.
  • Reliability — does it work correctly over time? Reliability is how consistently a system does the right thing, without failing or giving wrong results. A system can be available but unreliable — up all day, but double-booking seats (Act 13). Both matter: fans need the site to be up and right.
  • Fault tolerance — keep flying on one engine: fault tolerance is a system's ability to keep working when part of it fails. With four app servers behind a load balancer, one can crash and fans won't notice — the system tolerates that fault. It's achieved with redundancy: having more than one of each important part, so a spare can take over.
  • Single point of failure (SPOF) — the one bridge into town: if a town has only one bridge and it collapses, nobody can get in or out, however good the roads on either side are. A single point of failure is any one part whose failure stops the whole system. Finding and removing SPOFs is one of the most important habits in architecture.
Table — What the "nines" really mean
AvailabilityDowntime per yearDowntime per month
99% (two nines)about 3.7 daysabout 7.3 hours
99.9% (three nines)about 8.8 hoursabout 44 minutes
99.99% (four nines)about 53 minutesabout 4.4 minutes
99.999% (five nines)about 5 minutesabout 26 seconds
Table — Anna's single-point-of-failure hunt
BoxIf it disappears...SPOF?Fix
Load balancer (one)Nobody gets inYesTwo, with automatic takeover
App servers (four)Others carry the loadNoAlready redundant
PostgreSQL (one)No bookings at allYesStandby replica + failover + backups
Redis (one)Everyone logged out, cache goneYesRedis replica
QueueSMS paused, bookings finePartialReplica; messages wait safely
SMS providerTexts delayedNo (degrades)Queue holds messages until it's back
Table — Availability vs. reliability vs. fault tolerance
TermQuestion it answersExample
AvailabilityIs it up?The site answers 99.9% of the time
ReliabilityDoes it work correctly over time?No seat is ever sold twice
Fault toleranceDoes it keep working when a part fails?One app server dies; fans don't notice

Removing single points of failure

Load balancer ×2

one takes over if the other fails

redundant

App servers ×4

one can crash, others carry on

redundant data stores

PostgreSQL main + standby

automatic failover, plus backups

Redis main + replica

sessions survive

Queue + SMS workers

if SMS fails, messages wait — graceful degradation

Anna's SPOF hunt: she goes through the Sale Day diagram box by box, asking "if this disappears, what stops?" The load balancer: everything — SPOF. Fix: two load balancers, with the second taking over automatically if the first fails (cloud providers offer this built in). The app servers: nothing — there are four, so it's not a SPOF. PostgreSQL: all bookings stop — SPOF. Fix: a standby replica kept up to date that can be promoted to main within seconds (failover), plus regular backups. Redis: everyone is logged out and the cache disappears — SPOF for sessions. Fix: a Redis replica too. The queue: SMS messages pause, but bookings still work — a partial failure, acceptable for an hour, but worth a replica. The SMS provider: messages are delayed — the queue simply holds them until it's back.

Failing safely, not just failing less: some failures can't be prevented, so a good system decides what happens when they occur. If the SMS company is down, bookings continue and texts are sent later. If the cache is empty, requests go to the database, a bit slower. This is called graceful degradation: losing a feature, not the whole system.

John adds one more box to the diagram, far to the side: a second data centre (availability zone). "Even a whole building can lose power," he says. "That's Act 20." Samantha looks at the finished diagram and asks the only question she cares about: "So if one thing breaks on Sale Day, we keep selling?" For the first time, Anna can answer: "Yes. If one thing breaks."

Key Takeaway

Availability is the share of time a system is up (99.9% ≈ 8.8 hours of downtime a year); reliability is whether it works correctly over time. Fault tolerance — keeping going when a part fails — comes from redundancy, like a plane's second engine. A single point of failure is one part whose loss stops everything, like the only bridge into town: find every one, add a spare, and design failures to degrade gracefully.

Why This Matters

On Sale Day, ten minutes of downtime at 10:00 AM would cost BlueTicket its biggest client. Every serious system is designed around the assumption that parts will fail — and the SPOF question ("if this disappears, what stops?") is something you can ask about any system, from a website to a payment flow to a team that depends on one person who knows everything.

Anna's diagram has grown from one box to more than a dozen, with arrows everywhere. John tells her that reading diagrams like this — quickly, and asking the right questions — is a skill on its own. It's the last one in this Act.

Next