Caches, Load Balancers and Scaling

4.Open More Checkout Counters

A

In this chapter

We'll learn how systems grow to handle more users — scalability, scaling up (a bigger truck) vs. scaling out (more trucks), load balancers (the person directing shoppers to open checkouts), and application caches — and design how BlueTicket will serve 50,000 fans.

14–16 min

The Problem in Real Life

John writes a number on the whiteboard: 2,000 requests per second. "That's roughly what Sale Day's first minute will look like. Our one server, even after the queue fix, can handle maybe 300."

"So we buy a much bigger server?" Anna asks. "That's one option," he says. "There's a better one. Let's go shopping."

J

When one checkout counter has a huge line, you don't build a giant counter. You open more counters.

John

One Bigger Machine vs. Many Machines Sharing the Work

Demand will grow fast

Sale Day's traffic is several times more than one server can handle.

Work needs sharing

With several servers, something must decide which server handles each request.

Avoid repeating work

Thousands of fans asking for the same event page shouldn't mean thousands of identical database queries.

Scalability, Load Balancers and Caches

Scalability is a system's ability to handle more work — more users, more requests, more data — by adding resources, without being redesigned. A scalable system can grow from 300 requests per second to 3,000 mostly by adding machines.

The delivery truck analogy: a delivery company gets more orders than its truck can carry. It has two choices.

  • Vertical scaling (scaling up) — buy a bigger truck: give the one server more CPU cores, more RAM, faster disks (Act 02). It's simple — nothing in the code changes. But there's a limit to how big a truck you can buy, the biggest ones are very expensive, and it's still one truck: if it breaks down, nothing gets delivered.
  • Horizontal scaling (scaling out) — buy more trucks: run several copies of the app server side by side, and share the work between them. Need more capacity? Add another copy. One breaks down? The others keep delivering. Almost every large website scales this way.
  • What makes scaling out possible — stateless servers: for any copy to handle any request, no copy can keep a fan's important data only in its own memory. That's exactly why BlueTicket stores sessions in Redis (Acts 12–13) and bookings in PostgreSQL: the app servers are stateless, like interchangeable checkout counters. A fan's second request can land on a different server and everything still works.
  • Load balancer — the person directing shoppers to open counters: in a supermarket with ten checkout counters, someone at the front points each shopper to the shortest open line, and stops sending people to a counter that's closed. A load balancer does that for servers: every request arrives at the load balancer first, which forwards it to one of the healthy app servers — taking turns, or picking the least busy one. It regularly checks each server's health and stops sending traffic to any that don't respond.
Table — Vertical vs. horizontal scaling
FeatureVertical (scale up)Horizontal (scale out)
Truck versionBuy a bigger truckBuy more trucks
HowMore CPU, RAM, disk on one machineMore copies of the server
Code changes needed?Usually noneServers must be stateless
LimitThe biggest machine you can buyPractically very high
If one machine failsEverything stopsThe others keep going
Table — The supermarket and the load balancer
SupermarketLoad balancer
Shoppers arrivingRequests arriving
Checkout countersApp servers
Person pointing to the shortest open lineLoad balancer choosing a server
A counter closesHealth check fails — no more traffic sent there
Open more counters at rush hourAdd more servers during a sale
Table — What to cache on Sale Day
DataCache?Where
Images, styles, scriptsYes, longCDN + browser
Event name, date, venueYes, ~30 secondsRedis (application cache)
Logged-in sessionsStored, not cachedRedis
Seat availability, pricesNeverAlways from PostgreSQL

BlueTicket's Sale Day plan

50,000 fans

browsers and apps

most static requests stop here

CDN

images, styles, cached pages

everything else

Load balancer

health checks, shares the requests

spread across copies

App server 1

stateless

App server 2

stateless

App servers 3, 4

add more on the day

shared by every app server

Redis

sessions + event-page cache

PostgreSQL + read replica

the truth

Queue → SMS workers

slow work in the background

Caches inside the system: in Act 12 we cached copies in the browser and the CDN. Systems also keep an application cache, usually in Redis (Act 13), between the app servers and the database. When 20,000 fans open the festival's event page, the event's name, date and venue don't change — so the first request reads them from PostgreSQL and stores a copy in Redis for, say, 30 seconds. The next 19,999 requests read the copy from memory, and the database barely notices. As Act 12 taught: cache what rarely changes; never cache what must be exact, like seat availability and prices.

The database is the hardest part to scale: app servers are easy to copy because they're stateless. The database holds the state, so it can't just be copied freely. Common steps, in order: cache reads (above), use connection pools (Act 13), add read replicas — copies of the database that handle read-only queries while one main database handles all writes — and only much later, split the data across many databases. The System Design courses cover this in depth.

Anna's Sale Day plan: she draws it next to today's diagram. Fans → CDN for images and pages → a load balancer → four identical app servers (stateless) → Redis for sessions and the event-page cache → PostgreSQL (with a read replica), plus the queue and SMS workers from the last chapter. Four servers at about 300 requests per second each, plus the CDN and caches taking most of the load: comfortably over 2,000 requests per second, with room to add more servers on the day.

Key Takeaway

Scalability is handling more load by adding resources. Vertical scaling (a bigger truck) adds power to one machine — simple but limited and still a single point of failure; horizontal scaling (more trucks) runs many copies side by side, which needs stateless servers. A load balancer directs each request to a healthy server, like someone sending shoppers to open checkouts. Application caches keep repeated reads off the database, which is the hardest part to scale.

Why This Matters

Load balancers, stateless servers, caches and horizontal scaling are the standard toolkit behind every large website. Understanding them is the difference between a system that falls over at 2,000 users and one that calmly handles 50,000 — exactly what Sale Day will test. The System Design Fundamentals course builds this into full designs.

Four app servers can handle the load. But John points at the new diagram and asks a different kind of question: "What happens if the load balancer itself dies? Or the database?" Speed is one thing. Staying up is another.

Next