Servers, Cloud, Delivery and Security

4.Sold Out in 41 Minutes

A

In this chapter

Underneath the application, the infrastructure carries Sale Day — servers across two cloud zones, containers scaling from 4 to 22, a pipeline hotfix shipped in eight minutes, monitoring that spotted the problem in twenty seconds, and the security that kept bots out — reviewing Acts 14 to 22, explained as the stadium, staff and safety systems behind a big concert.

12–14 min

The Problem in Real Life

10:15. Liam notices a small bug on the checkout page: the total shows 178 instead of 178.00 for some fans. Harmless, but ugly. A year ago, someone would have fixed it on the production server by hand. Today, Liam opens a one-line pull request.

John reviews it in two minutes. The pipeline builds, tests and deploys the image. At 10:23 the fix is live, rolled out across every container while 30,000 fans are buying tickets — and nobody notices except the fans whose totals now look right.

J

On the busiest morning of the year, we deployed — and it was boring. That's the goal.

John

Machines Set Up by Hand vs. Infrastructure That Runs Itself

Capacity on demand

Traffic jumps from normal to 50,000 people in seconds — and back down an hour later.

Changing a running system

Even on Sale Day, small fixes must ship safely while fans are buying.

Seeing everything at once

Dozens of containers, two zones, many services — the team needs one clear view.

Servers, Cloud, Containers, CI/CD, Monitoring and Security

The stadium analogy: a big concert needs much more than the band. It needs the stadium (built to survive problems), extra staff who arrive when the crowd does, a crew who can change things on stage between songs without stopping the show, a control room watching every camera, and security at every gate. BlueTicket's infrastructure is all of those.

  • Servers and the cloud — the stadium (Acts 19, 20): everything runs on rented cloud servers in two availability zones in Europe (London). If a whole zone had failed during the sale, the other would have kept going, with the database's standby taking over (Act 20).
  • Containers and autoscaling — extra staff when the crowd arrives (Acts 20, 21): the app is a container image. As traffic rose, the platform added copies automatically: from 4 containers at 9:55 to 22 at 10:10, each starting in seconds. After 11:00, it scaled back down — so BlueTicket paid for 22 containers for about an hour, not all year.
  • CI/CD — changing the stage between songs (Acts 15, 18, 19, 21): Liam's fix went the normal way: branch, small PR, review, pipeline (build once, test inside the image, push), rolling deploy with health checks. Because the path was automatic and tested, using it during the sale was safe. No hand deploys — the lesson of Act 19, and of node-2 this morning.
  • Monitoring and alerting — the control room (Acts 17, 19, 21): the error-rate alert fired at 10:00:20, twenty seconds after the problem started. Dashboards for every layer let Anna rule each one in or out in seconds (chapter one). Logs, metrics and traces (Act 21) were all there when she needed them.
  • Security — guards at every gate (Act 22): HTTPS everywhere (the very thing that failed on node-2 — and the browsers were right to refuse it), rate limits and the waiting room against bots, restricted payment keys in a secrets manager, MFA on every admin account. None of it was visible to fans. All of it was working.
  • Professional practices — the crew's habits (Acts 16, 18, 24, 25): the README and runbooks meant anyone could act; small reviewed PRs meant fixes were safe; the organiser dashboard's live counter (Act 25) let venues watch their sales climb, instead of phoning Samantha.
Table — The infrastructure on Sale Day
PieceStadium versionTaught inSale Day
Two availability zonesA stadium built to survive problemsAct 20Ready if a zone failed
Container autoscalingExtra staff when the crowd arrivesActs 20, 214 → 22 containers → back to 4
CI/CD pipelineChanging the stage between songsActs 19, 21Fix shipped in 8 minutes
Monitoring + alertsThe control roomActs 17, 19, 21Alert after 20 seconds
SecurityGuards at every gateAct 22Bots blocked, HTTPS enforced
Docs + practicesThe crew's habitsActs 16, 18, 24, 25Anyone could act
Table — Sale Day timeline
TimeEvent
09:5850,000 fans in the waiting room
10:00:20Error-rate alert: 38% errors
10:01:30Servers healthy — problem is outside them
10:04node-2 certificate expired — found
10:05node-2 removed; errors near zero
10:10Autoscaled to 22 containers
10:23Checkout display fix live via the pipeline
10:41Sold out — 0 double bookings

10:41 — sold out: the last ticket for the City Music Festival sells 41 minutes after the gates opened. 50,000 tickets. Zero double bookings. Every paid order has a ticket. The error rate after 10:05: 0.03%. On the organiser dashboard, the live counter stops at exactly the venue's capacity.

Key Takeaway

The infrastructure carried Sale Day like a stadium and its crew: cloud servers across two zones survive building failures; containers autoscaled from 4 to 22 and back; the CI/CD pipeline shipped a fix safely in eight minutes mid-sale; monitoring alerted in twenty seconds and let every layer be checked in seconds; and security — HTTPS, rate limits, restricted keys, MFA — worked invisibly at every gate. Good infrastructure makes even Sale Day deploys boring.

Why This Matters

Modern infrastructure — cloud zones, containers, pipelines, monitoring and security — is what lets a small team handle a huge event calmly. Understanding how each piece contributes, and why hand-made exceptions are dangerous, is the foundation for any backend, DevOps or cloud role.

Sold out. The office erupts. When it calms down, John asks Anna to do one last thing before lunch: follow a single fan's ticket purchase through every layer — the whole year in one journey.

Next