In this chapter
On Sale Day morning, errors appear seconds after 10:00 — and Anna starts checking from the bottom of the stack: hardware, operating system, processes, runtime and the application itself, using what she learned in Acts 02 to 10. We'll review how those layers fit together, explained as inspecting a building from the foundations up.
The Problem in Real Life
Saturday, 9:58 AM. Sale Day. The whole team is in the office. The countdown board says NOW. On the big wall screen, the waiting room (Act 22) holds 50,000 fans. Samantha is standing so close to the screen she could touch it.
10:00:00 — the gates open. 10:00:20 — the dashboard turns orange: errors 38%. Some fans are buying tickets. Many others see nothing at all. Samantha turns around. Nobody says anything.
Anna is surprised by how calm she feels. A year ago she'd have panicked. Now she has a method. "I'll check every layer," she says, "from the bottom up." John nods and stays at his keyboard. It's her call.
Problem first. Name later. Check every layer.
Anna
Guessing Under Pressure vs. Checking Every Layer in Order
Pressure makes people guess
With 50,000 fans waiting, the temptation is to restart things and hope.
Many layers could be at fault
Hardware, operating system, runtime, network, load balancer, code, database — any one could fail.
Every minute counts
Tickets are selling — or not — right now.
From Hardware to Application: The Stack Under Every Program
The building inspection analogy: when cracks appear in a building, a good inspector doesn't start by repainting the walls. They start at the foundations, then the frame, then the plumbing and wiring, then the rooms. If the foundations are solid, they move up. Debugging a live system works the same way: check the lowest layers first, quickly, and rule them out.
Every program in this course sits on the same tower of layers. Here's the tower, from the bottom, with the Act where each one was taught.
- Hardware — the foundations (Act 02): CPU, memory and storage do the actual work. Anna checks the dashboards: CPU on the app containers is at 55% (busy, not maxed), memory is flat, disks have space. Healthy.
- Data — what's stored and moved (Act 03): everything is bits and bytes: text, prices, images. Nothing unusual in the data today — and prices are in cents, never floats (Acts 03 and 24).
- Software, builds and runtimes — the machinery (Act 04): the app is a Node.js program, built once by the pipeline into a container image (Acts 19, 21), running on the Node 20 runtime it carries inside. Same image as yesterday's successful test.
- Operating system and processes — the building manager (Act 05): the OS runs each container's processes, schedules them on CPU cores and gives them memory. No process is stuck at 100%, none has crashed or restarted. Healthy.
- Files, the terminal and logs (Act 06): the logs are where the clues are. She filters the last minute of logs by level (Act 17): plenty of normal INFO lines from the app — but far fewer requests reaching the app than fans trying to buy.
- Program logic, algorithms and data structures (Acts 07–09): the code paths for Sale Day were tested, and the seat map uses the fast hash-table lookup from Act 09. Requests that do arrive are answered in 80 ms. Healthy.
- Memory and execution (Act 10): no memory creeping up like the 3 AM leak — the graphs are flat. Healthy.
| Time | Layer checked | What she looked at | Result |
|---|---|---|---|
| 10:00:30 | Hardware | CPU, memory, disk graphs | Healthy |
| 10:00:45 | OS / processes | Container restarts, CPU per process | Healthy |
| 10:01:00 | Runtime / memory | Memory graph, image version | Healthy, same tested image |
| 10:01:15 | Application | Response times, error logs | Fast — but too few requests arriving |
| 10:01:30 | Conclusion | — | Problem is outside the servers |
The stack underneath BlueTicket's app — checked bottom up
Application (BlueTicket code)
Acts 07–09 · fast ✓
Runtime (Node.js 20 in a container)
Acts 04, 21 · healthy ✓
Operating system + processes
Acts 05, 06 · healthy ✓
Memory
Act 10 · flat ✓
CPU + memory + storage (hardware)
Act 02 · healthy ✓
10:01:30 — the conclusion from the bottom of the stack: every layer inside the servers is fine, and the requests that reach the app are fast. The problem is that many requests never reach the app at all. So the fault must be outside the servers — somewhere between a fan's phone and BlueTicket's code. "Not the code," she says out loud. "Moving out to the network." Samantha doesn't understand the words, but she understands the calm.
Key Takeaway
Every program sits on a tower of layers: hardware (CPU, memory, storage), data as bits, software built for a runtime, the operating system running processes, files and logs, the program's logic and data structures, and memory. Under pressure, check them like a building inspector, from the foundations up — quickly ruling each one in or out. When everything inside the servers is healthy but requests are missing, the fault is on the path to the servers.
Why This Matters
Real incidents are stressful, and the people who solve them fastest aren't the ones who know the most tricks — they're the ones with a method. Knowing the layers underneath every program, and checking them in order, turns a frightening moment into a checklist. That's the whole point of the first ten Acts.
Inside the servers, everything is healthy. Outside, half of 50,000 fans can't get through. Anna opens her phone, switches off Wi-Fi, and starts following a fan's request from the outside in.
