In this chapter
We'll learn the two ways to find out that something happened in another system — polling and webhooks — and how events, message queues and asynchronous processing make them reliable, then fix the lost-payment bug for good with a fallback check, explained as checking your mailbox versus a doorbell.
The Problem in Real Life
Now Anna understands the whole failure: the payment provider sends a webhook — a callback request — when a payment succeeds. BlueTicket's handler did fourteen seconds of work before answering, the provider gave up after ten, retried three times, and stopped. Nothing else ever checked.
John draws two arrows on the whiteboard. One from the provider to BlueTicket — with a red ✗ on it. Then a second arrow, going the other way, labelled "every 5 min: check payment status". "The doorbell can fail. So we also check the mailbox."
Never trust one message.
John
Waiting for One Message vs. A Doorbell, a Mailbox Check and a Queue
Slow answers fail
The webhook handler did all the work before answering, and the caller gave up.
Messages arrive twice
Retries mean the same event can arrive more than once — it must not create two orders.
Messages go missing
Even with retries, some messages never arrive. Something must notice and recover.
Webhooks, Polling, Events, Message Queues and Asynchronous Processing
The doorbell and mailbox analogy: you're expecting an important parcel. You can walk to the mailbox every five minutes to check — reliable, but wasteful, and you might be up to five minutes late. Or you can rely on the courier ringing the doorbell — instant and effortless, but if you're in the shower, or the bell's broken, you miss it. The smart thing: rely on the doorbell, and check the mailbox now and then, just in case.
- Polling — checking the mailbox: the consumer asks the provider again and again: "has payment
pay_7Hk2succeeded yet?" Simple and fully under your control, but it wastes requests when nothing has changed, and there's always a delay of up to one polling interval. - Webhooks — the doorbell: a webhook is the provider calling your URL when something happens:
POST https://api.blueticket.example/webhooks/paymentswith a JSON event likepayment.succeeded. Instant and efficient — but you're now the API provider for that endpoint, and you must follow the caller's rules (fast 2xx), and handle duplicates, delays and missing deliveries. - Securing webhooks — is it really the payment provider? anyone can send a POST to your URL pretending "payment succeeded!". So providers sign each webhook with a shared secret (an HMAC signature in a header). Your handler checks the signature before believing anything (Act 22).
- Events — "this happened": an event is a message saying that something happened:
payment.succeeded,order.created,ticket.scanned. In event-driven systems, one part announces events and any number of other parts react to them — the email service, the analytics service, the venue's report — without the announcer knowing who's listening. - Message queues — the order rail (Act 14): a queue stores messages safely until a worker takes them, and lets work happen later, at the workers' own pace. If a worker crashes, the message stays and is retried. Messages that keep failing go to a dead-letter queue — a "problem shelf" that a person checks, instead of retrying forever.
- Asynchronous processing — answer now, work later: instead of doing slow work inside a request, the server accepts the job, puts it on a queue, and replies right away (often 202 Accepted). Workers do the slow part in the background. The caller gets a fast answer; the work still happens.
- Idempotent processing — the same event twice is fine: because of retries, webhooks are delivered at least once — sometimes twice. So the handler records each event's ID, and if it has already processed
evt_88aa, it does nothing the second time. A database unique constraint on the event ID makes this safe even if two copies arrive at the same moment (Act 13). - Reconciliation — the mailbox check: a scheduled job compares the two systems' records and fixes any difference. Every 5 minutes: "ask the payment provider for every payment that succeeded in the last 2 hours; for each one without an order here, create it." If a webhook is ever lost, this catches it within minutes.
| Feature | Polling (mailbox) | Webhooks (doorbell) |
|---|---|---|
| Who starts the request | The consumer, repeatedly | The provider, when something happens |
| Delay | Up to one interval | Almost instant |
| Wasted requests | Many | None |
| Can be missed | No — you keep asking | Yes — needs retries and a fallback |
| You must also | Respect rate limits | Verify signatures, reply fast, handle duplicates |
| Feature | Friday | Now |
|---|---|---|
| Time to reply | ~14 seconds | A few milliseconds |
| Work done in the request | Order, PDF, email | Save event + queue it |
| Signature checked | Yes | Yes |
| Duplicate events | Could create two orders | Ignored by event ID |
| Lost webhook | Lost order | Recovered within 5 minutes + alert |
The new payment flow
Payment provider
payment.succeeded (signed)
Webhook handler
verify signature · save event once · queue · reply 200
Queue → worker
create order + tickets (idempotent)
Email queue
confirmation
PDF queue
ticket file
Reconciliation every 5 min
ask provider: paid but no order? → create it + alert
app.post("/webhooks/payments", async (req, res) => {// 1. Is it really from the payment provider?if (!payments.verifySignature(req.rawBody, req.headers["x-signature"])) {return res.status(400).send("bad signature");}const event = req.body;// 2. Save each event once (unique constraint on event_id)const isNew = await db.query("INSERT INTO payment_events (event_id, payload) VALUES ($1, $2) ON CONFLICT DO NOTHING",[event.id, event]).then((r) => r.rowCount === 1);// 3. Queue the slow work, reply immediatelyif (isNew) await queue.publish("payments", { eventId: event.id });res.status(200).send("ok");});
const paid = await payments.listPayments({ status: "succeeded", since: minutesAgo(120) });for (const payment of paid) {const order = await findOrderByPaymentId(payment.id);if (!order) {await createOrderFromPayment(payment); // same idempotent code as the workerlog.warn("Recovered by fallback check", { paymentId: payment.id });metrics.increment("payments.recovered"); // alert if > 0}}
BlueTicket's new payment flow: (1) the webhook handler checks the signature, saves the event (ignoring it if the event ID is already saved), puts a message on a queue, and replies 200 within milliseconds; (2) a worker takes the message, creates the order and tickets in one transaction, and puts "send email" and "make PDF" on their own queues; (3) every 5 minutes, a reconciliation job asks the payment provider for recent successful payments and creates any missing order — through the same idempotent code; (4) an alert fires if the reconciliation job ever finds more than zero missing orders, so the team knows webhooks are failing.
The test: on staging, Anna turns off the webhook endpoint completely and buys a ticket with a test card. Four minutes later the reconciliation job finds the payment, creates the order, and the ticket email arrives. The log says: "Recovered by fallback check ✓ — ticket created". The alert fires, exactly as designed.
Key Takeaway
Polling is checking the mailbox (simple, wasteful, delayed); webhooks are the doorbell (instant, but they can fail, arrive late or arrive twice). Verify webhook signatures, reply fast and do the work asynchronously through a message queue, process each event idempotently, and add a reconciliation job — a regular mailbox check — so a lost message is recovered automatically. Events let many parts react to something that happened without being tied together.
Why This Matters
Webhooks, queues and background jobs are everywhere in real systems — payments, shipping, chat, CI. Knowing that messages can be late, duplicated or lost, and designing with idempotency and reconciliation, is what separates systems that quietly lose money from systems that recover by themselves. It's also a favourite topic in system-design interviews.
The lost-order problem is solved. But Anna realises it began with something simple: nobody had read the provider's documentation carefully. Next, John shows her how to read API documentation properly — and how BlueTicket should write its own.
