Webhooks, Events and Message Queues

5.The Doorbell and the Mailbox

A

In this chapter

We'll learn the two ways to find out that something happened in another system — polling and webhooks — and how events, message queues and asynchronous processing make them reliable, then fix the lost-payment bug for good with a fallback check, explained as checking your mailbox versus a doorbell.

16–18 min

The Problem in Real Life

Now Anna understands the whole failure: the payment provider sends a webhook — a callback request — when a payment succeeds. BlueTicket's handler did fourteen seconds of work before answering, the provider gave up after ten, retried three times, and stopped. Nothing else ever checked.

John draws two arrows on the whiteboard. One from the provider to BlueTicket — with a red ✗ on it. Then a second arrow, going the other way, labelled "every 5 min: check payment status". "The doorbell can fail. So we also check the mailbox."

J

Never trust one message.

John

Waiting for One Message vs. A Doorbell, a Mailbox Check and a Queue

Slow answers fail

The webhook handler did all the work before answering, and the caller gave up.

Messages arrive twice

Retries mean the same event can arrive more than once — it must not create two orders.

Messages go missing

Even with retries, some messages never arrive. Something must notice and recover.

Webhooks, Polling, Events, Message Queues and Asynchronous Processing

The doorbell and mailbox analogy: you're expecting an important parcel. You can walk to the mailbox every five minutes to check — reliable, but wasteful, and you might be up to five minutes late. Or you can rely on the courier ringing the doorbell — instant and effortless, but if you're in the shower, or the bell's broken, you miss it. The smart thing: rely on the doorbell, and check the mailbox now and then, just in case.

  • Polling — checking the mailbox: the consumer asks the provider again and again: "has payment pay_7Hk2 succeeded yet?" Simple and fully under your control, but it wastes requests when nothing has changed, and there's always a delay of up to one polling interval.
  • Webhooks — the doorbell: a webhook is the provider calling your URL when something happens: POST https://api.blueticket.example/webhooks/payments with a JSON event like payment.succeeded. Instant and efficient — but you're now the API provider for that endpoint, and you must follow the caller's rules (fast 2xx), and handle duplicates, delays and missing deliveries.
  • Securing webhooks — is it really the payment provider? anyone can send a POST to your URL pretending "payment succeeded!". So providers sign each webhook with a shared secret (an HMAC signature in a header). Your handler checks the signature before believing anything (Act 22).
  • Events — "this happened": an event is a message saying that something happened: payment.succeeded, order.created, ticket.scanned. In event-driven systems, one part announces events and any number of other parts react to them — the email service, the analytics service, the venue's report — without the announcer knowing who's listening.
  • Message queues — the order rail (Act 14): a queue stores messages safely until a worker takes them, and lets work happen later, at the workers' own pace. If a worker crashes, the message stays and is retried. Messages that keep failing go to a dead-letter queue — a "problem shelf" that a person checks, instead of retrying forever.
  • Asynchronous processing — answer now, work later: instead of doing slow work inside a request, the server accepts the job, puts it on a queue, and replies right away (often 202 Accepted). Workers do the slow part in the background. The caller gets a fast answer; the work still happens.
  • Idempotent processing — the same event twice is fine: because of retries, webhooks are delivered at least once — sometimes twice. So the handler records each event's ID, and if it has already processed evt_88aa, it does nothing the second time. A database unique constraint on the event ID makes this safe even if two copies arrive at the same moment (Act 13).
  • Reconciliation — the mailbox check: a scheduled job compares the two systems' records and fixes any difference. Every 5 minutes: "ask the payment provider for every payment that succeeded in the last 2 hours; for each one without an order here, create it." If a webhook is ever lost, this catches it within minutes.
Table — Polling vs. webhooks
FeaturePolling (mailbox)Webhooks (doorbell)
Who starts the requestThe consumer, repeatedlyThe provider, when something happens
DelayUp to one intervalAlmost instant
Wasted requestsManyNone
Can be missedNo — you keep askingYes — needs retries and a fallback
You must alsoRespect rate limitsVerify signatures, reply fast, handle duplicates
Table — Friday's handler vs. the new handler
FeatureFridayNow
Time to reply~14 secondsA few milliseconds
Work done in the requestOrder, PDF, emailSave event + queue it
Signature checkedYesYes
Duplicate eventsCould create two ordersIgnored by event ID
Lost webhookLost orderRecovered within 5 minutes + alert

The new payment flow

Payment provider

payment.succeeded (signed)

webhook (doorbell)

Webhook handler

verify signature · save event once · queue · reply 200

async

Queue → worker

create order + tickets (idempotent)

more events

Email queue

confirmation

PDF queue

ticket file

the mailbox check

Reconciliation every 5 min

ask provider: paid but no order? → create it + alert

The new webhook handler
app.post("/webhooks/payments", async (req, res) => {
// 1. Is it really from the payment provider?
if (!payments.verifySignature(req.rawBody, req.headers["x-signature"])) {
return res.status(400).send("bad signature");
}
const event = req.body;
// 2. Save each event once (unique constraint on event_id)
const isNew = await db.query(
"INSERT INTO payment_events (event_id, payload) VALUES ($1, $2) ON CONFLICT DO NOTHING",
[event.id, event]
).then((r) => r.rowCount === 1);
// 3. Queue the slow work, reply immediately
if (isNew) await queue.publish("payments", { eventId: event.id });
res.status(200).send("ok");
});
The reconciliation job (every 5 minutes)
const paid = await payments.listPayments({ status: "succeeded", since: minutesAgo(120) });
for (const payment of paid) {
const order = await findOrderByPaymentId(payment.id);
if (!order) {
await createOrderFromPayment(payment); // same idempotent code as the worker
log.warn("Recovered by fallback check", { paymentId: payment.id });
metrics.increment("payments.recovered"); // alert if > 0
}
}

BlueTicket's new payment flow: (1) the webhook handler checks the signature, saves the event (ignoring it if the event ID is already saved), puts a message on a queue, and replies 200 within milliseconds; (2) a worker takes the message, creates the order and tickets in one transaction, and puts "send email" and "make PDF" on their own queues; (3) every 5 minutes, a reconciliation job asks the payment provider for recent successful payments and creates any missing order — through the same idempotent code; (4) an alert fires if the reconciliation job ever finds more than zero missing orders, so the team knows webhooks are failing.

The test: on staging, Anna turns off the webhook endpoint completely and buys a ticket with a test card. Four minutes later the reconciliation job finds the payment, creates the order, and the ticket email arrives. The log says: "Recovered by fallback check ✓ — ticket created". The alert fires, exactly as designed.

Key Takeaway

Polling is checking the mailbox (simple, wasteful, delayed); webhooks are the doorbell (instant, but they can fail, arrive late or arrive twice). Verify webhook signatures, reply fast and do the work asynchronously through a message queue, process each event idempotently, and add a reconciliation job — a regular mailbox check — so a lost message is recovered automatically. Events let many parts react to something that happened without being tied together.

Why This Matters

Webhooks, queues and background jobs are everywhere in real systems — payments, shipping, chat, CI. Knowing that messages can be late, duplicated or lost, and designing with idempotency and reconciliation, is what separates systems that quietly lose money from systems that recover by themselves. It's also a favourite topic in system-design interviews.

The lost-order problem is solved. But Anna realises it began with something simple: nobody had read the provider's documentation carefully. Next, John shows her how to read API documentation properly — and how BlueTicket should write its own.

Next