Tech #018•8 min read•31 August 2026 , Monday

Your API Can Retry. Your Business Can't.

A retry is technically safe only when repeating the request is business-safe — and most systems retry first, ask that question never.

Rajnish Kumar

Rajnish Kumar

Editor-in-Chief & Founder

Your API Can Retry. Your Business Can't. — Tech dispatch hero image

#01The Retry That Looks Harmless

Every engineer has written this code without thinking twice about it: a request goes out, the socket times out before a response comes back, and the client — a mobile app, a frontend fetch call, a backend service calling another backend service — quietly tries again. Nobody reviews this as a risky change. It's the textbook answer to "what if the network is unreliable," and the network is, in fact, always unreliable. A packet gets dropped, a load balancer recycles a connection mid-request, a pod gets rescheduled two milliseconds before it would have written the response. Retrying is the correct first instinct for almost every kind of failure a distributed system produces.

The trouble is that "the request failed" is not actually something a client can observe. A client can only observe that it did not receive a response. Those sound like the same fact. They are not, and the gap between them is where this entire piece lives.

A timeout tells you the response didn't arrive — it never tells you whether the work did or didn't happen.

#02But What If the First Request Already Succeeded?

Picture a checkout flow. A user taps "Pay." The client sends POST /payments. The payment gateway receives it, charges the card, writes a succeeded row to its own database — and then, on the way back, the response never makes it to the client. Maybe the mobile network drops mid-download, maybe a proxy in between times out first, maybe the client's own app gets backgrounded by the OS for half a second at exactly the wrong moment. From the server's point of view, the payment is done. From the client's point of view, nothing happened at all, because nothing came back.

Nothing in the sequence that follows is a bug in the traditional sense. The gateway does exactly what it was told, twice. The client does exactly what a well-behaved client is supposed to do when a request times out. The double charge isn't the result of broken code — it's the result of two correct pieces of code holding two different, equally reasonable beliefs about whether the first request succeeded, laid out step by step below.

The server knows the truth the instant it happens. The client only finds out if the network agrees to tell it.

  • User taps "Pay."
  • Payment succeeds on the server.
  • The response is lost in transit.
  • The client sees only a timeout.
  • The client retries the same request.
  • The server processes it as a brand-new payment.
  • The card gets charged twice.

#03Same Request, Different Consequences

Not every retry carries this risk, and the reason comes down to what HTTP method the request uses. GET /orders/8472 can be sent a hundred times in a row and nothing changes — reading the same order twice produces the same result both times, so a retry is free. POST /payments is a different kind of request entirely: each successful call is, by default, a brand-new instruction to "charge this card," and the server has no built-in way to tell "the same payment, sent again" apart from "a second, separate payment the user genuinely intended."

The HTTP specification actually has vocabulary for exactly this distinction, and it's worth knowing precisely because so few engineers ever read it past the method names themselves. PUT /users/8472 { "name": "Alex" } is idempotent by design — sending it five times still leaves the name set to "Alex," not five renames. DELETE /orders/8472 is idempotent too, in the sense that "already deleted" is an acceptable outcome for the second call. POST, on the other hand, is defined as not idempotent, precisely because its most common use — "create a new thing" — has no natural way to distinguish a resend from a genuine second creation. This is exactly why POST /payments is the request type that turns "let's just retry" from a reasonable default into a real financial risk.

A retried GET is a non-event. A retried POST is a decision the client is making on the business's behalf, whether it means to or not.

MethodSafe (no side effects)Idempotent (repeat = same effect)
GETYesYes
HEADYesYes
PUTNoYes
DELETENoYes
POSTNoNo
PATCHNoNo

#04The Real Problem Is Ambiguity

Strip away the payment example and the underlying problem is simpler, and more general, than "payments are dangerous." A client sitting on a timeout cannot distinguish two very different realities from where it stands — both produce the exact same symptom: silence where a response should have been. The server always knows which one occurred, because it has state the client doesn't have access to; the entire discipline this piece is building toward exists to close that gap for the client too, not to make failures rarer but to make it safe to act without knowing which of the two actually happened. Concretely, whenever a client sends a mutating request and doesn't get a response back, exactly two things could have happened:

Two completely different realities produce the exact same symptom on the client — and the client is the one deciding whether to retry.

  • The request never reached the server, or the server rejected it before doing any work — nothing happened, and a retry is exactly correct.
  • The request reached the server, the work completed, and only the response got lost on the way back — the operation already happened, and a retry duplicates it.

#05Idempotency Enters the Story

The formal answer to that ambiguity is : an operation is idempotent when performing it once and performing it multiple times leave the system in the same state. PUT and DELETE get this property for free from the HTTP spec's own definition. POST — creating a payment, placing an order, sending an email — does not, because "create" has no natural notion of "already done" built into it the way "set this value" or "remove this row" do.

The fix isn't to stop using POST for these operations; creating things is a real, necessary category of request. The fix is to make a naturally non-idempotent operation behave idempotently anyway, by giving the server enough information to recognize "I've already done this exact logical operation" and respond with the original result instead of performing the work a second time. That's the whole idea behind an idempotency key, and it's a pattern, not a protocol feature — nothing in HTTP grants it automatically the way it grants safety to GET.

Idempotency isn't a property some requests happen to have — it's a property engineers deliberately design into the ones that don't.

#06How an Idempotency Key Actually Works

The client generates a unique value — a UUID is the common choice — once per logical operation, before the first attempt, and attaches it to every retry of that same operation. On the server side, that key becomes the thing that gets checked before any payment logic runs at all: the first time it arrives, the server processes the payment normally, then stores the key alongside the response it generated — Stripe's public API, one of the best-known real-world implementations of this pattern, keeps that stored response around for roughly 24 hours before the key expires. If the exact same key arrives again inside that window, the server skips the charge entirely and returns the stored response from the first attempt, so a client retrying because it never saw the original response gets back the same "payment succeeded" confirmation instead of a second charge.

One detail is easy to get wrong and worth naming directly: a robust implementation doesn't just match on the key, it also stores a fingerprint of the request body and checks that a repeated key arrives with the same payload. Without that check, a client bug that accidentally reuses a key across two genuinely different payments would silently return the wrong stored response for the second one — the key alone isn't a safety guarantee, the key paired with the original request is. Concretely, a first attempt and its retry look like this:

An idempotency key doesn't stop the retry from being sent — it stops the retry from being acted on twice.

POST /payments
Idempotency-Key: order-8472-payment
{
"amount": 4999,
"currency": "INR",
"order_id": "8472"
}

#07Retry Isn't Only a Payments Problem

Payments make the risk viscerally obvious because the consequence has a currency symbol attached to it, but the underlying pattern — a mutating request, a lost response, an ambiguous retry — shows up anywhere a client can't tell "failed" from "succeeded, but I never heard back."

Wherever a request changes state and a client can lose the answer, the payments story replays with a different noun in place of "charge."

  • Order creation — a lost response to POST /orders can leave a customer with two identical orders and two identical shipments, discovered only when the second box arrives.
  • Ticket booking — a retried booking request against a limited-inventory event can double-book the same seat, or charge a customer for two seats they never intended to hold.
  • Email and notification sending — a retried "send confirmation email" call is not dangerous the way a payment is, but a customer receiving the same email four times reads as broken software regardless.
  • Inventory reservation — retrying a "reserve 3 units" call without deduplication can silently over-reserve stock that was never actually available, corrupting a count every downstream system trusts.
  • Subscription creation — a duplicated POST /subscriptions can enroll a customer in the same recurring plan twice, generating a second recurring charge that surfaces weeks later as a support ticket, not an immediate error.
  • Webhook processing — most webhook providers document at-least-once delivery explicitly, meaning the same event can legitimately arrive more than once by design; a receiver that isn't built to deduplicate by event ID will process a "refund issued" event twice and refund the customer twice.

#08Designing for Safe Retries

None of this argues against retrying — a system that gives up on the first timeout is worse, not safer, since most timeouts really are the harmless kind. The actual discipline is making sure a retry is safe before it happens, not hoping it turns out fine after.

Retry logic and idempotency logic are two separate decisions — one about when to try again, the other about what happens if the attempt before it actually landed.

  • Generate the idempotency key client-side, once, before the first attempt — not fresh on every retry, which defeats the entire mechanism.
  • Apply this only to unsafe, non-idempotent methods (POST, and any custom action-style endpoint) — GET and PUT already don't need it.
  • Give stored keys a real expiry window rather than keeping them forever; a day or two is usually enough to cover realistic client retry behavior without growing an unbounded table.
  • Back the key check with a real database constraint — a unique index on the idempotency key column — so the guarantee holds even under concurrent retries racing each other, not just in the common single-threaded case.
  • Treat retries with an and a little random jitter between attempts as the default, so a struggling downstream service faces a spread-out trickle of retries instead of every client retrying in lockstep the instant it recovers.
  • For webhook consumers specifically, deduplicate on the provider's own event ID — the fix for at-least-once delivery lives on the receiving side, not by asking the provider to promise something HTTP was never built to guarantee.

#09The Deeper Lesson

Step back from the mechanics and the real shift this piece has been arguing for is a change in what "reliable" even means. It's tempting to treat reliability as a property of uptime — the service didn't crash, the request eventually got a response, the retry worked. That framing quietly assumes failure is the enemy and success is the goal. A more useful framing treats failure as a given — networks drop packets, servers restart, responses go missing, all of that is simply the operating environment every distributed system runs in — and asks a different question instead: when failure happens, is what comes next safe?

That's the real distinction between a system that merely retries and one that's actually reliable. Retrying is easy; almost any client can be made to try again. What's hard, and what actually protects a business, is making sure that trying again produces the same outcome as trying once — so that the engineer who added a retry loop for resilience never becomes the reason a customer got charged twice, booked the same seat twice, or received a refund twice.

Reliability isn't about making systems never fail. It's about making failure safe to recover from.

Your API can retry as many times as the network makes necessary. Whether your business can survive that retry was decided the moment someone designed — or forgot to design — for it.

Found this useful? Share it

Have a technical response or architectural perspective to share with the engineering desk?

Submit Engineering Feedback