Tech #019•12 min read•7 September 2026 , Monday

The Hardest Part of Idempotency Happens After the Database Commit

An idempotency key protects the client-facing boundary. It does nothing for the boundary that actually breaks — the one between your own database and the system you're calling next.

Rajnish Kumar

Rajnish Kumar

Editor-in-Chief & Founder

The Hardest Part of Idempotency Happens After the Database Commit — Tech dispatch hero image

#01What "Idempotent" Actually Means Here

Before going further, it's worth being plain about the word itself, because most of the confusion in this space starts with people using it loosely. An operation is idempotent if doing it once and doing it five times leave the system in exactly the same state. Setting a light switch to "on" is idempotent — flip it to "on" a hundred times and the room is exactly as lit as it was after the first flip. Toggling the switch is not idempotent — do that a hundred times and you're in the dark. "Charge this card $50" is not naturally idempotent either: run it twice and the customer is out $100, not $50. An idempotency key is the trick that makes a naturally non-idempotent action, like a charge, behave like the light switch instead of the toggle — the server remembers "I already did this exact request" and returns the same result instead of doing it again. That part is well understood, well documented, and genuinely solved by most payment providers and API frameworks today. This piece is about the part that isn't: what happens inside your own system, after the key has already done its job at the front door.

#02The Diagram That Looks Fine Until It Isn't

Every idempotency writeup — this site's own included — walks through the same picture: a client sends POST /payments, the response gets lost on the way back, the client retries, and an idempotency key stops the server from charging the card twice. That picture is correct, and it's also incomplete, because it quietly assumes the server itself is one single, atomic thing. It isn't. Trace one layer deeper and a second, uglier version of the same problem shows up — this time between your own server and whatever it calls next. Concretely, it plays out as one continuous chain: the client's POST /payment reaches the API, the API commits the attempt to its own database, the API then calls out to the payment provider, and somewhere on that outbound leg — before the provider ever sees the request, or after it's already charged the card — the network drops. The client sees nothing but a failed request, retries the exact same call, and the API is left facing that same request a second time with no idea which of the two outcomes actually happened the first time.

The database commit in that chain is real and durable — that row exists, permanently, the moment it lands. Everything below it is not yet certain. The call to the payment provider can fail before it's sent, fail after the provider processes it but before the response comes back, or succeed completely with the confirmation lost in transit — and from where your API is standing, all three look identical: an exception, a timeout, silence. There is no error message that distinguishes "the card was never charged" from "the card was charged and I just never heard about it." Both show up as the exact same stack trace.

The database commit is a fact. Everything after it, until proven otherwise, is a guess.

The chain a lost response actually breaks

Client

POST /payment

reaches the API

API

commits the attempt to its own database

the commit is durable

Payment provider

the outbound call goes out

on the outbound leg

Network

drops before or after the charge

nothing comes back

Client

sees a failed request, retries

#03The Question Nobody's Retry Logic Actually Answers

Strip away the payments example and what's left is one specific, uncomfortable question, and the entire discipline in this piece exists to give a system a real answer to it instead of a shrug:

If the operation succeeded but the caller never learned that it succeeded, what exactly should the system do when the request comes back?

Most retry logic never actually answers this. It answers a different, easier question — "should I send the request again?" — and treats that as the whole problem. A typical retry wrapper checks the HTTP status code, maybe checks whether the error looks "retryable," waits a bit, and tries again. None of that logic asks the harder question first: does anyone actually know what happened last time? Sending the request again is only safe once you already know what "again" means for the state your own database is sitting in. An idempotency key on the client-facing endpoint answers that question for the client's retry — the client doesn't need to know or care what happened internally, it just needs the same answer every time it asks. But nothing has answered it yet for what the API itself is supposed to do about the half-finished attempt sitting in its own database, the moment before it decides whether to call the payment provider again.

#04Why an Idempotency Key Alone Doesn't Reach This Far

The client-facing idempotency key genuinely solves its own problem: a second POST /payments with the same Idempotency-Key header gets the stored response from the first attempt instead of a second charge. That's real protection, and it's not in question here. What it protects is the boundary between the client and your API — it says nothing about the boundary one layer further in, between your API and the payment provider it calls internally.

That inner boundary is a genuinely different kind of problem for one structural reason: the payment provider cannot participate in your database transaction. Your COMMIT and the provider's own decision to charge the card are two separate events, on two separate systems, connected by nothing more reliable than a network call in between. There is no version of BEGIN; write the order; call Stripe; COMMIT; that makes both of those atomic together — the moment a network hop is involved, the two-phase-commit guarantee any single database gives you for free is gone, and you're back to exactly the same ambiguity the client faced, just one level deeper in the stack. Wrapping the external call inside the same database transaction doesn't fix this either — a transaction can't "roll back" a charge that a third party's server already processed, and holding a database transaction open while waiting on a slow network call is its own separate problem (locked rows, exhausted connection pools) that tends to make things worse, not safer.

An idempotency key stops your client from causing a duplicate. It does nothing to stop your own retry logic from causing one against a system that isn't yours.

#05This Isn't Just a Payments Problem

Payments make the clearest example because the cost of a duplicate is immediate and visible — a customer sees two charges and complains the same day. But the exact same "commit, then call something outside your control" shape shows up anywhere a system needs to trigger an action it can't take back. Provisioning a cloud resource works this way: your API records "user requested a new virtual machine," then calls the cloud provider to actually create it — if that call times out, did the VM get created or not, and is retrying about to leave the user with two of them on their bill? Activating a SIM card on a telecom network is the same shape again: the request gets recorded, a call goes out to the network's provisioning system, and a lost response leaves the API unsure whether the SIM is live or still pending. Sending a one-time password by SMS, kicking off a background export job, calling a third-party fraud-check API before approving a loan — every one of these has a database write on one side and a call to something you don't control on the other, and every one of them has the exact same gap in the middle. The specific fix in this piece was built around a payment example because it's the most familiar one, but the pattern itself applies anywhere that shape exists.

#06What "Indeterminate" Actually Means

Stripe's own error-handling documentation is unusually direct about naming this state instead of hand-waving past it. Under connection errors — a network problem between your server and Stripe — its guidance isn't "assume it failed and retry," and it isn't "assume it succeeded and move on." It's a third instruction most integrations skip entirely:

That's the whole shape of the correct answer, stated plainly by the company with the most production experience watching this exact failure happen at scale: a lost response is not evidence of failure. It's an absence of evidence, and the only way to convert an absence of evidence into an actual answer is to go ask — either by polling the provider's own record of what happened, or by waiting for the provider to tell you asynchronously via a webhook. Stripe's own idempotency-key mechanism is explicitly framed around this same state: retry with the same key "until you receive a clear success or failure," not until you get any response at all. Notice what that instruction actually implies — it assumes your own system has somewhere to keep retrying from, a record that survives between attempts. Without that record, "keep retrying until you get a clear answer" isn't something your code can even do.

"I don't know yet" is a legitimate state for a distributed system to be in. Pretending it isn't is where duplicate charges come from.

“"Treat the result of the API call as indeterminate. That is, don't assume that it succeeded or that it failed. To find out if it succeeded, you can: Retrieve the relevant object from Stripe and check its status. Listen for webhook notification that the operation succeeded or failed."”

#07The Fix Starts Before the External Call, Not After

The mistake that makes this hard to recover from isn't the network failure — that's just the environment every distributed system has to live in. The mistake is not having anywhere durable to record "I was about to do this" before the risky call goes out. If the only record of an attempted charge lives in a local variable that dies with the crashed process, there is nothing left to reconcile against once the process comes back — the system genuinely has no memory of having tried, and the next request that arrives looks exactly like the first one ever made.

The fix is a durable, three-state (at minimum) record, written to the database in the same transaction as everything else the request needs to commit, before the external call is made at all:

The ordering here is the entire trick: PROCESSING gets written and committed before the request to the payment provider goes out, not after. That single ordering decision is what turns a crash mid-call from "we have no idea what happened" into "we know exactly what we don't know" — the next process to look at that row can see a PROCESSING attempt with no resolution yet, and knows precisely what it needs to go check, rather than staring at a gap with no record of an attempt ever having been made at all. This is a small, deliberate trade: one extra database write on every attempt, in exchange for never again facing a request with zero history behind it.

StateMeaningSafe to retry the external call?
PENDINGAttempt recorded, external call not yet madeYes — nothing was sent
PROCESSINGExternal call in flight, outcome unknownNo — must reconcile first
SUCCEEDED / FAILEDOutcome confirmedNo — return the stored result

#08What This Looks Like in Code

The state machine above is easier to trust once it's not just a table but something closer to real code. The sketch below leaves out real error handling and real database calls for clarity, but the ordering of operations — write PENDING, write PROCESSING, then call out, then resolve — is the actual point, and that ordering has to survive intact no matter what language or framework it's written in.

Two details in that sketch matter more than the rest of the code around them. First, a NetworkError never writes FAILED directly — writing FAILED on a connection error is exactly the mistake this whole piece is arguing against, because the charge might have actually succeeded on the provider's side. Second, a retry that lands while the row is still PROCESSING doesn't make a second call to the provider at all — it goes straight to reconcile(), which is the only code path allowed to move a PROCESSING row forward.

function handlePayment(idempotencyKey, amount):
existing = db.findAttempt(idempotencyKey)
if existing and existing.state in [SUCCEEDED, FAILED]:
return existing.result # already resolved — return it, call nobody
if existing and existing.state == PROCESSING:
result = reconcile(existing) # ask the provider what actually happened
return result
attempt = db.createAttempt(idempotencyKey, amount, state=PENDING)
db.updateAttempt(attempt.id, state=PROCESSING) # committed before the risky call
try:
response = paymentProvider.charge(amount, idempotencyKey)
db.updateAttempt(attempt.id, state=SUCCEEDED, result=response)
return response
catch NetworkError:
# do NOT mark this FAILED — the charge may have gone through anyway
return { status: "processing", message: "check back shortly" }

#09Reconciliation: The Piece Most Systems Skip

A PROCESSING row that nothing ever looks at again is just a more detailed way of being stuck. The state machine above only pays for itself once something actually resolves that row — either a background reconciliation job that periodically asks the payment provider "what actually happened to attempt X," or a webhook handler that receives the provider's own asynchronous answer and writes it back. Either path converts PROCESSING into a real SUCCEEDED or FAILED, and either path has to run whether or not the client ever comes back to retry — a client that gives up and never retries at all shouldn't leave a payment permanently stuck in limbo, quietly consuming support tickets weeks later when someone finally notices.

One detail here is easy to get backwards: the reconciliation step, the thing whose entire job is fixing an uncertain state, has to be idempotent itself. If reconciliation discovers a successful charge that was never recorded and "fixes" it by applying the update, that fix has to be safe to run more than once too — a reconciliation job that isn't careful here can process the same discovered outcome twice and create the exact duplicate-charge bug it was built to prevent, just one layer further downstream. Idempotency isn't a property you add once at the front door; it has to hold at every boundary where a retry — automatic or manual, client-driven or system-driven — can plausibly happen.

A PROCESSING state that nothing ever resolves isn't safer than not tracking state at all. It's just a more precisely documented way of never finding out.

#010Back to the Original Sequence, With an Answer

Walk the same failure through a system built this way, and the ambiguity at the end of the chain closes. The client's POST /payment still reaches the API the same way. The API still writes to its own database before doing anything risky — except now it writes twice: a PENDING row first, then a PROCESSING row immediately before calling the payment provider, both committed. The call to the provider still goes out, and the network still fails exactly the same way it did before. The client still sees nothing but a failed request and still retries with the same idempotency key. But this time, when that retry reaches the API, there's a real row to consult instead of a guess to make: it sees PROCESSING, reconciles with the provider to find out what actually happened, resolves that row to SUCCEEDED or FAILED, and returns the real answer.

Nothing about the network got more reliable. The failure still happens exactly the same way it always did. What changed is that the API's retry path now has a real, durable answer to consult instead of a decision to guess at — the same request that used to arrive with no way to answer it now arrives at a specific row that says exactly what's still unresolved and exactly what to go check before doing anything else.

The same failure, with a durable answer to consult

Client

POST /payment

reaches the API

API

writes a PENDING row

committed before the risky call

API

writes PROCESSING, then calls the provider

on the outbound leg

Network

drops the same way

sees a failed request

Client

retries with the same idempotency key

finds a real row to consult

API

reconciles PROCESSING to SUCCEEDED or FAILED

#011Designing for This from Day One

None of this requires a rewrite — it requires deciding these things before the first external call goes out, not after the first duplicate-charge ticket comes in. Treat the list below as the minimum bar for any code path that writes to your own database and then calls something you don't control.

  • Write the attempt's state to your own database, in the same transaction as everything else that request needs, before calling an external system that can't participate in that transaction.
  • Treat a timeout or connection error from that external system as indeterminate, never as a failure — Stripe's own guidance names this explicitly, and it generalizes to any external call your system doesn't control.
  • Where the provider supports it, give the call two independent paths back to a resolved state: an active reconciliation check (poll the provider) and a passive one (a webhook/callback handler) — not every provider offers both, so the active poll is the one path to build regardless, with the passive webhook added wherever it's available.
  • Make the reconciliation path itself idempotent — it's the thing responsible for fixing an inconsistent state, and if it isn't safe to run twice, it can recreate the exact bug it exists to prevent.
  • Never let PROCESSING be a state with no expiry or escalation path — an attempt stuck there past a reasonable window is an incident, not a background detail, and should page someone rather than sit silently.
  • Keep the client-facing idempotency key and the internal reconciliation state as two separate mechanisms solving two separate boundaries — collapsing them into one hides exactly the layer this piece is about.

#012The Deeper Lesson

The client-facing idempotency key gets most of the attention because it's the boundary a customer notices — a duplicate charge on a bank statement is impossible to ignore. But it protects exactly one hop in a chain that usually has at least two, and the hop it doesn't protect is the one where your own system hands control to something it doesn't own and can't make atomic with its own commit. A system that's rigorous about the client-facing key and silent about everything past its own database has only solved the half of this problem a customer can see — the other half just fails less visibly, in a support queue or a reconciliation report weeks later instead of on a bank statement the same day.

The database commit was never the hard part. It's transactional, durable, and done the instant it happens. Everything that comes after it — the call your system makes next, on the strength of a fact only your database is sure of yet — is where the real discipline has to live.

#013Sources

Every technical claim above is traceable to the specific documentation below, not to general commentary about distributed systems.

  • Stripe — Error handling: Connection errors — Stripe's own guidance to treat a network failure as indeterminate, and to resolve it via object retrieval or webhook rather than assuming success or failure
  • Stripe — Idempotent requests — the idempotency-key mechanism this piece builds past, including "retry with the same key until you receive a clear success or failure"
  • MDN — Idempotent and RFC 7231 §4.2.2 — the formal definition of idempotency this piece's state-machine design has to satisfy at the reconciliation step, not just at the client-facing endpoint

Found this useful? Share it

Have a technical response or architectural perspective to share with the engineering desk?

Submit Engineering Feedback