DEV Community

gentlyding
gentlyding

Posted on

Your Retry Just Charged the Customer Twice

"Just retry on failure" is the most expensive line of code you'll ship this year. It looks harmless — a network blip, a 500, you fire the request again, everyone's happy. Until the first request didn't actually fail. It timed out. The server got it, processed it, and died before it could tell you. Your retry is now a second charge, a second email, a second order.

This isn't a corner case. It's the default behavior of almost every HTTP client out of the box.

A timeout is not a "no"

The trap is conceptual. We treat a failed request as "the operation didn't happen," so re-running it feels safe. But a timeout means exactly one thing: you don't know what happened. The server may have:

  • rejected it before doing anything,
  • done the work and failed only on the response,
  • done the work and you'll never see the success because the connection dropped.

Only the last two cases hurt you on retry. And you can't tell them apart from the client. So "retry on timeout" is really "retry when the outcome is unknown" — which is precisely when re-executing is dangerous.

The honest statement a retry makes is: "I'm fine with this operation happening twice." If you're not fine with that, a bare retry is a bug.

The only real fix: an idempotency key

Idempotency means "doing it N times has the same effect as doing it once." The standard way to get there across a network is an idempotency key: a caller-supplied unique identifier for one logical operation. The server records it. If the same key comes back, the server returns the recorded result instead of running the work again.

POST /charges
Idempotency-Key: 9f1c2e3a-...      # generated by the client, one per charge attempt

# first call: server executes, stores {key -> 201 result}, returns it
# timeout on the client side, client retries with the SAME key
# second call: server finds the key, returns the STORED result, never charges again
Enter fullscreen mode Exit fullscreen mode

The key is the contract. As long as you resend the same key, the server guarantees one execution.

Where the key actually lives

The most common mistake is tying the key to the HTTP request instead of the intent. They are not the same thing.

  • A retry is the same intent sent again → must reuse the same key.
  • A new attempt by the user (clicked "pay" again after the page hung) → must be a new key, because it's a new intent.

If you generate the key inside the HTTP layer on every send, your retries get fresh keys and you're back to double-charging. The key has to be decided at the business level: one key per "thing the user is trying to accomplish," generated before any network call, then pinned to every retry of that attempt.

A practical shape: combine a stable business reference with a random component so collisions are astronomically unlikely:

key = f"{order_id}:{uuid4()}"     # or just uuid4() if you have no natural ref
Enter fullscreen mode Exit fullscreen mode

The business reference also helps you reason about it later ("why did this order get two different keys?").

What the server stores, and for how long

The server needs a small table: idempotency_key -> (status, stored_response, created_at). On an incoming request:

  1. If the key exists and is completed, return the stored response (HTTP status + body) as-is.
  2. If the key exists and is in_progress, you have a concurrent retry — return 409 Conflict (or lock and wait, depending on your tolerance).
  3. If the key is unknown, execute, store the result, return it.

The TTL matters. It must outlive the longest realistic retry window — if a client retries 30 seconds later, the key has to still be there, or you've lost the guarantee. But it can't live forever; pick something like 24 hours for payments, shorter for ephemeral actions. Past the TTL, the key expires and a late retry would re-execute — acceptable, because by then the retry storm is long over.

One subtlety: if the stored result was an error (say the first attempt returned 500 after partially failing), do you re-run on a replay? Usually no — return the stored error. The whole point is deterministic replay. If you re-execute on a stored 5xx, you've defeated the mechanism. Store the outcome, return the outcome.

Retries done right (because you still need them)

Idempotency handles the "did it run twice" question. Retries handle "will it eventually succeed." They are separate concerns, and you need both.

  • Backoff with jitter. Fixed-interval retries synchronize every client into a stampede the moment a dependency hiccups. Exponential backoff spreads them; jitter prevents them re-synchronizing on the next round.
  • Cap the attempts. Three to five, then stop and surface the failure. Infinite retry loops are how a minor outage becomes a permanent one.
  • Classify errors. A 4xx (validation, auth) will fail the same way forever — don't retry it. Retry only on timeouts, 5xx, and connection errors. And even then, only for operations that are idempotent or carry a key.
  • Never retry a non-idempotent write blindly. If there's no idempotency key in play, a write retry is a gamble. Either add the key or don't retry.

A minimal retry shape:

attempt = 0
while attempt < MAX:
    try:
        return call_with_key(req, key)     # key stays constant across the loop
    except (Timeout, ConnectionError, ServerError):
        attempt += 1
        sleep(base * 2**attempt + random_jitter())
raise RetryExhausted(req)
Enter fullscreen mode Exit fullscreen mode

Note the key is constant inside the loop. That's the whole trick.

The recursion footgun

The nastiest version is retry logic that retries itself: a wrapper around the HTTP client that retries, wrapped by another layer that also retries, wrapped by a queue consumer that redelivers on failure. Three independent retry policies multiply into "this runs 5×5×5 = 125 times," most of them after the operation already succeeded. If the call isn't idempotent, you've built a double-charge machine with three knobs.

Pick one retry boundary — usually the outermost — and make everything inside it assume the call may execute at most once per key.

Treat it as an API contract

Idempotency isn't a library you drop in. It's a contract between caller and server: the caller promises a stable key per intent; the server promises one execution per key. Design it at the boundary where side effects happen — payments, emails, provisioning, anything a user would notice happening twice. Everything else (backoff, jitter, error classification) is just hygiene around that contract.

Ship the retry. But ship the key first — otherwise you're not retrying, you're duplicating.

Top comments (1)

Collapse
 
launchgatecheck profile image
Launch Gate •

I'd qualify the "clicked pay again after the page hung -> new key" rule. If the first outcome is still unknown, the second click may be the same purchase intent, not a new charge. Would you reconcile the original attempt before allowing a fresh key? A useful regression case is: payment commits, response is lost, customer reloads and clicks again. The UI and server should agree whether that's recovery of the original order or an explicitly requested second purchase.