DEV Community

Cover image for To Retry or Not to Retry? That Is the Question.

To Retry or Not to Retry? That Is the Question.

Daniel Balcarek on October 08, 2026

This is a submission for the Kaggle Benchmarking Challenge Those who read my articles know that a lot of them are actually benchmarks of something...
Collapse
 
sinarezaei profile image
Sina Rezaei •

One thing I'd add to the benchmark is that β€œshould I retry?” is sometimes the wrong unit of evaluation. The more interesting question is β€œhow much retry capacity does this logical request still have?”

Take a payment workflow behind an API gateway, order service, and payment provider. The payment call might have a 2s remaining deadline when it reaches the provider. If the payment client independently decides it has 3 retries left, those retries can already be pointless because the original request has almost no budget left.

I’d make the benchmark carry a request deadline and retry budget through the scenario:

logical_request_deadline = T+2s
elapsed = 1.4s
remaining_budget = 600ms
attempts_used = 2
retryable = true
Enter fullscreen mode Exit fullscreen mode

A model might correctly identify the error as transient and still make the wrong operational decision by retrying after a backoff that exceeds the remaining deadline.

The same gets nastier with layered retries. If the gateway, service, and downstream client each have their own retry policy, a single logical request can fan out into a lot more physical attempts. Google SRE calls out exactly this amplification problem, along with deadline propagation and retry budgets.

That could make a really interesting v2 dimension: not just whether the model picked the correct retry decision, but whether it preserved the request's remaining latency and retry budget across the whole call chain.

Collapse
 
gramli profile image
Daniel Balcarek •

That's a really good point, I actually have something slightly similar in my Cached Cart Options scenario, where the decision depends on cache expiration and stale-if-error, but your example definitely takes it a step further.

I really like the idea that a retry might be perfectly valid in isolation but still cause problems for the whole request. Definitely an interesting direction for v2. Thanks for the detailed suggestion!

Collapse
 
sinarezaei profile image
Sina Rezaei •

Yeah, I think that’s where the v2 could get really interesting.

I’d probably make the benchmark distinguish between retry safety and retry viability. A retry can be completely safe from an idempotency perspective, but still be a bad decision because the logical request no longer has enough budget to complete.

For example, imagine:

API Gateway
  -> Order Service
      -> Payment Service
          -> PSP
Enter fullscreen mode Exit fullscreen mode

The gateway gives the request a 3s deadline. By the time it reaches the PSP, only 700ms remain. The payment operation is idempotent and the PSP returns a transient 503, so retrying is technically safe. But if the retry policy waits 1s before the next attempt, the correct decision should still be NO, because the retry cannot complete within the remaining deadline.

That also gives you a nice way to test retry amplification. If each layer independently gets 3 attempts, the logical request can turn into 27 physical attempts before reaching the dependency. Google’s SRE guidance explicitly calls out this kind of multiplicative retry behavior and recommends propagating deadlines and controlling retries at the appropriate layer.

So for v2, I’d love to see the model given something like deadline_remaining, attempts_used, retry_budget_remaining, and expected_backoff, then score not only whether it retries, but whether the retry is actually viable for the logical request.

That would make the benchmark less about recognizing retryable errors and more about reasoning about the request as a whole.

Thread Thread
 
gramli profile image
Daniel Balcarek •

Thanks for taking the time to explain it in such detail, this is exactly the kind of discussion that makes publishing an experiment worthwhile.

I'm definitely saving this for v2.

Collapse
 
arhancanli profile image
Arhan Canli •

The Unsafe Payment Retry row is the most informative one, but check how its label is defined before reading it as a model error. A 503 usually comes from a load balancer or an overloaded tier that never ran the handler, so YES_AFTER_DELAY with a Retry-After header is a defensible answer; NO is the conservative policy answer because a 503 doesn't prove the handler didn't run. The Reason lines are where you can see which models knew that and which just followed the header.

I'd split the scoring in two. One column for "would this have caused a duplicate effect" (the NO class, where a miss is costly), and one for YES vs YES_AFTER_DELAY, which is only about timing. Right now a model that says YES on the idempotent DELETE and one that says YES_AFTER_DELAY on an unsafe POST lose the same single point, though only one of them could charge a customer twice.

On the leaderboard: with 14 scenarios, one item is 7 points. The 95% Wilson interval for 14/14 is roughly 78-100%, for 12/14 roughly 60-96%, and even Haiku's 9/14 is about 39-84%. So the ordering between models is not separated by this run; the per-scenario pattern (which items the models miss in common) is the real result. Running each scenario 5 times per model would also show whether the Retry-After miss is stable or a coin flip for the models that got it right.

Collapse
 
gramli profile image
Daniel Balcarek •

Thanks for the detailed feedback. I agree that separating retry safety from retry timing would make the benchmark much more meaningful. It's something I started thinking about after seeing the YES vs YES_AFTER_DELAY results.

For the payment scenario, I still think NO is the right answer when there's no idempotency guarantee. A 503 doesn't prove the payment was processed, but it doesn't prove the opposite either, and that's exactly the risk I wanted to test.

I actually ran the benchmark three times and interestingly, the Unsafe Payment Retry results were completely consistent across all three runs. More runs would definitely be better, but this is still a small experiment anyway.

Definitely some good ideas here for a potential v2!

Collapse
 
arhancanli profile image
Arhan Canli •

Agreed on NO for the payment case: with no idempotency key, a 503 is ambiguous about whether the charge went through, so a blind retry can double-charge.

One caution on the three identical runs: if the settings were the same, matching results show the model is stable on that prompt, not that the answer is right. That is still useful for the payment item, since a stable NO is what you want there. For v2, a variant of the same scenario with an Idempotency-Key header present would show whether the model flips to YES only when it should.

Thread Thread
 
gramli profile image
Daniel Balcarek •

Not exactly the same scenario (the Retry-After header is missing), but I do have both Safe Payment Retry and Unsafe Payment Retry in the benchmark. Still, testing more variations could be interesting for v2.

Collapse
 
sylwia-lask profile image
Sylwia Laskowska •

Wow, what a great article!!! Not only is it incredibly practical for engineers and architects, but it also highlights something bigger: nope, we still can't blindly trust coding agents, no matter what some people try to convince us of. πŸ˜…

Now I'm seriously tempted to send this to a few self-proclaimed "experts". 🀣

Collapse
 
alexcodebytes profile image
Oleksandr •

Spot on alignment, Sylwia. Shifting from blind faith to rigorous architectural inspection is exactly how we prevent critical system failures. πŸ› οΈ AI is an awesome compiler assistant, but a terrible independent software architect.

Collapse
 
sinarezaei profile image
Sina Rezaei •

Exactly. πŸ˜… What I find interesting here is that the agent can produce perfectly reasonable-looking retry logic and still make the wrong architectural decision.

The dangerous part isn't necessarily bad code. It's code that looks correct because it follows an obvious signal, while missing the state and constraints around the request.

That's where benchmarks like Daniel's are useful. They test the reasoning behind the code, not just whether the generated code compiles. And yes, I can already think of a few β€œAI replaces engineers” arguments that would get interesting after running this benchmark. πŸ˜‚

Collapse
 
gramli profile image
Daniel Balcarek •

Oh, thank you! Glad you liked it, Sylwia.

Yep, totally agree. Kaggle doesn't allow all the latest models, but even the best ones can make mistakes, so a good code review is still necessary.

I think the benchmark has some practical value, but this first version still has its limitations. The YES vs YES_AFTER_DELAY decisions are debatable, and more runs would help too. So I'm sure those self-proclaimed experts would find something to argue about. 🀣

Btw, how was Frontkon? Did you become even more famous? 😁

Collapse
 
sylwia-lask profile image
Sylwia Laskowska •

Hahahaha, I have no idea if I'm any more famous now! 🀣 But apparently, the head of the conference has appointed me Frontkon's ambassador to Poland, and my mission for next year is to bring 60 Polish girls with me! 🀣🀣🀣

Thread Thread
 
gramli profile image
Daniel Balcarek •

60?! That's a huge task! 🀣 That's practically a medium-sized company! 🀣🀣

Collapse
 
webdeveloperhyper profile image
Web Developer Hyper •

It seems that expensive AI models are not only expensive but also produce good results. Nice challenge! πŸ˜€

Collapse
 
gramli profile image
Daniel Balcarek •

That's true! But for everyday coding tasks, efficiency might be more valuable than getting the absolute best results. 😁

Collapse
 
koda2026 profile image
Harun - solo dev •

daniel, this is a brilliant angle for an ai benchmark. testing models on contextual reasoning (idempotency, http methods, state) rather than just raw code generation is exactly what the industry needs right now.

your "unsafe payment retry" trap is the perfect example of why blind ai automation is dangerous. as someone building an ai coding ecosystem (koda), this is my biggest fear: an agent confidently executing a non-idempotent mutation just because it saw a retry-after header, completely ignoring the lack of an idempotency key.

i also completely agree with @arhancanli's point about splitting the scoring. a model guessing the wrong timing (yes vs yes_after_delay) is a minor flaw, but a model causing a duplicate charge is a catastrophic failure. weighting safety heavily over timing would make this benchmark even more powerful for v2.

thanks for proving that good old-fashioned architectural principles (like idempotency and deterministic fallbacks) still trump raw model size! 🐯

Collapse
 
gramli profile image
Daniel Balcarek •

Thanks! Yep, as we discussed with @arhancanli, separating safety from timing would make the benchmark more meaningful. Definitely something worth exploring in v2.

Collapse
 
koda2026 profile image
Harun - solo dev •

glad we're on the same page! separating safety from timing is going to make v2 incredibly valuable for the community. looking forward to seeing the results. keep up the great work! 🐯

Collapse
 
alexcodebytes profile image
Oleksandr •

Benchmarking AI on system resilience logic like idempotency keys is an elite-tier experiment configuration. Most models can regurgitate basic code syntax, but they completely throw a kernel panic when calculating the edge-case risks of duplicate POST payloads on a 503 error. πŸš€

Collapse
 
gramli profile image
Daniel Balcarek •

Thanks!

Collapse
 
mrsaynothing profile image
Mr Say Nothing •

The retry decision gets scarier one layer down: the timeout-after-server-applied case, where the retry succeeds twice. We ended up treating idempotency as a precondition for retrying at all β€” no idempotency key, no retry, fail loud instead. Curious whether your benchmark cases covered the double-apply failure or stayed on the should-you-retry side.

Collapse
 
gramli profile image
Daniel Balcarek •

Yes, I have something similar in my Unsafe Payment Retry scenario. The request returns 503 with a Retry-After header, but there's no idempotency key and we don't know whether the payment was already processed. So it covers the risk of double application, but not an explicit double-apply failure. The benchmark focuses on whether the model recognizes that risk and decides not to retry.

Collapse
 
khaithisran profile image
thisran •

This is such a good question because retry logic looks so simple until you put it inside a real system πŸ˜‚. I really like the point about retrying being more than just β€œtry again” context, remaining time, cost, and whether the operation is even safe to repeat can completely change the answer. Definitely one of those topics that seems straightforward until production starts having opinions. Really enjoyed this one!

Collapse
 
gramli profile image
Daniel Balcarek •

Yep, production always has the final say! πŸ˜‚ Thanks for reading!

Collapse
 
bryan_hamilton profile image
Bryan Hamilton •

Master Yoda says... Do, or do not. There is no try.

Collapse
 
gramli profile image
Daniel Balcarek •

But never said, there retry is not.

Collapse
 
bryan_hamilton profile image
Bryan Hamilton •

Touche

Collapse
 
xulingfeng profile image
xulingfeng •

Hmm, why are they still using DeepSeek-R1? It's a bit outdated. They should try the latest DeepSeek-V4.1 Flash.
and Good Luck!

Collapse
 
gramli profile image
Daniel Balcarek •

Unfortunately, Kaggle only offers DeepSeek-R1 and thanks! 😁

Collapse
 
hemapriya_kanagala profile image
Hemapriya Kanagala •

I didn’t realize there was so much to think about when deciding whether to retry a request πŸ˜„The results were interesting too. All the best with the competition, Daniel πŸ˜€

Collapse
 
gramli profile image
Daniel Balcarek •

Haha, there's always more to it than it seems! πŸ˜„ Thanks for reading and for the good wishes! 😁