DEV Community

Cover image for To Retry or Not to Retry? That Is the Question.
Daniel Balcarek
Daniel Balcarek Subscriber Community Curator

Posted on

To Retry or Not to Retry? That Is the Question.

Kaggle Benchmarking Challenge Submission

This is a submission for the Kaggle Benchmarking Challenge

Those who read my articles know that a lot of them are actually benchmarks of something. Mostly .NET related, but still, someone could say that this challenge should be pretty close to what I usually do.

The opposite was true.

Building a code benchmark and benchmarking AI models are two different worlds, so when I first saw this challenge, I had no idea what exactly I should benchmark. Comparing models on coding tasks felt too generic, and I didn't want to create a benchmark just for the sake of having one.

Then I looked at the topics of my last couple of articles. A lot of them were about APIs, resilience, failures, and how systems behave when something goes wrong. And that gave me an experiment idea: What if I benchmark AI models on one very simple question: To Retry or Not to Retry?

A 503 Service Unavailable does not automatically mean that retrying is safe. A POST request may already have been processed. An idempotency key can completely change the answer. A timeout may happen before the server receives anything or after it has already changed some state.

So the HTTP status code alone is often not enough. The model has to understand the whole situation.

What I Benchmarked

The idea is simple. I prepared several API failure scenarios containing information about the request, the response, and some additional context. The model has to return two things:

  • a decision: YES, NO, or YES_AFTER_DELAY
  • a one-sentence explanation of why

For example:

POST /payments
503 Service Unavailable

An idempotency key was supplied and the API guarantees
duplicate requests with the same key are not processed twice.

Retry?

Enter fullscreen mode Exit fullscreen mode

A developer will probably immediately say:

YES_AFTER_DELAY

The important part is not only the 503. The context tells us that an idempotency key was supplied and repeated requests with the same key will not process the payment twice. But will an AI model notice the same thing? And more importantly, what happens when the scenario is less obvious?

For this small experiment, I prepared scenarios where the answer depends on details such as HTTP method, status code, idempotency, rate limiting etc.

Every model gets the same response format:

Decision: YES | NO | YES_AFTER_DELAY
Reason: <one sentence>
Enter fullscreen mode Exit fullscreen mode

For scoring, I only use the Decision. This keeps the benchmark simple and deterministic. The Reason does not affect the score. I collect it because it can show whether the model actually understood the scenario or simply arrived at the correct answer for the wrong reason. That also makes the benchmark easy to compare between models. Each scenario has an expected decision, so the final result can simply be calculated as the percentage of retry decisions the model got right.

But why benchmark this at all?

Imagine that you are integrating an external service and want to make the call more resilient. In today's world of AI-assisted coding, there is a very good chance that a coding agent will do the work. The AI can easily generate a retry policy. The more interesting question is: Will it retry the right requests? Because retrying something that should not be retried can be much worse than not retrying at all.

That's what I wanted to measure. Now let's have a look at the scenarios.

Scenarios

I prepared 14 scenarios, each based on an actual task used in the Kaggle benchmark. Originally, all of them were part of the article, but that made it too long for a small experiment. So I decided to move the full scenarios into a standalone app, where you can browse every task together with its expected answer and explanation. And if you want to try them yourself first, there is also a test for humans. After all, why not benchmark some humans too? 😁

Try the human benchmark or browse all scenarios: To Retry or Not to Retry?

Here are the scenarios used in the experiment:

  • Safe Payment Retry - A payment request fails with 503 Service Unavailable, but an idempotency key guarantees that retrying cannot create a duplicate payment.
  • Unsafe Payment Retry - A payment request returns 503 Service Unavailable with Retry-After, but there is no idempotency key, so retrying could create a duplicate charge.
  • Idempotent PUT - A PUT request is sent successfully, but the connection is lost before the response arrives.
  • Rate-Limited Order - An order request is rate-limited before processing and returns 429 Too Many Requests with Retry-After.
  • Payment Details Conflict - A multi-step payment flow reaches /payments/details, which returns 409 Conflict with transient-error: false.
  • Non-Idempotent PATCH - A PATCH request increases inventory, but the connection is lost before the client knows whether the change was already applied.
  • Idempotent DELETE - A session deletion request is sent, but the connection is lost before the response arrives.
  • Service Unavailable - A GET request receives 503 Service Unavailable together with a Retry-After header.
  • Stale Resource Version - A resource is updated by another client, so a PATCH using an old ETag fails with 412 Precondition Failed.
  • Rate Limit Without Retry-After - A safe GET request repeatedly receives 429 Too Many Requests, but the server does not provide a Retry-After header.
  • I'm a Teapot - A coffee request receives the legendary 418 I'm a Teapot response.
  • Eventual Consistency - A newly created resource immediately returns 404 Not Found when another service tries to use it.
  • Expired Filter Workflow - A temporary filter times out several times and eventually returns 404 Not Found, so the original filter can no longer be used.
  • Cached Cart Options - A request for cart options fails with 504 Gateway Timeout, but a recently expired cached response is still available through stale-if-error.

Models Tested

Picking the models was pretty straightforward. First, I picked models I use a lot while working: GPT-5.6 Luna and GPT-5.6 Sol. From my experience, Luna is very efficient. It makes more mistakes than Sol, but it also uses significantly fewer tokens.

Then I added Gemini 3.7 Flash and Gemini 3.1 Pro Preview. I’ve used both quite a lot for text generation and CSS styling.

I also included Claude Sonnet 5 and Claude Haiku 4.5. I used Claude a lot for coding before, but these days I’ve mostly switched to GPT-5.6 because it feels more efficient for my workflow.

Lastly, I added GLM-5 and DeepSeek-R1 mostly out of curiosity to see how well they would perform.

I also tried to include Grok, but I immediately hit 404 Not Found errors with both Grok 4.5 and Grok 4.6, so they are not included in the results.

The goal wasn’t to test every available model, but to compare a mix of models I already use with a few I was curious about.

Findings

Okay, first let's have a look at the leaderboard and the results from the first run.

Kaggle leaderboard

We can see that the best model, with a score of 1.00, was GPT-5.6 Sol, followed by Gemini 3.7 Flash with one mistake. Then came four models with two mistakes each, DeepSeek-R1 with three, and Claude Haiku 4.5 finished last with five.

When we look closer, we can see that most of the models made a mistake in the Unsafe Payment Retry task. The task itself is tricky because of the Retry-After header, but repeating that request could result in a duplicate charge. Only GPT-5.6 Sol and Gemini 3.1 Pro Preview answered correctly. The rest of the models returned something similar to:

Decision: YES_AFTER_DELAY
Reason: The server explicitly requests waiting 10 seconds before attempting the request again.
Enter fullscreen mode Exit fullscreen mode

So most of the models fell into the trap of following the obvious HTTP signal, even though it conflicted with the wider context. Retry-After says “retry,” but a payment without idempotency says “maybe don't.” And to be honest, I am glad that most of the models failed. A small win for the author, who can still trick AI sometimes. 😄

I was also surprised that some models failed on tasks with idempotent HTTP methods: Idempotent PUT and Idempotent DELETE. But they did not fail completely. Most of them answered YES_AFTER_DELAY because they assumed a temporary network issue that could be resolved after a delay.

These YES versus YES_AFTER_DELAY cases show that a model can understand that retrying is safe but disagree about when to retry. That is different from misunderstanding the scenario. For version 2, I could therefore allow multiple valid Decision values for some scenarios. So in these cases, AI actually trained me a little.

Three Runs, Not Just One

That was only the first run, but I didn't want to build the whole experiment around a single run, so I ran the same benchmark two more times to check model stability and collect more data.

Model Run 1 Run 2 Run 3 Avg.
GPT-5.6 Sol 14/14 14/14 13/14 13.67
Gemini 3.7 Flash 13/14 13/14 13/14 13.00
Gemini 3.1 Pro Preview 12/14 12/14 13/14 12.33
GLM-5 12/14 13/14 12/14 12.33
Claude Sonnet 5 12/14 12/14 12/14 12.00
GPT-5.6 Luna 12/14 12/14 12/14 12.00
DeepSeek-R1 11/14 10/14 10/14 10.33
Claude Haiku 4.5 9/14 9/14 9/14 9.00

Overall, the results were actually quite stable. Out of the 112 model-scenario combinations (8 models × 14 scenarios), 102 had the same correct/incorrect outcome in all three runs, which is about 91%. But the interesting part is that the same score did not always mean the same behavior. Some models were stable in the number of correct decisions while changing which scenarios they got wrong. A score of 12/14 in all three runs does not necessarily mean that the model made the same decisions every time.

This was especially visible in Expired Filter Workflow, where four models changed their result between runs: GPT-5.6 Sol, Gemini 3.1 Pro Preview, GLM-5, and DeepSeek-R1. Even GPT-5.6 Sol, which scored 14/14 in the first two runs, changed its answer in the third run from YES to YES_AFTER_DELAY. It still decided that the request should be retried, but disagreed about when. This is exactly the type of scenario that made me question whether exact YES versus YES_AFTER_DELAY scoring is always the right approach.

The additional runs also made the Unsafe Payment Retry finding much stronger. It was answered correctly only 6 times out of 24 attempts, just 25%. Even more interestingly, the result was completely consistent across runs. GPT-5.6 Sol and Gemini 3.1 Pro Preview got it right all three times, while the other six models missed it in every run. So this looks less like random model variation and more like a systematic trap in how the models interpreted the scenario.

On the other hand, seven of the 14 scenarios were answered correctly in all 24 attempts. The straightforward retry cases were therefore not really what separated the models. The differences started to appear when retry timing, wider context, or multi-step state became important.

Model Efficiency

Now let's have a look at efficiency. I will use the graph from the third run because the models ended up in roughly the same areas of the chart across all three runs.

To Retry or Not to Retry?: Score vs. Total Cost Pareto

GPT-5.6 Luna seems to be the most efficient model, and I am not surprised. I mostly use it at work because it is cheap and usually gets me close to the final solution.

Claude Haiku 4.5 is also in the efficient corner, but it had the worst results of all the tested models and was still more expensive than GPT-5.6 Luna.

The rest of the models are in the upper-right corner, and I was really surprised that DeepSeek-R1 ended up as the second most expensive model. So I looked at its output and found that DeepSeek-R1 answered with anywhere from 4,000 to 9,000 characters, even though it was supposed to answer only with Decision and Reason. All of the other models followed that instruction, so in this case the cost difference wasn't only about token pricing, but also about how well the model followed the requested output format.

What I Learned

When I designed the experiment, I was actually worried that the scenarios would be too easy for today's AI models. The results showed that this is not always the case. Yes, I could definitely improve the benchmark, especially the ambiguity between YES and YES_AFTER_DELAY. But even the strongest models still made mistakes once the decision depended on more than just the obvious HTTP signal. Retry decisions in real systems are not always simple. AI can definitely help us reason through the difficult cases, but the wider context still matters, and blindly following the model's answer can be risky.

My Benchmark

To Retry or Not to Retry? Benchmark

Kaggle was completely new to me and, to be honest, I am glad that I discovered it through the DEV.to challenge. I expected the setup to be much more complicated, but creating this small experimental benchmark directly in the UI was surprisingly straightforward. I will definitely explore Kaggle more, especially after seeing how much there is to explore beyond traditional ML competitions.

Top comments (32)

Collapse
 
sinarezaei profile image
Sina Rezaei •

One thing I'd add to the benchmark is that “should I retry?” is sometimes the wrong unit of evaluation. The more interesting question is “how much retry capacity does this logical request still have?”

Take a payment workflow behind an API gateway, order service, and payment provider. The payment call might have a 2s remaining deadline when it reaches the provider. If the payment client independently decides it has 3 retries left, those retries can already be pointless because the original request has almost no budget left.

I’d make the benchmark carry a request deadline and retry budget through the scenario:

logical_request_deadline = T+2s
elapsed = 1.4s
remaining_budget = 600ms
attempts_used = 2
retryable = true
Enter fullscreen mode Exit fullscreen mode

A model might correctly identify the error as transient and still make the wrong operational decision by retrying after a backoff that exceeds the remaining deadline.

The same gets nastier with layered retries. If the gateway, service, and downstream client each have their own retry policy, a single logical request can fan out into a lot more physical attempts. Google SRE calls out exactly this amplification problem, along with deadline propagation and retry budgets.

That could make a really interesting v2 dimension: not just whether the model picked the correct retry decision, but whether it preserved the request's remaining latency and retry budget across the whole call chain.

Collapse
 
gramli profile image
Daniel Balcarek •

That's a really good point, I actually have something slightly similar in my Cached Cart Options scenario, where the decision depends on cache expiration and stale-if-error, but your example definitely takes it a step further.

I really like the idea that a retry might be perfectly valid in isolation but still cause problems for the whole request. Definitely an interesting direction for v2. Thanks for the detailed suggestion!

Collapse
 
sinarezaei profile image
Sina Rezaei •

Yeah, I think that’s where the v2 could get really interesting.

I’d probably make the benchmark distinguish between retry safety and retry viability. A retry can be completely safe from an idempotency perspective, but still be a bad decision because the logical request no longer has enough budget to complete.

For example, imagine:

API Gateway
  -> Order Service
      -> Payment Service
          -> PSP
Enter fullscreen mode Exit fullscreen mode

The gateway gives the request a 3s deadline. By the time it reaches the PSP, only 700ms remain. The payment operation is idempotent and the PSP returns a transient 503, so retrying is technically safe. But if the retry policy waits 1s before the next attempt, the correct decision should still be NO, because the retry cannot complete within the remaining deadline.

That also gives you a nice way to test retry amplification. If each layer independently gets 3 attempts, the logical request can turn into 27 physical attempts before reaching the dependency. Google’s SRE guidance explicitly calls out this kind of multiplicative retry behavior and recommends propagating deadlines and controlling retries at the appropriate layer.

So for v2, I’d love to see the model given something like deadline_remaining, attempts_used, retry_budget_remaining, and expected_backoff, then score not only whether it retries, but whether the retry is actually viable for the logical request.

That would make the benchmark less about recognizing retryable errors and more about reasoning about the request as a whole.

Collapse
 
arhancanli profile image
Arhan Canli •

The Unsafe Payment Retry row is the most informative one, but check how its label is defined before reading it as a model error. A 503 usually comes from a load balancer or an overloaded tier that never ran the handler, so YES_AFTER_DELAY with a Retry-After header is a defensible answer; NO is the conservative policy answer because a 503 doesn't prove the handler didn't run. The Reason lines are where you can see which models knew that and which just followed the header.

I'd split the scoring in two. One column for "would this have caused a duplicate effect" (the NO class, where a miss is costly), and one for YES vs YES_AFTER_DELAY, which is only about timing. Right now a model that says YES on the idempotent DELETE and one that says YES_AFTER_DELAY on an unsafe POST lose the same single point, though only one of them could charge a customer twice.

On the leaderboard: with 14 scenarios, one item is 7 points. The 95% Wilson interval for 14/14 is roughly 78-100%, for 12/14 roughly 60-96%, and even Haiku's 9/14 is about 39-84%. So the ordering between models is not separated by this run; the per-scenario pattern (which items the models miss in common) is the real result. Running each scenario 5 times per model would also show whether the Retry-After miss is stable or a coin flip for the models that got it right.

Collapse
 
gramli profile image
Daniel Balcarek •

Thanks for the detailed feedback. I agree that separating retry safety from retry timing would make the benchmark much more meaningful. It's something I started thinking about after seeing the YES vs YES_AFTER_DELAY results.

For the payment scenario, I still think NO is the right answer when there's no idempotency guarantee. A 503 doesn't prove the payment was processed, but it doesn't prove the opposite either, and that's exactly the risk I wanted to test.

I actually ran the benchmark three times and interestingly, the Unsafe Payment Retry results were completely consistent across all three runs. More runs would definitely be better, but this is still a small experiment anyway.

Definitely some good ideas here for a potential v2!

Collapse
 
arhancanli profile image
Arhan Canli •

Agreed on NO for the payment case: with no idempotency key, a 503 is ambiguous about whether the charge went through, so a blind retry can double-charge.

One caution on the three identical runs: if the settings were the same, matching results show the model is stable on that prompt, not that the answer is right. That is still useful for the payment item, since a stable NO is what you want there. For v2, a variant of the same scenario with an Idempotency-Key header present would show whether the model flips to YES only when it should.

Thread Thread
 
gramli profile image
Daniel Balcarek •

Not exactly the same scenario (the Retry-After header is missing), but I do have both Safe Payment Retry and Unsafe Payment Retry in the benchmark. Still, testing more variations could be interesting for v2.

Collapse
 
sylwia-lask profile image
Sylwia Laskowska •

Wow, what a great article!!! Not only is it incredibly practical for engineers and architects, but it also highlights something bigger: nope, we still can't blindly trust coding agents, no matter what some people try to convince us of. 😅

Now I'm seriously tempted to send this to a few self-proclaimed "experts". 🤣

Collapse
 
alexcodebytes profile image
Oleksandr •

Spot on alignment, Sylwia. Shifting from blind faith to rigorous architectural inspection is exactly how we prevent critical system failures. 🛠️ AI is an awesome compiler assistant, but a terrible independent software architect.

Collapse
 
gramli profile image
Daniel Balcarek •

Oh, thank you! Glad you liked it, Sylwia.

Yep, totally agree. Kaggle doesn't allow all the latest models, but even the best ones can make mistakes, so a good code review is still necessary.

I think the benchmark has some practical value, but this first version still has its limitations. The YES vs YES_AFTER_DELAY decisions are debatable, and more runs would help too. So I'm sure those self-proclaimed experts would find something to argue about. 🤣

Btw, how was Frontkon? Did you become even more famous? 😁

Collapse
 
sylwia-lask profile image
Sylwia Laskowska •

Hahahaha, I have no idea if I'm any more famous now! 🤣 But apparently, the head of the conference has appointed me Frontkon's ambassador to Poland, and my mission for next year is to bring 60 Polish girls with me! 🤣🤣🤣

Thread Thread
 
gramli profile image
Daniel Balcarek •

60?! That's a huge task! 🤣 That's practically a medium-sized company! 🤣🤣

Thread Thread
 
sylwia-lask profile image
Sylwia Laskowska •

Hahaha, let's leave those 60 women aside for now! (Although I have to admit, they could probably use a few more women there! 😅)

I was just sitting through an incredibly boring call where a backend developer was literally doing live debugging in front of everyone xDDD (and it wasn't going particularly well).

Anyway, that somehow got me thinking about your article, AI agents, and conference CFPs, and I realized this topic could actually evolve into a pretty awesome talk for KubeCon! They have a co-located event called Agentics Day: MCP + Agents.

So, if you're into public speaking at all, drop me a message on LinkedIn! 😁

Thread Thread
 
gramli profile image
Daniel Balcarek •

Wow, I definitely didn't expect my Retry article to lead to a potential KubeCon talk! 🤣

But seriously, I'm really honored that you'd consider me as a potential speaker and think my articles could make a good conference talk! Are you referring to my WebMCP game articles? Public speaking at conferences is actually one of my future goals, but to be honest, I'm not sure I'm ready yet. I still need to work on my public speaking skills a bit, especially since I'm not an extrovert at all😅

BUT I'd definitely love to hear more about your idea, though.

Thread Thread
 
sylwia-lask profile image
Sylwia Laskowska •

Hahaha, just to be clear, I'm not on the KubeCon program committee or anything! 🤣 I just had an idea that I think could make a really interesting talk. But maybe I'll tell you more about it privately, before someone steals it! 😂

Thread Thread
 
gramli profile image
Daniel Balcarek •

Yeah, I get it, but I'm still honored! 😁

Okay, I'll message you on LinkedIn, because thieves are everywhere! 🤣

Collapse
 
sinarezaei profile image
Sina Rezaei •

Exactly. 😅 What I find interesting here is that the agent can produce perfectly reasonable-looking retry logic and still make the wrong architectural decision.

The dangerous part isn't necessarily bad code. It's code that looks correct because it follows an obvious signal, while missing the state and constraints around the request.

That's where benchmarks like Daniel's are useful. They test the reasoning behind the code, not just whether the generated code compiles. And yes, I can already think of a few “AI replaces engineers” arguments that would get interesting after running this benchmark. 😂

Collapse
 
koda2026 profile image
Harun - solo dev •

daniel, this is a brilliant angle for an ai benchmark. testing models on contextual reasoning (idempotency, http methods, state) rather than just raw code generation is exactly what the industry needs right now.

your "unsafe payment retry" trap is the perfect example of why blind ai automation is dangerous. as someone building an ai coding ecosystem (koda), this is my biggest fear: an agent confidently executing a non-idempotent mutation just because it saw a retry-after header, completely ignoring the lack of an idempotency key.

i also completely agree with @arhancanli's point about splitting the scoring. a model guessing the wrong timing (yes vs yes_after_delay) is a minor flaw, but a model causing a duplicate charge is a catastrophic failure. weighting safety heavily over timing would make this benchmark even more powerful for v2.

thanks for proving that good old-fashioned architectural principles (like idempotency and deterministic fallbacks) still trump raw model size! 🐯

Collapse
 
gramli profile image
Daniel Balcarek •

Thanks! Yep, as we discussed with @arhancanli, separating safety from timing would make the benchmark more meaningful. Definitely something worth exploring in v2.

Collapse
 
koda2026 profile image
Harun - solo dev •

glad we're on the same page! separating safety from timing is going to make v2 incredibly valuable for the community. looking forward to seeing the results. keep up the great work! 🐯

Collapse
 
webdeveloperhyper profile image
Web Developer Hyper •

It seems that expensive AI models are not only expensive but also produce good results. Nice challenge! 😀

Collapse
 
gramli profile image
Daniel Balcarek •

That's true! But for everyday coding tasks, efficiency might be more valuable than getting the absolute best results. 😁

Collapse
 
alexcodebytes profile image
Oleksandr •

Benchmarking AI on system resilience logic like idempotency keys is an elite-tier experiment configuration. Most models can regurgitate basic code syntax, but they completely throw a kernel panic when calculating the edge-case risks of duplicate POST payloads on a 503 error. 🚀

Collapse
 
gramli profile image
Daniel Balcarek •

Thanks!

Collapse
 
khaithisran profile image
thisran •

This is such a good question because retry logic looks so simple until you put it inside a real system 😂. I really like the point about retrying being more than just “try again” context, remaining time, cost, and whether the operation is even safe to repeat can completely change the answer. Definitely one of those topics that seems straightforward until production starts having opinions. Really enjoyed this one!

Collapse
 
gramli profile image
Daniel Balcarek •

Yep, production always has the final say! 😂 Thanks for reading!

Collapse
 
xulingfeng profile image
xulingfeng •

Hmm, why are they still using DeepSeek-R1? It's a bit outdated. They should try the latest DeepSeek-V4.1 Flash.
and Good Luck!

Collapse
 
gramli profile image
Daniel Balcarek •

Unfortunately, Kaggle only offers DeepSeek-R1 and thanks! 😁

Collapse
 
mrsaynothing profile image
Mr Say Nothing •

The retry decision gets scarier one layer down: the timeout-after-server-applied case, where the retry succeeds twice. We ended up treating idempotency as a precondition for retrying at all — no idempotency key, no retry, fail loud instead. Curious whether your benchmark cases covered the double-apply failure or stayed on the should-you-retry side.

Collapse
 
hemapriya_kanagala profile image
Hemapriya Kanagala •

I didn’t realize there was so much to think about when deciding whether to retry a request 😄The results were interesting too. All the best with the competition, Daniel 😀

Some comments may only be visible to logged-in visitors. Sign in to view all comments.