This is a submission for the Kaggle Benchmarking Challenge
Those who read my articles know that a lot of them are actually benchmarks of something...
For further actions, you may consider blocking this person and/or reporting abuse
One thing I'd add to the benchmark is that βshould I retry?β is sometimes the wrong unit of evaluation. The more interesting question is βhow much retry capacity does this logical request still have?β
Take a payment workflow behind an API gateway, order service, and payment provider. The payment call might have a 2s remaining deadline when it reaches the provider. If the payment client independently decides it has 3 retries left, those retries can already be pointless because the original request has almost no budget left.
Iβd make the benchmark carry a request deadline and retry budget through the scenario:
A model might correctly identify the error as transient and still make the wrong operational decision by retrying after a backoff that exceeds the remaining deadline.
The same gets nastier with layered retries. If the gateway, service, and downstream client each have their own retry policy, a single logical request can fan out into a lot more physical attempts. Google SRE calls out exactly this amplification problem, along with deadline propagation and retry budgets.
That could make a really interesting v2 dimension: not just whether the model picked the correct retry decision, but whether it preserved the request's remaining latency and retry budget across the whole call chain.
That's a really good point, I actually have something slightly similar in my Cached Cart Options scenario, where the decision depends on cache expiration and stale-if-error, but your example definitely takes it a step further.
I really like the idea that a retry might be perfectly valid in isolation but still cause problems for the whole request. Definitely an interesting direction for v2. Thanks for the detailed suggestion!
Yeah, I think thatβs where the v2 could get really interesting.
Iβd probably make the benchmark distinguish between retry safety and retry viability. A retry can be completely safe from an idempotency perspective, but still be a bad decision because the logical request no longer has enough budget to complete.
For example, imagine:
The gateway gives the request a 3s deadline. By the time it reaches the PSP, only 700ms remain. The payment operation is idempotent and the PSP returns a transient 503, so retrying is technically safe. But if the retry policy waits 1s before the next attempt, the correct decision should still be NO, because the retry cannot complete within the remaining deadline.
That also gives you a nice way to test retry amplification. If each layer independently gets 3 attempts, the logical request can turn into 27 physical attempts before reaching the dependency. Googleβs SRE guidance explicitly calls out this kind of multiplicative retry behavior and recommends propagating deadlines and controlling retries at the appropriate layer.
So for v2, Iβd love to see the model given something like
deadline_remaining,attempts_used,retry_budget_remaining, andexpected_backoff, then score not only whether it retries, but whether the retry is actually viable for the logical request.That would make the benchmark less about recognizing retryable errors and more about reasoning about the request as a whole.
Thanks for taking the time to explain it in such detail, this is exactly the kind of discussion that makes publishing an experiment worthwhile.
I'm definitely saving this for v2.
The Unsafe Payment Retry row is the most informative one, but check how its label is defined before reading it as a model error. A 503 usually comes from a load balancer or an overloaded tier that never ran the handler, so YES_AFTER_DELAY with a Retry-After header is a defensible answer; NO is the conservative policy answer because a 503 doesn't prove the handler didn't run. The Reason lines are where you can see which models knew that and which just followed the header.
I'd split the scoring in two. One column for "would this have caused a duplicate effect" (the NO class, where a miss is costly), and one for YES vs YES_AFTER_DELAY, which is only about timing. Right now a model that says YES on the idempotent DELETE and one that says YES_AFTER_DELAY on an unsafe POST lose the same single point, though only one of them could charge a customer twice.
On the leaderboard: with 14 scenarios, one item is 7 points. The 95% Wilson interval for 14/14 is roughly 78-100%, for 12/14 roughly 60-96%, and even Haiku's 9/14 is about 39-84%. So the ordering between models is not separated by this run; the per-scenario pattern (which items the models miss in common) is the real result. Running each scenario 5 times per model would also show whether the Retry-After miss is stable or a coin flip for the models that got it right.
Thanks for the detailed feedback. I agree that separating retry safety from retry timing would make the benchmark much more meaningful. It's something I started thinking about after seeing the YES vs YES_AFTER_DELAY results.
For the payment scenario, I still think NO is the right answer when there's no idempotency guarantee. A 503 doesn't prove the payment was processed, but it doesn't prove the opposite either, and that's exactly the risk I wanted to test.
I actually ran the benchmark three times and interestingly, the Unsafe Payment Retry results were completely consistent across all three runs. More runs would definitely be better, but this is still a small experiment anyway.
Definitely some good ideas here for a potential v2!
Agreed on NO for the payment case: with no idempotency key, a 503 is ambiguous about whether the charge went through, so a blind retry can double-charge.
One caution on the three identical runs: if the settings were the same, matching results show the model is stable on that prompt, not that the answer is right. That is still useful for the payment item, since a stable NO is what you want there. For v2, a variant of the same scenario with an
Idempotency-Keyheader present would show whether the model flips to YES only when it should.Not exactly the same scenario (the
Retry-Afterheader is missing), but I do have both Safe Payment Retry and Unsafe Payment Retry in the benchmark. Still, testing more variations could be interesting for v2.Wow, what a great article!!! Not only is it incredibly practical for engineers and architects, but it also highlights something bigger: nope, we still can't blindly trust coding agents, no matter what some people try to convince us of. π
Now I'm seriously tempted to send this to a few self-proclaimed "experts". π€£
Spot on alignment, Sylwia. Shifting from blind faith to rigorous architectural inspection is exactly how we prevent critical system failures. π οΈ AI is an awesome compiler assistant, but a terrible independent software architect.
Exactly. π What I find interesting here is that the agent can produce perfectly reasonable-looking retry logic and still make the wrong architectural decision.
The dangerous part isn't necessarily bad code. It's code that looks correct because it follows an obvious signal, while missing the state and constraints around the request.
That's where benchmarks like Daniel's are useful. They test the reasoning behind the code, not just whether the generated code compiles. And yes, I can already think of a few βAI replaces engineersβ arguments that would get interesting after running this benchmark. π
Oh, thank you! Glad you liked it, Sylwia.
Yep, totally agree. Kaggle doesn't allow all the latest models, but even the best ones can make mistakes, so a good code review is still necessary.
I think the benchmark has some practical value, but this first version still has its limitations. The YES vs YES_AFTER_DELAY decisions are debatable, and more runs would help too. So I'm sure those self-proclaimed experts would find something to argue about. π€£
Btw, how was Frontkon? Did you become even more famous? π
Hahahaha, I have no idea if I'm any more famous now! π€£ But apparently, the head of the conference has appointed me Frontkon's ambassador to Poland, and my mission for next year is to bring 60 Polish girls with me! π€£π€£π€£
60?! That's a huge task! π€£ That's practically a medium-sized company! π€£π€£
It seems that expensive AI models are not only expensive but also produce good results. Nice challenge! π
That's true! But for everyday coding tasks, efficiency might be more valuable than getting the absolute best results. π
daniel, this is a brilliant angle for an ai benchmark. testing models on contextual reasoning (idempotency, http methods, state) rather than just raw code generation is exactly what the industry needs right now.
your "unsafe payment retry" trap is the perfect example of why blind ai automation is dangerous. as someone building an ai coding ecosystem (koda), this is my biggest fear: an agent confidently executing a non-idempotent mutation just because it saw a
retry-afterheader, completely ignoring the lack of an idempotency key.i also completely agree with @arhancanli's point about splitting the scoring. a model guessing the wrong timing (
yesvsyes_after_delay) is a minor flaw, but a model causing a duplicate charge is a catastrophic failure. weighting safety heavily over timing would make this benchmark even more powerful for v2.thanks for proving that good old-fashioned architectural principles (like idempotency and deterministic fallbacks) still trump raw model size! π―
Thanks! Yep, as we discussed with @arhancanli, separating safety from timing would make the benchmark more meaningful. Definitely something worth exploring in v2.
glad we're on the same page! separating safety from timing is going to make v2 incredibly valuable for the community. looking forward to seeing the results. keep up the great work! π―
Benchmarking AI on system resilience logic like idempotency keys is an elite-tier experiment configuration. Most models can regurgitate basic code syntax, but they completely throw a kernel panic when calculating the edge-case risks of duplicate POST payloads on a 503 error. π
Thanks!
The retry decision gets scarier one layer down: the timeout-after-server-applied case, where the retry succeeds twice. We ended up treating idempotency as a precondition for retrying at all β no idempotency key, no retry, fail loud instead. Curious whether your benchmark cases covered the double-apply failure or stayed on the should-you-retry side.
Yes, I have something similar in my Unsafe Payment Retry scenario. The request returns 503 with a Retry-After header, but there's no idempotency key and we don't know whether the payment was already processed. So it covers the risk of double application, but not an explicit double-apply failure. The benchmark focuses on whether the model recognizes that risk and decides not to retry.
This is such a good question because retry logic looks so simple until you put it inside a real system π. I really like the point about retrying being more than just βtry againβ context, remaining time, cost, and whether the operation is even safe to repeat can completely change the answer. Definitely one of those topics that seems straightforward until production starts having opinions. Really enjoyed this one!
Yep, production always has the final say! π Thanks for reading!
Master Yoda says... Do, or do not. There is no try.
But never said, there retry is not.
Touche
Hmm, why are they still using DeepSeek-R1? It's a bit outdated. They should try the latest DeepSeek-V4.1 Flash.
and Good LuckοΌ
Unfortunately, Kaggle only offers DeepSeek-R1 and thanks! π
I didnβt realize there was so much to think about when deciding whether to retry a request πThe results were interesting too. All the best with the competition, Daniel π
Haha, there's always more to it than it seems! π Thanks for reading and for the good wishes! π