This is a submission for the Kaggle Benchmarking Challenge
What I Benchmarked
A while back I spent time manually poking at Gemini, Chat...
For further actions, you may consider blocking this person and/or reporting abuse
This one started as a very informal hunch and turned into an actual benchmark. The Grok result genuinely surprised me β I expected reasoning mode to make the model more likely to catch a planted mistake, not more likely to follow it.
The sample is small, so Iβm not claiming βreasoning mode is badβ from 15 questions. But the difference was big enough to make me want to investigate further.
Have you seen reasoning mode make a model more confident in a wrong path rather than more likely to catch it?
From operating a multi-model routing layer: the underrated variable here is provider variance over time β model behavior shifts on the vendor side without any change in your code. Testing the same prompt across several models (we keep 32 behind one key at heypico.ai) is how you tell real signal from one model's quirk.
The open question you end with actually came up for me last month. I had a bug in an agent pipeline and kept misreading the reasoning trace as the model working it out, when it was just building commitment to a bad intermediate state. Ended up taking twice as long to debug because the shown work looked considered. If the trace commits harder to wrong paths in reasoning mode, that is a real problem for using it as a first-pass debugging signal.
Ran the paper you cite before reading the rest. arXiv 2307.13702 is real, and you describe the intervention the way the authors do: inject a wrong step mid-trace, force the model to continue, check whether the answer follows the error. Their own headline points your way too ("as models become larger and more capable, they produce less faithful reasoning on most tasks"), so reasoning mode making a model more bound to a bad step is not a wild result.
The piece I'd sharpen, because it's the part a judge will press on: your two extreme rows are the softest evidence, not the strongest.
0/15 does not say Claude never follows an injected error. At 15 trials it is consistent with a true rate up to about 18% (95% one-sided). Same at the other end: 15/15 is consistent with a true rate as low as ~82%. "Never" and "always" are the two claims a 15-item sample cannot carry.
The row carrying your headline is Grok Reasoning 11/15 vs Non-Reasoning 2/15. That one holds (Fisher exact, p ~ 0.002, 5.5x). So the defensible claim is "toggling reasoning mode moved this model's faithfulness ~5x", with DeepSeek-R1's 15/15 read as "consistent with high faithfulness," not as the test confirming the test.
Not a knock on the writeup; it's more careful than most benchmark posts I read. If you ever have a number headed into a deck or a leaderboard and want it traced to source first, that is the thing I do.
Yeah, absolutely β and thatβs pretty much why I was careful throughout the writeup about not treating the 15-question sample as concrete evidence. I explicitly wanted the 0/15 and 15/15 results to be read as observations from this benchmark rather than βClaude always does Xβ or βDeepSeek always does Y.β
The Grok 11/15 vs 2/15 result is the part that genuinely surprised me and is why I framed the ~5x difference as the headline rather than making broader claims about reasoning mode.
Really appreciate you actually going back to the paper and checking the numbers, though. The distinction around the extreme rows is a useful way to make the limitations even clearer. π
Fair β you didn't overclaim, and that's exactly why the Grok row carries the piece.
One other thing: if you've got a number sitting in something you're about to act on β a writeup, a pitch, a decision β that you only ever saw secondhand, send it over. I trace it to the primary source and post the verdict. Free.
This was a really interesting read.
I honestly expected reasoning mode to help catch mistakes, not make them more βstickyβ π
The Claude result surprised me too. Iβd be curious to see this tested on more models.
Nice work β would be cool to chat more about AI benchmarking sometime.
Well, its first time I created a kaggle benchmark. And I wanted to include more models but thought, that would make my write up more complex, so I sticked to known models, also its free tier so wanted to finish the task in free credits.π
But thanks for the read, have a great day!
The 11/15 vs 2/15 result is wild, but the control suggested in the comments makes it even more interesting.
If Grok follows a correct injected trace just as often, then this may be less about reasoning mode making errors stick and more about how strongly it anchors to supplied reasoning.
That feels like the experiment worth running next.
The 5x jump with reasoning mode on is a genuinely uncomfortable result, since it suggests the visible chain is rationalizing the answer rather than producing it. I have started scoring faithfulness separately from accuracy for exactly this reason, because a model can land the right answer while narrating a path it never took. Did the models that followed your injected mistake also sound more confident in the final answer?
Really interesting experiment. The 5x difference from toggling reasoning mode is especially surprising, and I like that youβre careful not to overstate what a 15-question sample can prove. The limitation around interpreting visible reasoning as actual verification is a great point for further testing.
The strongest addition to this design is a control arm you don't have yet: rerun the same questions with an intact trace β feed the model its own correct step instead of the corrupted one. If reasoning-mode Grok tracks valid shown work just as obediently, the 11/15 is anchoring on provided reasoning, not error-blindness. Those have different fixes: error-blindness is a calibration problem, anchoring is context-competition β and you can tell them apart with exactly the pipeline you already built.
Same for position: split the corruption into early vs late injections. Late errors usually bind harder (less remaining budget to reroute), and if your 11 flips concentrate on late injections, the 'reasoning makes mistakes more binding' story gets a mechanism.
On Claude 0/15: label the transcripts before interpreting. Rerouting around the injected step (noticing the error) and silently recomputing from the problem statement (never using the trace) look identical at the answer level but are different mechanisms β only the first one is verification.
On your open question: a visible thinking trace is evidence of what the model binds to, not a check on correctness. Debug with it, verify against ground truth.
The 'Adding Mistakes' intervention design is the cleanest thing here. Most faithfulness tests I've seen basically ask 'does the model think it's right', not 'will it follow a wrong step if you shove it partway through the chain'. Those are totally different questions.
The Grok result is the real headline though: same weights, reasoning mode toggled, 5x more likely to follow the injected error. That's almost an indictment of reasoning mode as a reliability feature β the extended chain gives the corruption more surface to propagate through.
Curious whether the faithfulness score correlates with task complexity or degrades uniformly regardless of problem type?
One thing I'd push back on is the construct itself: the injected step was never the model's own reasoning, so a model that ignores it isn't necessarily being unfaithful β it might just be robust to context tampering. Read that way, this benchmark is closer to measuring 'trace susceptibility' than faithfulness, and the 5x Grok swing becomes a story about reasoning-mode models binding harder to displayed context (RL rewards coherent rollouts, so the model commits to whatever's in the trace). The sharp implication is for agents: a model's own scratchpad and tool summaries get re-injected as context every step, so a reasoning model that binds to corrupted intermediate state is exactly the failure mode scratchpad-based prompt injection exploits. Did you happen to check whether the follows clustered on questions where the model had already committed to a partial answer before the injection point?
Thatβs actually something I tested in the benchmark β the full notebook and benchmark are included with the post, so you can see exactly where the injection happens and how the model responds before/after it.
The reason I framed this around faithfulness is based on that observed behavior, rather than assuming that ignoring the injected step automatically means βunfaithful.β If you spot something in the benchmark that changes that interpretation, Iβd genuinely be interested in seeing it.
Dhruv, this was really easy to follow π The way you explained the benchmark and how you tested the models made it much easier to understand. Very interesting results too. Good luck with the challenge!
It's my first time using kaggle benchmark and I actually wanted a platform to test models simultaneously, cause it also helps in college and stuff. So, Found it.
Thanks for the read! Have A great night!π
Good luck for this challange
Thanks di!π
Cool post bro.
Bro, its a submission post.π
The Grok on/off pair is one of the cleanest single-variable comparisons in this challenge. I had the same pair in a chart-reading benchmark and saw the opposite trade: reasoning mode helped Grok 4.20 read monitoring panels (0.646 to 0.694) at about 11 times the cost.
Put next to your result, it looks like reasoning makes Grok commit harder to whatever its reasoning says. That helps when the reasoning starts from correct observations and hurts when it is handed a planted error.
Did the reasoning-mode traces ever notice the 57 vs 47 discrepancy before following it? That would separate "didn't check" from "checked and deferred to the shown work".
15 questions is a thin sample to draw behavioral conclusions from, and the arithmetic corruption example is the easiest case β the model can see 25 + 22 β 57 with zero domain knowledge. The real test is whether the injected error survives in the logic/deduction bucket where the wrong step looks plausible all the way through. I ran something similar with 80 prompts and the gap between arithmetic and logic was where the interesting failures lived β models that self-corrected on math still followed bad premises on multi-step deduction because the error wasn't locally checkable. Would love to...
Sharp benchmark, and the open question at the end is the one that actually matters: a visible chain-of-thought is a self-report, not independent evidence. The "Adding Mistakes" intervention is testing something specific β whether the trace is causally load-bearing or just narration bolted on after the answer's already decided. A model that's 15/15 "faithful" to an injected error isn't being unreasonable, it's just telling you the trace is the computation, for better or worse.
The part I'd want isolated next: whether "reasoning mode" is teaching the model to trust its own prior steps more, or just making the context longer and giving a wrong early token more downstream tokens to anchor to. Those predict different fixes β one says don't trust the trace, the other says the trace is fine but the model needs a way to re-derive from the premise instead of extending from the last line it wrote. The Grok reasoning/non-reasoning pair is the right control for that; worth also running the inverse β inject a correct but unusual mid-step and see if reasoning mode anchors that hard to it too, or only to the wrong ones.
Either way: if a visible thinking trace can be this confidently wrong about its own reliability, treating it as ground truth in an eval harness or a debugging session is the mistake, not the trace itself. It's evidence to cross-check, not a receipt.
Ok ****