DEV Community

Cover image for 8 LLMs, 480 Questions, 1 Kaggle Benchmark: Who Can Explain a Traffic Drop?

8 LLMs, 480 Questions, 1 Kaggle Benchmark: Who Can Explain a Traffic Drop?

Nishikanta Ray on September 27, 2026

This is a submission for the Kaggle Benchmarking Challenge What I Benchmarked The short version I build InsightTrack, an ope...
Collapse
 
deanlee profile image
Dean Lee •

The split between tool routing and diagnosis mirrors how automated risk engines fail. Calling the correct pricing routine is cheap and mechanical. Interpreting whether the resulting variance is structural drift or high-frequency microstructure noise requires evaluating the underlying data-generating process.

When smaller models fail symmetrically, either inventing causal narratives for random walks or dismissing a forty-percent drop as variance, they behave like uncalibrated volatility filters. For an analytics product, a missed regime shift carries asymmetric downside compared to a noisy false positive. Saving ninety percent on inference unit economics becomes an expensive tradeoff if the user absorbs the tail risk of undetected decay.

Collapse
 
nishikantaray profile image
Nishikanta Ray •

Exactly — the tool-routing part is relatively easy; the hard part is interpreting what the data actually means. The asymmetric cost of missing a real regime shift is especially important for analytics products.

Collapse
 
deanlee profile image
Dean Lee •

That asymmetry is what makes the cost curve steep. When a tool call misfires, retry logic can catch the exception. When a diagnostics model mistakes structural churn for seasonality, the blast radius compounds silently until an operator intervenes.

Thread Thread
 
nishikantaray profile image
Nishikanta Ray •

Exactly. Operational errors are often recoverable; reasoning errors can quietly propagate because the system treats the wrong diagnosis as truth. That’s why knowing when not to infer anything is just as important as getting the diagnosis right.

Collapse
 
respect17 profile image
Kudzai Murimi •

Scoring whether a model knows when nothing happened is a great test design, most benchmarks only check if it gets the right answer, not if it correctly says there isn't one. The 35% reasoning-off number on cheaper models is the real headline.

Collapse
 
nishikantaray profile image
Nishikanta Ray •

Yeah, exactly. Testing whether a model can correctly recognize “nothing happened” feels like a much more realistic measure of reliability than benchmarks that always assume an answer exists. That 35% reasoning-off result is particularly interesting.

Collapse
 
micheypico profile image
Micheal Heypico •

From operating a multi-model routing layer: the underrated variable here is provider variance over time — model behavior shifts on the vendor side without any change in your code. Testing the same prompt across several models (we keep 32 behind one key at heypico.ai) is how you tell real signal from one model's quirk.

Collapse
 
nishikantaray profile image
Nishikanta Ray •

Exactly. Provider variance is an underrated variable too — model behavior can shift without any change in your code. Running the same benchmark across multiple models is a good way to separate real signal from one model’s quirks.

Collapse
 
mansio profile image
Mikhail •

Same two failure classes showed up in our memory-verification work on coding agents: fail-open models with false-accept 0.24–0.38, fail-closed ones silently dropping true claims. Different domain, same split — good to see it replicate outside code.

One question, since you still have the Grok ±reasoning pair harnessed: does the 100→35 gap hold on data-reading and SQL, or is it diagnosis-specific? We found CoT buys nothing on fact-verification arms, which suggests the effect is task-shape dependent rather than model-dependent — your rig could settle it from the other side for the cost of one more run.

Collapse
 
nishikantaray profile image
Nishikanta Ray •

That’s a great question. I haven’t tested the 100→35 gap on SQL/data-reading yet, so I’d be careful extrapolating from the diagnosis results. A cross-task run with the same harness would be a useful next experiment.