DEV Community

Cole Halton
Cole Halton

Posted on

Jeff put a 2B classifier next to Jev: 96% on labels, 62.6% on JudgeBench

The Jeff repo is a family of tiny "decision models" that do zero-shot classification and return a calibrated probability per option from a single forward pass. No text generation, no parsing. The README ships a head-to-head table against Jev across five benchmarks plus JevBench's hard tier, and I keep going back to it because of the shape of the numbers, not the headline.

What the models actually are

Jeff-Qwen3.5-0.8B, Jeff-Qwen3.5-2B, and Jeff-Gemma4-E2B, all Apache 2.0. Trained entirely on local hardware: one RTX PRO 6000, roughly two hours for the 0.8B and 3.5 for the 2B. About 22 ms per decision on that GPU, 28 ms on an M4 Max via MLX. The request format matches Jev (the v1/systemone endpoint, with choice, noul, and score question types), so you can slot it in wherever you already call Jev. The training code starts from the AutoJev recipe.

The overall number is the boring part. Jeff-Qwen3.5-2B lands at 83.1, Jev at 83.0. Read only that and you'd conclude a 2B model matches a much larger one.

The per-slice table says something else.

The per-slice table is the real eval

Benchmark Jeff-Qwen3.5-2B Jev (published) AutoJev-27B
Overall (5 benchmarks) 83.1 83.0 84.9
Financial PhraseBank 96.3 77.0 84.2
RAGTruth 88.9 77.3 88.9
WinoGrande 79.0 90.7 83.3
BBH 68.0 94.3 82.8
JudgeBench 64.6 78.6 78.9
JevBench hard (separate) 53.3 73.3 70.3

On the classification and grounding slices Jeff wins by a wide margin: 96.3 vs 77.0 on Financial PhraseBank, 88.9 vs 77.3 on RAGTruth. On the reasoning-heavy slices it falls off a cliff: 64.6 vs 78.6 on JudgeBench, 53.3 vs 73.3 on JevBench hard, 68.0 vs 94.3 on BBH.

The author is upfront about it in the README: these are very small models, and at this size "their reasoning won't match Jev's." That's the honest framing. But watch how the aggregate hides the tail. One overall score averaged across five benchmarks makes a specialist look like a drop-in replacement, when what it actually does is give up the reasoning end to win the label end. The five-benchmark average is also just a weight choice: swap in more reasoning slices and the 83.1 collapses toward the JudgeBench number.

The 0.8B tells the same story. It scores 79.1 overall, but 96.4 on Financial PhraseBank, 86.1 on RAGTruth, then 64.0 on BBH and 62.6 on JudgeBench. Even the smallest model nails the classification slices and drops 30-plus points on reasoning. That gap doesn't shrink with the model size in this family. Size buys you a couple points overall and almost nothing on the reasoning tail.

There's a lighter version of the same pattern in the repo's games test. The 0.8B played Frogger, Doom, and Pac-Man zero-shot, choosing moves by plain-language consequences. It edged the hand-coded rule bot on Frogger (10.3 crossings vs 10.25) and tied on Doom (6.55 kills), then lost badly on Pac-Man (57.0 pellets vs 94.1). Cheap model clears the easy game and gives up the harder one. Same shape as the benchmark table: the win is real, the loss is real, and a single blended score would have shown neither.

The claim that needs a second look

Two things in the README deserve a skeptical read before you take the table at face value.

First, the comparison line itself: "Jeff's published figures were measured on a different sample of the same benchmarks." That's not a controlled head-to-head. Jev's numbers and Jeff's numbers come from different runs on different samples, so the deltas are directional, not exact. I want the side-by-side to be the same items scored by both models in one harness. Until it is, the per-slice shapes are believable and the decimal points are not.

Second, the fine-tune claim. The README says a short fine-tune "moved held-out accuracy from 31.7% to 95.8% in under half an hour on one GPU." That's a big jump and it's the strongest number in the repo. It's also a fine-tune on your own examples, which is the polite way of saying the zero-shot number on your task might be 31.7%. The 95.8% is the number after you've handed it labeled data. If you're shopping for a decision model to drop into a review gate, that's the number that decides whether zero-shot is good enough for your labels or whether you're signing up to label a few hundred examples first.

Why this matters if your reviewer gates on a judge

Most AI code review tools make a per-finding call that's shaped like classification: is this finding real, does this diff match the rule we wrote, is the change in scope. A small calibrated decider is a good fit for exactly those. The probabilities are fixed, so the same diff gets the same verdict twice, which is the entire reason deterministic judges beat an LLM judge that flips run to run.

The reasoning-shaped decisions (is this actually a bug, what's the right fix, is this whole approach wrong) are where the 2B drops 14 points on JudgeBench. Those still want the big model.

So the practical move is to split the gate. Route the classification-shaped decisions (label the finding, does it match the rule, is it in scope) to a typed decider you can pin and re-run. Keep the reasoning-shaped ones on the big judge. Then report a number per slice instead of one aggregate accuracy that quietly averages away the part you're bad at.

That framing also explains a failure I keep seeing in tool marketing: a reviewer reports 96% and nobody measured recall. Accuracy on a slice where the model is good says nothing about the slice where it's bad, and if your rule gate lives on the bad slice, the number you were shown is decoration. I made that argument in an AI reviewer reported 96% accuracy and nobody measured recall, and the Jeff table is a clean illustration of how a good aggregate hides a weak slice.

How to copy the method without copying the hype

  • Fixed prompt set and fixed seed, both deciders on the same inputs. The Jeff/Jev comparison doesn't do this yet, and it's the one upgrade that would turn the table from directional into truth.
  • Report per-slice, not one average. The Jeff table works because every row is a slice. An overall row on top of it is fine as long as it isn't the only row.
  • Check calibration, so probability and not just top-1 accuracy. A confident wrong 0.95 is worse than a hedged 0.6 when you're gating a PR. This is also why a typed decider with a real probability beats an LLM asked to "return a confidence."
  • Commit the results as JSON so the claim is re-runnable, not a screenshot. Reproducibility is the whole point of these models running locally.

And measure the cost you're optimizing for. 22 ms per decision on a workstation GPU, no text to parse, means you can afford to run the decider on every file in a PR rather than sampling. That changes the benchmark: if the cheap decider is fast enough to score everything, you're not comparing accuracy at a fixed budget anymore, you're comparing coverage. A judge you can only afford to run on 1% of diffs is measuring the wrong thing entirely.

Reviewers that run on your PRs, Kodus, CodeRabbit, Qodo, Greptile, all sit on some judge somewhere. The question I want answered for any of them is identical to the one this table answers for Jeff: which decisions in the pipeline are classification-shaped, and do you have a per-slice number for them, or just one accuracy figure. If a vendor can't name the slice, they don't have the number.

What I'd actually take from this repo

The interesting thing about Jeff isn't that a 2B model matches Jev overall. It's that the overall number is the least informative line in the table. A reader who stops at 83.1 vs 83.0 walks away thinking "small model, same quality, way cheaper." A reader who reads the slices walks away with a routing decision: use the small typed decider for labels, keep the large model for reasoning, and never trust a blended score again.

The receipt here is the README table. It's small, it's reproducible, and it lists exactly which slice you're buying and which one you're giving up. I'd rather have that than a vendor scorecard with one bold number and no slices under it.

Top comments (0)