DEV Community

Mayank Laddha
Mayank Laddha

Posted on

Judging an LLM judge?

I have been reading some blog posts about LLM as a judge and was building a small evaluator to evaluate the judge itself .

My method is simple:

The dataset is:

task

rubric

ideal response

negative response
Enter fullscreen mode Exit fullscreen mode

The idea is then to test different models as judges for things like:

repeated-run consistency

position bias

sensitivity to verbosity

accuracy / ability to prefer the better response
Enter fullscreen mode Exit fullscreen mode

Here, “negative response” doesn’t necessarily mean a wrong answer. It can just be a response that is less preferred according to the rubric.

I have an initial version with around 200 lines of code
https://github.com/maylad31/judgeDjudge

But I’m more interested in discussing the idea.

If you have used LLM judges in practice, are there other failure modes or better ways of testing them?

Happy to hear criticism or suggestions or positive things about my method/code.

Top comments (2)

Collapse
 
raknaos profile image
Raknaos •

Testing the judge with ideal plus negative pairs is the right skeleton. The failure mode I'd add to your list: a judge can be perfectly consistent and still wrong — agreement across repeated runs measures stability, not accuracy. Keeping a small anchor set whose labels you actually trust, and scoring the judge against those rather than against itself, is what separates the two.

Two traps at this dataset size: verbosity bias tends to vanish once responses are length-matched, and position bias is uneven across model pairs, so it's worth reporting per judge instead of in aggregate. How many ideal/negative pairs do you run per model? Under roughly fifty, the consistency numbers are mostly noise.

Collapse
 
mayank_laddha_ml profile image
Mayank Laddha •

yeah, i am still researching more on it. thanks!