I need to admit something first. For about six weeks this summer, I ran an LLM-as-judge eval suite against an agent I maintain, watched the score go up, watched it go down, and made decisions based on that single number every single time. I never once opened the judge’s reasoning text. I looked at the score, I looked at the trend line, and I moved on. It felt responsible. I had an eval suite. Most people I know still don’t.
It wasn’t responsible. It was theater.
The agent in question takes a pull request diff and a bit of repo context and turns it into a structured, ready-to-file ticket: title, summary, affected components, suggested acceptance criteria. It’s the kind of task that’s easy to demo and genuinely hard to evaluate, because “is this a good ticket” is a judgment call, not a string match. So like a lot of people building agents right now, I reached for LLM-as-judge: have a second model read the input, read the agent’s output, and score how well it did.
I’d looked at DeepEval early on, since it’s become something of a default for this kind of work. It’s a solid library, but it’s Python, and my agent and its surrounding tooling live in .NET. Rather than bridge two runtimes for what is, underneath the abstractions, a chat completion call and a JSON parse, I hand-rolled a small evaluator in C#. If you’re on DeepEval, or Promptfoo, or any of the other eval frameworks, the lesson in this article applies just as much to you. The library isn’t the problem. The habit of trusting the number without reading the reasoning is the problem, and it’s a library-agnostic habit.
For six weeks the score sat around 0.78 on a 0 to 1 scale, drifted a little with each prompt change, and I treated that drift as ground truth. Then I had a slow Friday, pulled up the raw judge outputs instead of the aggregated score, and started actually reading what the model had written before it produced that number. It took maybe forty minutes to go through twenty cases. It was one of the more useful forty minutes I’ve spent on this project all year.
What I Was Actually Evaluating
Before I get into what I found, it’s worth being specific about the setup, because the failure modes I hit are tied to how the harness was built, and yours will have its own version of these even if the shape is different.
Each eval case is an input diff, a short description of the expected behavior (“the agent should identify the correct repo, produce a title under 80 characters, and list at least two acceptance criteria”), and the agent’s actual output. The judge model receives all three and returns a score from 0.0 to 1.0 plus a free-text explanation. I was running the judge on a self-hosted Qwen 2.5 model through Ollama, partly out of cost discipline and partly because I wanted the eval loop to run entirely on my own machine without depending on an external API staying up or staying priced the same. Everything in this article runs against that local setup; if you’d rather point the same code at a hosted model, it’s a one-line change to the judge implementation later on.
The scores got logged to a CSV, and a small dashboard script plotted the rolling average after every run. That dashboard was the only thing I looked at. The full judge response, reasoning included, was sitting in a JSON Lines file the whole time. I just never opened it.
The Day I Actually Read the Reasoning
Once I started reading instead of skimming the number, three distinct kinds of problem showed up, and none of them were things the aggregate score could have told me.
The first was the one that stung the most: cases where the judge’s own explanation directly contradicted its score. One case scored 0.90, which in my rubric means “matches expected behavior, no meaningful issues.” The reasoning read, almost verbatim, “the ticket is missing acceptance criteria and references the wrong repository, though the summary line is well formatted.” That is not a 0.90. That is closer to a 0.2. I found four more like it in that first batch of twenty, always in the same direction: strongly critical reasoning attached to a high score. I have a guess about why, which I’ll get to in the calibration section, but the honest answer is I don’t fully know, and that uncertainty is itself worth taking seriously.
The second problem was a stale rubric. One of my metrics checked whether the ticket followed “the standard template,” but the standard template referenced in the judge prompt was the template from an earlier version of the agent. I’d updated the agent’s output format three prompt revisions ago to drop a “Reproduction Steps” section that didn’t make sense for most of the tickets it generates, but I never updated the judge’s rubric to match. So for weeks, the judge was quietly penalizing the agent for correctly not including a section I’d deliberately removed. The score dip I’d attributed to “the new prompt made things worse” was actually the eval suite grading against a spec that no longer existed.
The third was a parameter I thought I controlled and didn’t. I run every judge call with temperature: 0.0, expecting deterministic scoring for the same input. When I reran ten of the flagged cases a second time to see if the mismatch was reproducible, four of them came back with different scores and noticeably different reasoning text than the first pass. The temperature setting was being sent correctly in the request; I checked the payload. But the specific quantized build of the model I had pulled didn't honor it consistently, something I only found because I went looking, not because anything failed loudly. A silently-ignored parameter is worse than a documented limitation, because there's no error message pointing you at it.
None of these three problems showed up as a change in the aggregate score. They showed up as noise I’d been averaging away.
Building a Diagnostic Harness
Once I knew what to look for, I didn’t want to keep finding it by manually scrolling through JSON files, so I built a small harness that runs alongside the existing eval suite and does three things: logs the full reasoning next to the score (which the original suite already did, it just wasn’t surfaced anywhere), flags cases where the reasoning’s sentiment and the score’s direction disagree, and produces a short report of the worst offenders for me to read by hand.
The core types are deliberately plain:
public record EvalCase(
string Id,
string Input,
string ExpectedBehavior,
string ActualOutput);
public record JudgeVerdict(double Score, string Reasoning);
public interface IJudge
{
Task<JudgeVerdict> EvaluateAsync(EvalCase evalCase, CancellationToken ct = default);
}
The judge implementation talks to a local Ollama instance over HTTP, using structured JSON output so the score and reasoning come back as a real object rather than something I have to regex out of free text:
public class OllamaJudge : IJudge
{
private readonly HttpClient _http;
private readonly string _model;
public OllamaJudge(HttpClient http, string model = "qwen2.5:14b-instruct")
{
_http = http;
_model = model;
}
public async Task<JudgeVerdict> EvaluateAsync(EvalCase evalCase, CancellationToken ct = default)
{
var prompt = $"""
You are grading whether an AI agent's output correctly follows the
expected behavior described below. Be specific and critical in your
reasoning; do not soften your explanation to match a score you have
not yet decided on.
Input given to the agent:
{evalCase.Input}
Expected behavior:
{evalCase.ExpectedBehavior}
Actual agent output:
{evalCase.ActualOutput}
Score the actual output from 0.0 (completely wrong) to 1.0 (fully
correct). Respond with JSON only, in exactly this shape:
{{"score": <number between 0 and 1>, "reasoning": "<at least two full sentences>"}}
""";
var payload = new
{
model = _model,
messages = new[] { new { role = "user", content = prompt } },
format = "json",
stream = false,
options = new { temperature = 0.0 }
};
var response = await _http.PostAsJsonAsync(
"http://localhost:11434/api/chat", payload, ct);
response.EnsureSuccessStatusCode();
var body = await response.Content
.ReadFromJsonAsync<OllamaChatResponse>(cancellationToken: ct);
var options = new JsonSerializerOptions { PropertyNameCaseInsensitive = true };
return JsonSerializer.Deserialize<JudgeVerdict>(body!.Message.Content, options)!;
}
private record OllamaChatResponse(OllamaMessage Message);
private record OllamaMessage(string Role, string Content);
}
If you’re on a hosted model instead, the only thing that changes is the constructor and the HTTP call target; the rest of the harness, and everything below, is identical either way. That’s the point of putting the judge behind IJudge in the first place.
The harness itself just runs the cases and keeps the full verdict, not just the score:
public record EvalRecord(EvalCase Case, JudgeVerdict Verdict, DateTime RanAt);
public class EvalHarness
{
private readonly IJudge _judge;
public EvalHarness(IJudge judge) => _judge = judge;
public async Task<List<EvalRecord>> RunAsync(IEnumerable<EvalCase> cases)
{
var results = new List<EvalRecord>();
foreach (var c in cases)
{
var verdict = await _judge.EvaluateAsync(c);
results.Add(new EvalRecord(c, verdict, DateTime.UtcNow));
}
return results;
}
}
Nothing here is clever. That’s deliberate. The whole point of this exercise was that I’d built something clever (an aggregate dashboard) on top of something I never actually inspected (the raw reasoning), and the fix was to stop skipping the boring, readable middle layer.
Catching Reasoning and Score Disagreement
I don’t think you can reliably automate “does this reasoning actually support this score” with another LLM call without just moving the trust problem one level up. What I wanted instead was a cheap, dumb, transparent heuristic that would flag candidates for a human to read, not a replacement for reading them.
The approach I landed on: keyword-based sentiment on the reasoning text, compared against the direction of the score. It is a heuristic in the least flattering sense of the word. It will misfire on sarcasm, on reasoning that lists both strengths and weaknesses in genuinely balanced proportion, and on domain-specific negative words that aren’t actually negative in context (“the ticket correctly flags this as a breaking change”). I want to be explicit about that, because the value of this tool is entirely in how cheaply it narrows down what a human reads next, not in any claim to accuracy on its own.
public static class ReasoningMismatchDetector
{
private static readonly string[] NegativeWords =
{
"fail", "failed", "failure", "incorrect", "wrong", "missing",
"does not", "doesn't", "did not", "didn't", "inconsistent",
"hallucin", "contradicts", "ignores", "omits", "unclear", "confus"
};
private static readonly string[] PositiveWords =
{
"correct", "accurate", "matches", "complete", "clear", "consistent",
"satisfies", "fulfills", "aligned", "appropriate", "well formatted",
"well structured"
};
// Returns a value from -1 (entirely negative-sounding) to 1 (entirely
// positive-sounding). A reasoning string with no keyword hits returns 0,
// meaning "no signal," not "neutral."
public static double SentimentScore(string reasoning)
{
var text = reasoning.ToLowerInvariant();
int neg = NegativeWords.Count(w => text.Contains(w));
int pos = PositiveWords.Count(w => text.Contains(w));
if (neg + pos == 0) return 0.0;
return (double)(pos - neg) / (pos + neg);
}
// Flags cases where the reasoning's sentiment and the score's direction
// disagree by more than the threshold. This is a keyword heuristic, not
// a sentiment classifier, and it exists purely to triage cases for a
// human to read, not to replace that reading.
public static bool IsMismatch(JudgeVerdict verdict, double threshold = 1.0)
{
var sentiment = SentimentScore(verdict.Reasoning);
if (sentiment == 0.0) return false;
var normalizedScore = (verdict.Score * 2) - 1; // map 0..1 onto -1..1
var disagreement = Math.Abs(sentiment - normalizedScore);
return disagreement >= threshold;
}
}
Running this over the same batch of cases that first tipped me off surfaced the exact ticket I’d found by hand, plus two more I’d missed:
CASE ID SCORE SENTIMENT DISAGREEMENT REASONING (TRUNCATED)
------------------------------------------------------------------------------
tkt-0042 0.90 -0.67 1.57 "missing acceptance criteria and
references the wrong repo, though
the summary line is well formatted"
tkt-0107 0.85 -0.50 1.35 "ignores the priority field
entirely, title is otherwise fine"
tkt-0031 0.20 0.60 0.80 "correctly captures scope and
components, wording is a bit terse"
tkt-0031 is a good example of the heuristic's limits: it's flagged, but reading it, the judge was reasonably scoring down for terseness on a case where terseness genuinely mattered. Not every flag is a bug. That's expected, and it's fine, because the tool's job is to shrink a haystack, not to find the needle itself.
The Report
The report generator just sorts by disagreement and prints the top handful, so a five-minute morning check replaces the forty-minute manual scroll I did the first time:
public static class MismatchReport
{
public static string Generate(IEnumerable<EvalRecord> records, int topN = 15)
{
var flagged = records
.Select(r => new
{
r.Case.Id,
r.Verdict.Score,
Sentiment = ReasoningMismatchDetector.SentimentScore(r.Verdict.Reasoning),
r.Verdict.Reasoning
})
.Where(r => r.Sentiment != 0.0)
.Select(r => new
{
r.Id,
r.Score,
r.Sentiment,
Disagreement = Math.Abs(r.Sentiment - ((r.Score * 2) - 1)),
r.Reasoning
})
.OrderByDescending(r => r.Disagreement)
.Take(topN);
var sb = new StringBuilder();
sb.AppendLine("CASE ID SCORE SENTIMENT DISAGREEMENT REASONING");
foreach (var r in flagged)
{
var truncated = r.Reasoning.Length > 60
? r.Reasoning[..60] + "..."
: r.Reasoning;
sb.AppendLine(
$"{r.Id,-10} {r.Score,-7:0.00} {r.Sentiment,-11:0.00} {r.Disagreement,-14:0.00} {truncated}");
}
return sb.ToString();
}
}
I run this after every eval batch now, not just when something looks off. It became a five-minute step, which is the only reason I actually keep doing it.
How Many Judge Outputs You Actually Need to Check
Finding bugs by reading twenty cases by hand is useful once, but it doesn’t tell you whether your judge is trustworthy in general, only that it was wrong on those twenty. For that you need a real calibration pass: a sample of judge outputs, scored independently by a human, compared against the judge’s own scores.
My rule of thumb, and I’ll say plainly this is a rule of thumb rather than a statistically airtight sample size, is at least 30 to 50 cases for an initial calibration, pulled with deliberate stratification across score bands rather than pure random sampling. Cases near your pass/fail boundary are the ones where judge and human are most likely to disagree, and pure random sampling under-represents them if most of your traffic clusters at the high end, which mine does.
To compare human scores against judge scores in a way that accounts for chance agreement, I use Cohen’s kappa. It’s a simple statistic, and it’s much more informative than raw percent agreement, because percent agreement doesn’t correct for two raters who’d agree by coincidence just from both leaning toward the same common answer.
public static class AgreementMetrics
{
// Both rating lists must use the same category indices, e.g. bucket
// continuous scores into fail=0, borderline=1, pass=2 before calling this.
public static double CohensKappa(
IReadOnlyList<int> raterA, IReadOnlyList<int> raterB, int categoryCount)
{
if (raterA.Count != raterB.Count)
throw new ArgumentException("Rating lists must be the same length.");
int n = raterA.Count;
var confusion = new int[categoryCount, categoryCount];
for (int i = 0; i < n; i++)
confusion[raterA[i], raterB[i]]++;
double po = 0;
for (int k = 0; k < categoryCount; k++)
po += confusion[k, k];
po /= n;
var rowTotals = new double[categoryCount];
var colTotals = new double[categoryCount];
for (int i = 0; i < categoryCount; i++)
for (int j = 0; j < categoryCount; j++)
{
rowTotals[i] += confusion[i, j];
colTotals[j] += confusion[i, j];
}
double pe = 0;
for (int k = 0; k < categoryCount; k++)
pe += (rowTotals[k] / n) * (colTotals[k] / n);
return (po - pe) / (1 - pe);
}
}
I bucket my continuous 0 to 1 scores into three categories (fail below 0.5, borderline 0.5 to 0.8, pass above 0.8) before running this, since Cohen’s kappa is built for categorical ratings, not continuous ones, and forcing a bucket also matches how I actually use the score downstream.
For interpreting the result, I use the Landis and Koch scale, which is a common reference point even though it’s a rough guide rather than a hard law of statistics:
KAPPA RANGE INTERPRETATION
--------------------------------------
< 0.00 worse than chance
0.00 - 0.20 slight agreement
0.21 - 0.40 fair agreement
0.41 - 0.60 moderate agreement
0.61 - 0.80 substantial agreement
0.81 - 1.00 near-perfect agreement
My personal threshold, for a judge I’m willing to let gate something automatically (blocking a merge, failing a build), is substantial agreement, kappa at or above 0.6. For a judge I’m only using descriptively, to rank cases for my own attention rather than to make a pass/fail call, I’ll tolerate moderate agreement, in the 0.4 to 0.6 range, as long as I keep reading the flagged cases myself rather than trusting the aggregate.
When I ran this calibration on my own setup, using 40 stratified cases, I landed at kappa of 0.51. Moderate. Good enough to keep using descriptively, not good enough to let gate anything on its own.
What I Do Differently When Agreement Comes Back Low
A 0.51 kappa isn’t a single problem with a single fix, so I don’t treat it as one. I work through three things in order.
First, I try a different judge model on the same calibration set, holding everything else constant, purely to find out whether the disagreement is a model capability issue or a prompt issue. Swapping my local Qwen model for a larger one, or briefly pointing the same IJudge interface at a hosted model, tells me quickly whether the ceiling is the model or the instructions I'm giving it. If a bigger or different model closes most of the gap, that's model capability. If it doesn't move much, the prompt is the problem.
Second, assuming it’s the prompt, I rewrite the judge’s rubric to be more concrete, usually by adding two or three worked examples directly in the prompt, one of a case that should score high and one that should score low, with the reasoning spelled out. Vague rubrics (“is this a good ticket”) invite vague, inconsistent reasoning. Specific rubrics (“does the ticket name a real repository in this list, does it include at least two acceptance criteria, is the title under 80 characters”) give the judge less room to drift.
Third, for any single sub-check that has a clear right answer, I stop asking the judge at all. Whether the ticket references a repository that actually exists in my org is not a judgment call, it’s a lookup, and I moved it into a deterministic code check that runs alongside the LLM judge rather than inside it. The judge is genuinely useful for the parts of “good ticket” that require judgment. It’s a liability for the parts that don’t.
My Rule Now
I calibrate the judge against human scoring roughly once a quarter, and immediately after any change to the judge prompt, the judge model, or the agent’s own prompt, since any of those three can shift the relationship between score and reality without warning. Between calibrations, I read the mismatch report after every batch, which takes minutes, not hours, now that the harness does the sorting for me.
The habit I’ve actually changed is smaller than all of this tooling, though. It’s that “the eval suite passed” no longer means “the agent is good” to me. It means the agent produced outputs that a specific model, under a specific prompt, with a specific and only partially-verified relationship to human judgment, scored favorably. That’s a real signal, and it’s worth having. It is not, on its own, the thing I used to treat it as.
Tags: llm-evaluation, dotnet, ai-agents, ollama, deepeval, software-testing, machine-learning
Top comments (0)