DEV Community

Cover image for Green Scores Without a Pinned Harness Are Screenshots
Igor Eduardo
Igor Eduardo

Posted on

Green Scores Without a Pinned Harness Are Screenshots

Everyone is talking about Evaluation Cards and reproducible eval benches — the EvalEval × UK AISI write-up on Hugging Face, plus a week of discourse where green numbers still hide the wrong harness (and DeepEval-style “which half is broken?” chatter keeps landing). That heat is a live market object, not a niche footnote.

Here is what I would require before I trusted a green eval score for high-stakes retrieval or generation — and what I would not claim yet.

The stake

Teams will ship the wrong default.

A dashboard can show a comforting pass rate while the report never pins how the score was earned: which protocol, which compute budget, which dataset slice, and whether the number measured retrieval, generation, or a blended soup. Celebrate that number and you optimize a screenshot. The next engineer cannot reproduce the gate, and the next incident cannot tell which half failed.

If your domain is identifier-heavy, multi-hop, or regulated-adjacent, a fluent answer after an unpinned harness is not a win. It is a silent miss with a pretty badge.

Opinion (one sentence)

A green score without a pinned harness — protocol, compute, dataset slice, and which half of the pipeline it measures — is not an eval contract; it is a screenshot.

That is a preference with a limiter, not a SOTA claim. I have not run EvalEval in production. AISI card fields and anyone else’s leaderboard numbers stay theirs. What I am formalizing is the harness pin as a first-class part of eval-as-contract, from the same reliable-systems lane I already publish on.

What a pinned harness looks like (menu, not recipe)

Evaluation Cards earned attention because they make the boring fields visible: protocol, compute, data cut, scoring rules. You do not need their exact schema to steal the shape.

Field Question it answers Failure if skipped
Protocol What procedure produced this number? “We eval’d it” with no replay path
Compute What budget / hardware / call limits? Incomparable scores across teams
Dataset slice Which queries / docs / holdouts? Student-subset vibes framed as prod
Pipeline half Retrieval, generation, or joint? Green faithfulness while recall is broken (or the reverse)

The DeepEval-style lesson that keeps circulating is the same shape: a single RAG score can hide which half failed. Pin the half. Prefer two honest columns over one blended crown.

The check you can run without my internals

You do not need my private stacks (and I will not publish them). Steal this menu-only check:

  1. Pin the harness block before the score. Same report header every time: protocol name/version, compute class, dataset slice id, and pipeline half (retrieval / generation / joint). If any field is missing, the number is not ready to celebrate.
  2. Split halves when you report. For the same query set, publish retrieval quality and generation/faithfulness as separate columns. One green blended score is not enough — “which half is broken?” should be answerable from the table alone.
  3. Freeze the evaluator before you chase labels. If the judge, metric, or prompt changed mid-board, say so. Moving the goalposts mid-week and claiming progress is theater.
  4. Keep a missing-evidence arm. Same task with the gold document removed (or retrieval forced empty). If generation still narrates a confident answer, the eval contract is incomplete — regardless of how green the happy-path column looks.

No thresholds copied from anyone’s blog as “our production numbers.” No paste recipe. Calibrate cutoffs on your corpus; where missing evidence is costly, treat empty retrieval as a failure in the harness, not as a soft success.

Why query-class honesty transfers

On Portuguese clinical text, BM25 and dense retrieval solved different query classes; fusing them beat either alone on a public 500-query study. Exact terms and conceptual phrasing fail in different ways. That finding is checkable: open code at nomad-link-id/hybrid-rag-pipeline, companion write-up on Dev.to.

The lesson that transfers to Evaluation Cards is not “clone our fusion.” It is eval honesty on query classes and pipeline halves. If your harness cannot say which class and which half produced the green cell, you are not ready to ship the number into a decision.

What I will not claim

  • That EvalEval, UK AISI cards, or any one vendor harness is universally correct for every workload. Workload-honest eval first.
  • That we “ran EvalEval inside Cortexa/DocMinds.” We did not. Trend-jacking with fake usage is spam.
  • That a pinned harness alone proves end-to-end answer quality. Downstream faithfulness and missing-doc policy remain separate gates.
  • Internal thresholds, prompts, or clone guides for private products. Menu only: pin the fields, split the halves — not the spice blend.

Why this formalizes who we are

The market is amplifying Evaluation Cards and green eval dashboards. The thesis I want indexed next to that heat:

  • Eval as contract — protocol, compute, slice, and pipeline half are gates, not vibes.
  • Reliable systems — production AI fails on harness mismatch and wrong-half celebration — not only on model choice.
  • Retrieve-first — when the retrieval half is unpinned, generation fluency is not evidence.

If a sharp reader walks away believing this engineer will not celebrate a green score without a pinned harness, the post did its job. Follows that come from that are the point — not volume.

Soft pointers (contribution first)

Stop here. No recipe. No “we proved SOTA.” Preference + limiter + public trail.


By Igor Eduardo · Austin, TX · Engineer of reliable search and AI systems for high-stakes science · https://igoreduardo.com

Top comments (1)

Collapse
 
kaziava profile image
Hardcore Engineer •

Igor, "a green score without a pinned harness is a screenshot" is the
precondition assertion wearing an eval-contract coat, and reading it through
my own lens makes your whole thesis click into place. In my golden set a
negative test that passes because the retriever never surfaced the trap is a
green result whose precondition was never asserted — the test ran, the
verdict came back, and the path that was supposed to be exercised was never
touched. Your screenshot framing is the same shape one layer up: a green
score arrives without the harness pin that asserts what had to be true for it
to mean anything, and the number that was supposed to be a contract is just a
picture. That is not an eval result; it is a green lie that looks like a
dashboard.

What strikes me is how cleanly your four pinned fields map onto my own
mechanics. Protocol answers "what procedure produced this number" — the same
way my embedder_tag answers "which model earned this retrieval verdict." A
score without a protocol pin is a pass with no route tag: you can see the
green cell, you cannot see which branch of procedure earned it, and the
branch may have changed mid-board while the headline still read "improved."
Your compute field is the budget precondition my cost per task asserts —
tokens burned, not price promised. Your dataset slice is the trap chunk id
made visible: which queries and documents were in the harness, so the next
engineer can reproduce the gate instead of guessing. And your pipeline half
is exactly my two gates — retrieval precondition versus generation contract —
stated as a column instead of a gate, so "which half is broken" is answerable
from the table alone instead of buried in a blended score.

The bridge I could not resist building from my side is this: your
missing-evidence arm is my "test did not run" verdict in eval-contract form.
A green faithfulness score next to a broken retrieval is a lie that looks
like success — the generation narrates confidently, the answer is fluent, and
the path that was supposed to bring the evidence was never exercised. Your
control — same task with the gold document removed, assert the generation
refuses or flags absence — is exactly the negative test I run in my golden
set: a question whose correct answer is "not in documents," so any confident
answer fails without anyone judging content. Your missing-evidence arm is the
protocol-layer version of my third verdict, and it is doing the same work:
catching the green lie that looks like a functioning pipeline.

And here is the part I underlined twice: your "query-class honesty transfers"
section is the same insight as my anchor mortality, just pointed at a
different layer. Your Portuguese clinical study showed that BM25 and dense
solved different query classes, so fusion beat either alone — and the lesson
is not "clone our fusion" but "pin which class produced the green cell." My
anchors are mortal on purpose: a verbatim substring that resolves to an id is
a stamp, a paraphrase that merely sounds like one is a stamp that lost its
id, and it fails quietly in exactly the way a blended score fails quietly.
Your query-class pin is the same discipline — the context that makes the
assertion meaningful must travel with the result, instead of being assumed.

So the question I genuinely want your read on, continuing your own framing
rather than proposing a new one: your four pinned fields are the precondition
of a trustworthy score, and your check is "if any field is missing, the
number is not ready to celebrate." Is there an equivalent for the eval layer
itself — not "green score" or "pass rate" but something like the share of
eval runs that actually pinned all four fields — protocol, compute, dataset
slice, pipeline half — before counting the result as a contract? Because if a
screenshot without a pinned harness is what just blinked, then the metric
that catches the blink first is the one that asserts the harness precondition
before counting the score. Without that stamp, "our eval passed" is a green
lie wearing a dashboard sticker.

I suspect you are closer to that metric than the post lets on, the same way
you were closer to the two-gates idea than the thread reply admitted. The
pattern with you is that you build the mechanic first and name it two posts
later, so I am watching for the naming.