DEV Community

Cover image for My Harness Used One Label for Three Different Failures.

My Harness Used One Label for Three Different Failures.

Self-Correcting Systems on September 14, 2026

Three fixtures, three separate calls into the same reducer. Here is the complete failure_reasons each one returned, unedited: unreadable arrivin...
Collapse
 
vinhnguyenthanhdn profile image
Vinh Nguyen •

The split moves the boundary, but EXEC_COMPARATOR_ERROR inherits the exact property you just diagnosed: it still carries no subject. Both canonicalJsonBytes calls sit inside the one try, so a throw from canonicalJsonBytes(prepared.expectedExecArguments) — your own frozen fixture, nothing the model sent — lands under the same name as a throw on the arriving object.

The old name blamed the arguments for a comparator failure. The new one blames the comparator for a fixture failure, and a reader fills that in as our tooling broke the same way they used to fill in the model deviated. It closes in the same shape as the parse split: canonicalize the expected side once before the try that wraps the actual side, which you want anyway since it is prepared data being recomputed on every call.

Collapse
 
kenielzep97 profile image
Self-Correcting Systems •

you're right, and the part that stings is that the post already argued against you and the argument is dead.

what i wrote: "an un-canonicalizable expectation and a genuine comparator bug are both on my side and are not distinguishable from the outside. i declined to split them, because inventing a distinction the code cannot detect is the defect i was fixing."

your hoist is the detection. canonicalize the expected side separately and the code knows exactly which operand failed. so the reason i gave for not splitting was wrong, not just incomplete.

confirmed in the code. both calls share the try:

} else if (!canonicalJsonBytes(actual).equals(canonicalJsonBytes(prepared.expectedExecArguments))) {
Enter fullscreen mode Exit fullscreen mode

and the test i shipped to prove the new name puts the bad value in the expected object, so i wrote an assertion certifying a fixture defect as a machinery failure. it is also in reducer.mjs at 131-137, same shape, which the post does not mention at all.

one thing your fix does not do, and i would rather say it than let it read as solved. that reducer runs after the model. hoisting the expected side separates the attribution but the throw still happens post-hoc, so a bad fixture has already cost a relay call. the version that actually belongs is validating the expected bytes at prepare time, before the model is consulted, which is the same preflight obligation from the other thread.

both unbuilt. the split as published moved the boundary instead of removing it.

Collapse
 
orca_forge profile image
orca forge •

I run a fallback pool of six LLM providers behind LiteLLM. The provider errors arrive already well-named: 401 bad key, 402 billing, 404 model gone, 429 rate limited. The collapse was in the consumer — my retry policy had one bucket, "provider failed", with num_retries: 5 on top of it. Retry only helps the "wait and it clears" kind; 402 and 404 failed immediately, every time.

The difference from yours is who reads the label. EXEC_ARGUMENTS_MISMATCH misleads a human, and a human can at least feel that a name is off. Mine went to a retry policy, which can't feel anything. One dead provider in six meant 1/6 of requests failing indefinitely, and the logs never said "dead". They said "failed, retried, failed".

Collapse
 
kenielzep97 profile image
Self-Correcting Systems •

Yours is honestly worse than mine. When a human reads a bad label they can at least get a weird feeling about it. Your retry policy just did whatever the bucket told it to. One dead provider out of six failing forever while the logs keep saying failed, retried, failed is the same words meaning two completely different things. Pulling 402 and 404 out as never retry and leaving 429 as wait it out is the call the label should have made for you in the first place.

Collapse
 
xuks124 profile image
xuks124 •

Reading this reminded me of a failure class I have hit twice, and it is worth naming because it hides behind clean logs: the component that is running but no longer connected to reality.

In my case it was a risk guard. The rule ran on schedule, the comparison was correct, the test suite was green - and the number it compared against came from a snapshot that a different code path had quietly stopped refreshing. Nothing threw. The guard was decorative for weeks, and the only reason we caught it is that someone asked "what is that number right now?" and the honest answer was 0.

Two cheap habits that catch this class:

  1. Assert on the effective value, from the same object production reads. Not a recomputed copy. If production compares state.margin_used, the test asserts state.margin_used - the divergence then shows up as a red test instead of a silent one.
  2. Log the value that was actually used, every cycle, not the constant it was supposed to come from. margin_used=0.00 in a log line beats a month of "the code looks correct".

The third cousin of this bug is state that dies with the process: everything looks right until a restart, at which point the counter resets and the limit politely re-opens. The only test I trust for that one is destructive - let the limit trip, kill the process with something open, restart it, and watch what it does next.

I wrote up the restart-persistence variant (with the fixes) here, if it is useful: xuks124.github.io/vigildesk/blog/r... - the code is MQL5, the shape of the bug is not.

Thanks for the write-up; "My Harness Used One Label for Three Different Failures." is a better title than most posts in this space earn.

Collapse
 
mickyarun profile image
arun rajkumar •

A failure in the checking stage reads as a deviation in the thing being checked. That's the sentence, and it generalises well past your reducer.

The payments version: a status callback arrives, parses cleanly, and carries nothing that separates success from silence. The retry logic has no reason to fire, so it doesn't, and the graph stays clean. Same shape as your middle line — one name covering rejected, compared-and-differed, and never-compared — and the damage is the same, because triage starts from the wrong subject.

@vinhnguyenthanhdn's point about EXEC_COMPARATOR_ERROR inheriting the same subjectlessness is the one I'd fix first. A label earns its keep when it names who failed, not which stage was running when it happened.

Collapse
 
kenielzep97 profile image
Self-Correcting Systems •

you picked the one to fix first and it is built. Vinh was right that EXEC_COMPARATOR_ERROR inherited
the same subjectlessness, because both canonicalize calls sat in one try, so a throw on my own
frozen expectation came out as my machinery failing.

four names now, on my tree and not on origin yet:

EXEC_EXPECTATION_INVALID  my expectation cannot be canonicalized
EXEC_ARGUMENTS_INVALID    what arrived could not be parsed
EXEC_ARGUMENTS_MISMATCH   both sides usable, they differ
EXEC_COMPARATOR_ERROR     the comparison did not complete
Enter fullscreen mode Exit fullscreen mode

and one thing past his fix: hoisting only above the inner try leaves the expectation check behind
the arriving-args guard, so my defect hides behind a bad payload. both report independently now.

what it still does not do: that reducer runs after the model. a bad expectation of mine has already
cost the call by the time it is named. catching it at prepare time is the version that belongs and
it is not written.

your payments version is the sharper one because there is no exception at all. a callback that
parses cleanly and carries nothing separating success from silence means the retry has no reason to
fire, so the graph stays clean and nothing is ever red. mine at least threw. yours is the version
where absence reads as a pass and no stack trace exists to argue with.

"a label earns its keep when it names who failed, not which stage was running" is the harder
standard and COMPARATOR_ERROR still only clears the second half. it names the stage. it does not
say whether the stage or its input was at fault.

Collapse
 
nark3d profile image
Adam Lewis •

Three reasons sharing one label would have cost me a morning, because a fixture failing on the wrong invariant sends you to the wrong call. Naming each failure at the point it is raised is also what makes the harness worth keeping after the code moves on. The step it leaves out is that the label only stays honest while the invariant behind it still matches the reducer, so the name is a claim you have to re-check whenever the reducer changes. Keep a label nobody re-reads and you may as well have one label for three failures.

Collapse
 
mudassirworks profile image
Mudassir Khan •

the one catch for three failure stages pattern is one we’ve shipped and paid for. the version that bit us hardest was in an agent eval where we wrapped both model output parsing and expected output comparison in the same catch. both landed as COMPARISON_FAILED, and for two sprints we blamed the model when the real bug was in our comparison serializer. splitting each stage into its own named catch was the fix, but the harder lesson: any eval code with a catchall is hiding which party is guilty. what pm25coder named is real — a subjectless label invites you to assign blame to the wrong actor. are you exposing separate failure codes by stage now, or keeping one label with a richer payload?

Collapse
 
kenielzep97 profile image
Self-Correcting Systems •

Separate codes by stage. I went the way you did.

The commit is dd1a654 if you want to look. The catch that used to swallow everything
now has its own outcome:

} catch {
failures.add('EXEC_COMPARATOR_ERROR');

Unreadable arriving args are EXEC_ARGUMENTS_INVALID. Usable args that differ stay
EXEC_ARGUMENTS_MISMATCH. Three names where that catch used to write one.

One caveat before you go looking, because you'd find it anyway: that split is on a
branch, not on main. origin/main line 162 is still the old catch writing
EXEC_ARGUMENTS_MISMATCH, and neither new name appears there at all. So I've written the
fix and not shipped it, which is a different sentence than the one I started with.

I considered one label with a richer payload and rejected it for a boring reason. A
payload field gets read when someone is already suspicious. A state name gets read every
time, including in a dashboard six months later by someone who wasn't there. If the thing
that identifies the guilty party is optional to look at, it will be optional.

Your two sprints blaming the model is the exact cost I was trying to name. The label
didn't just hide the cause, it pointed at a suspect.

Collapse
 
mihai_leanzero profile image
Mihai Perdum •

Had one of these that cost more than a naming headache. A bulk write path on an old client engagement was gated on "this query shows zero issues" as the negative that licensed skipping the write. Zero turned out to mean two different things: genuinely empty, or my account's permissions couldn't see that project's issues at all. Same signal, same downstream branch, only one of the two readings was true, and the write had already gone out to real records by the time that surfaced. The rule since then is that a negative licensing an action has to be proven on the exact same object, not inferred from a query that might just be blind to it - a working query on a different project isn't proof.

Your EXEC_ARGUMENTS_MISMATCH case and mine are the same shape from different angles - yours is one name absorbing three distinct causes, mine was one boolean absorbing two. Same question either way: what does this signal actually rule out, and did I check that, or just assume it.

Collapse
 
kenielzep97 profile image
Self-Correcting Systems •

"A negative licensing an action has to be proven on the exact same object."

I'm taking that one. It's the general form and I only had the instance.

Yours is worse than mine and I want to be clear about why, because I don't think it's just
scale. My bad label sent a human down a wrong path and a human eventually noticed. Yours
had zero mean both "nothing there" and "you can't see," and the second reading is
indistinguishable from the first from inside the process. Permission blindness doesn't
raise, it returns empty. So the code was correct, the query was correct, and the write was
wrong.

Your closing question is the one I'd put above my desk. What does this signal actually
rule out, and did I check that, or assume it. An empty result rules out nothing until you
can prove you were looking at the thing you meant to look at.

I have this exact hole open in my own work right now, one layer up. My validator confirms
a citation is present and well formed. It does not confirm the cited record is the one
that was actually retrieved.

Collapse
 
jkming profile image
jkming •

The asymmetric registries worry me more than the naming did: the wider list is the one that fails quietly, which is exactly backwards. Before the full refactor, one cheap guard has worked well for me — a test that greps every failures.add('...') literal in the source and asserts each one is in the registry. That turns silent-drop drift into a red CI run instead of an invisible failure mode. The naming split also separates triage ownership nicely: INVALID is the caller's problem, MISMATCH is the model's, COMPARATOR_ERROR is yours. That alone saves the next reader an hour per incident.

Collapse
 
kenielzep97 profile image
Self-Correcting Systems •

i ran your guard on my local tree before answering, which is a different tree from origin. 34
distinct failures.add literals under scripts/, 37 unique names across the three registries, zero
orphans. so it is green today, which makes it a regression guard
rather than a bug finder, and that is still worth having.

then the counts did not match and the gap was the useful part. three registered names are never
added by failures.add at all:

BASE_MISMATCH        failures.push   pr2/preflight.mjs:33
DEPENDENCY_MISMATCH  failures.push   pr2/preflight.mjs:36
INPUT_INVALID        literal array   pr2/run.mjs:98
Enter fullscreen mode Exit fullscreen mode

so the guard as written would have missed those three, because there are three emission shapes and
it keys on one call name. the version that holds has to key on "a reason entering the failure set"
rather than on failures.add, or it inherits the same blind spot it was built to close.

on triage ownership, two things. there is a fourth name now, EXEC_EXPECTATION_INVALID, and it is
mine, because the expected and arriving sides were sharing a catch and a bad frozen fixture of mine
reported as a comparator error. so INVALID is no longer only the caller's problem, it is two names
with two owners.

the harder one is MISMATCH. it establishes that the call disagrees with my expectation. it does not
establish that the model was wrong, because the expectation can be the stale side, and in the post
before this one that is exactly what happened. so MISMATCH is not the model's column either. it is
the column that says two authorities disagree and does not say which one governs.

built locally, not pushed, so none of the four names are on origin yet. Vinh found it in this
thread.

Collapse
 
sameerqaisar17 profile image
Sameer Qaiser •

The line that stopped me: "A failure in the checking stage reads as a deviation in the thing being checked."

I'm a beginner — I just started writing Python tutorials this week. So I don't have a harness, I don't have fixtures, and I definitely don't have three catch blocks doing the same job. But this post taught me something I'll carry into every function I ever write:

A name is a claim about who's responsible.

When I catch an error and label it "invalid input," I'm not just logging a fact — I'm telling my future self (or anyone reading the logs) that the input was wrong. If the real failure was in my own parsing logic, I've just lied to myself in a way that's hard to detect.

The part about writing a test after the implementation that just asserts what the implementation does — that's the trap. I've done this already, even in my tiny Python scripts. I write code, it runs, I say "okay it works" and move on. But I never asked: "Does this test fail if the code is wrong?"

The control test that passes on both commits is such a clean idea. It's not there to prove the fix — it's there to prove the fix didn't break the thing it wasn't supposed to touch.

Saved this. The question at the end — "When this fires, whose fault does a reader assume it is?" — is going in my notes.

Collapse
 
kenielzep97 profile image
Self-Correcting Systems •

"a name is a claim about who's responsible" is the whole post in one line and you did not need a
harness to get there. i needed three catch blocks and someone else's comment.

and you already found the trap that cost me the most. writing the code, watching it run, then
writing the assertion from what it did.

to be exact about why that hurts, because i had it too broad at first: the test is not incapable of
failing. it will fail later, the day someone changes that behavior. what it cannot do is catch the
bug that was there when you wrote it, because it took its answer from the bug. so it locks the wrong
behavior in and then defends it. i did that twice on this patch and both times someone else caught
it.

the thing that works at any size, including a small script: before you write the assertion, write
down what the wrong answer would look like, then break the function on purpose and watch the test
go red. if you cannot make it fail, it is not testing, it is describing.

and keep one case that is supposed to stay green next to it. otherwise a fix that rejects
everything still passes a suite that only checks for rejections.

on "invalid input" specifically, since you named it. ask what you actually observed. you observed
that parsing did not succeed. you did not observe who produced the bad value. keeping the name at
what you saw costs one word and it is the difference between a log you can trust in six months and
one that quietly blames the wrong thing.

good luck with the tutorials. the instinct you showed here is the part people skip.

Collapse
 
jo-do profile image
Jo Do •

One label for three failure classes is a telemetry bug that reads as a tooling bug. The reducer collapsed distinct causes into a shared cardinality bucket, and every triage session after that started from the wrong premise. Labels are a contract with your future debugging self; merging three failures into one name is saving a constant and buying a week.

Collapse
 
kenielzep97 profile image
Self-Correcting Systems •

"a telemetry bug that reads as a tooling bug" is the cleaner framing, and it explains why it lasted. nothing was broken. the run did what it was built to do and reported it under a name that could not carry the distinction, so there was never a red flag to chase.

one correction on the mechanics: the shared bucket was EXEC_ARGUMENTS_MISMATCH, not a cardinality code. the two cardinality reasons in those arrays are noise from minimal fixtures with no sandbox event and no tool response, and they fire on every row regardless.

the triage premise point is the part with teeth, and it is worse than a single misread run. every count built on that name inherited the collapse. a pass rate over receipts, a human skimming for patterns, anything grouping by reason code, all of it was folding three causes into one bucket and then reasoning about the bucket.

"labels are a contract with your future debugging self" is the line i wish had been in the post.

Collapse
 
capestart profile image
CapeStart •

I liked that you left the stale commit message visible instead of cleaning it up. It makes the debugging process feel much more honest.

Collapse
 
kenielzep97 profile image
Self-Correcting Systems •

leaving it was the only honest option once i noticed. the alternative was a force push to rewrite a
public sha so a sentence would look right, and then the article would have been standing on a
receipt i had quietly edited to agree with it.

it also cost nothing. the correction is one paragraph and the stale line is still there for anyone
who wants to check which one is true.