DEV Community

Cover image for 56 fault-injection tests passed. The one that injected nothing failed.
ashg2099
ashg2099

Posted on

56 fault-injection tests passed. The one that injected nothing failed.

Negative controls reveal flaky HLL estimators

I was building a tool that detects when data quietly changes meaning — a vendor switching units, a source dropping a field, an undocumented enum appearing. The kind of failure where every test passes and every job is green.

Claims about detection are cheap, so I built a benchmark. 56 seeded defects across fault type, magnitude, time window and pipeline layer. Each one has a known root cause. The tool profiles the pipeline, detects drift, walks the lineage graph, and names the node where the problem started. Score it against the node I actually broke.

It scored 55/56. I was pleased with myself for about a day.

Then I added the controls

A benchmark made only of faults can only tell you one thing: does the detector fire? It cannot tell you whether it fires too much. A detector that screams on every run scores 100% on that benchmark and is completely useless in production, because nobody reads an alert channel that cries wolf.

So I added four negative controls. Scenarios where the correct answer is silence:

  • control-null — rebuild and reprofile with no change at all
  • control-subthreshold-tip — tips up 3%, under the 5% threshold
  • control-subthreshold-extra — extras up 2%
  • control-subthreshold-tip-near-limit — tips up 4.5%, just under the line

A control passes only if it raises no high or critical signal.

control-null failed. Three high-severity signals on a run where nothing had changed.

What was actually happening

Two separate causes, compounding.

HyperLogLog. I was computing distinct counts with approx_count_distinct. It is fast, and for dashboards the approximation is fine. But HLL is a probabilistic sketch, and its estimate is not stable across runs — I measured variation up to 30% between two identical runs on the same data. My distinct-count threshold was 10%. The estimator's own noise was three times louder than the signal I was trying to detect.

Floating point. The second cause is the one I would not have guessed. DuckDB parallelises aggregates, so sum() and avg() accumulate in a non-deterministic order across threads. Floating-point addition is not associative: (a + b) + c and a + (b + c) can differ in the last bits. Two mathematically identical groups could therefore produce values that differed at the fifteenth decimal place — and that was enough to change which values counted as distinct.

The fix

Count exactly, and round floats before counting:

python
def _distinct_expr(col: str, data_type: str) -> str:
    if _is_float(data_type):
        return f"count(distinct round({col}, 6))"
    return f"count(distinct {col})"
Enter fullscreen mode Exit fullscreen mode

The same reasoning later applied to min and max, which I was recording as text:

python
def _bound_expr(fn: str, col: str, data_type: str) -> str:
    if _is_float(data_type):
        return f"round({fn}({col}), 6)::varchar"
    return f"{fn}({col})::varchar"
Enter fullscreen mode Exit fullscreen mode

Without that second one, a float min of 22575.66999999999 in one run and 22575.669999999995 in the next reads as a changed value. It is not a changed value. It is the same number, added up in a different order.

Two identical runs now produce zero signals.

What I actually took away

The obvious lesson is "use exact counts". That is not the interesting one.

The interesting one is that determinism is a precondition for detection, not a nice-to-have. A drift detector compares two measurements and calls the difference a signal. If your measurement process has its own variance, you have built a random number generator with a threshold on it. Every false positive it produces is indistinguishable from a real finding — and you will only find out after someone has stopped trusting the alerts.

And the second: fifty-six tests that expected something to happen never caught this. One test that expected nothing to happen caught it immediately.

That asymmetry generalises well beyond data tools. Most test suites are built entirely out of "given this input, assert this output". Very few contain "given no change, assert no output". The second kind is cheap to write and catches a category of bug the first kind is structurally blind to — anything where your system is noisy rather than wrong.

If you are building anything that detects, classifies, or alerts, write the test that expects silence. It will not pass on the first try.

The tool is Upstrace — column-level drift detection and lineage-based root-cause analysis for dbt projects. pip install upstrace; MIT-licensed. The full benchmark, including the one scenario that still fails, is in the repo.

Top comments (7)

Collapse
 
reidmarlow profile image
Reid Marlow •

The DuckDB parallel aggregate ordering issue is a nasty trap. Ran into the exact same floating-point associativity drift on windowed metrics across parquet partitions where grouping keys were unstable across chunk boundaries.The silence test advice matches what happens with agent evaluation harnesses too. People wire up fifty scenario benchmarks expecting a patch or tool call, but the baseline control where the repository needs zero changes often triggers speculative edits because the model refuses to emit an empty diff.

Collapse
 
ashg2099 profile image
ashg2099 •

The fix in the end was to stop estimating. count(distinct round(col, 6)) —
exact, deterministic, and the six decimal places absorb precisely the kind of
last-digit movement you're describing. Slower on wide tables, but a metric you
can't reproduce isn't really a metric.

Your agent-eval parallel is the more interesting half, and I hadn't thought of it
that way. "The model refuses to emit an empty diff" is the same shape as my
detector refusing to report zero signals. Both systems are being scored on
producing output, so producing nothing feels to them like failing.

Which suggests the negative control isn't really testing the system's accuracy at
all. It's testing whether the system is allowed to do nothing.

Collapse
 
irusik profile image
Ира Иващенко •

I really liked the idea of a “test that expects silence.” Sometimes checking that nothing changes can be even more important than checking for an expected result. Especially in systems where a small variation can look like a real problem and eventually make people lose trust in the alerts. A great reminder that we should test not only what a system does, but also what it shouldn’t do.

Collapse
 
ashg2099 profile image
ashg2099 •

Thanks. That framing came out of the failure, not before it. I built the four
negative controls expecting them to be a formality, and one of them immediately
started screaming at data that hadn't changed at all.

The part I didn't expect is how differently you debug a silence test. When a
detection test fails, you go and look at the detector. When a silence test
fails, you have to go and look at your own measurement, because you know for a
fact the data didn't move. That's how I ended up finding the estimator noise; the tool was behaving correctly; my metric was the thing that was wrong.

Collapse
 
raknaos profile image
Raknaos •

The control-null failure is the whole story here. 55/56 on seeded faults says nothing about a detector you'd trust in production, and HyperLogLog made it worse in the sneaky way: the estimator's own run-to-run noise (you measured up to 30% variation between identical runs) was three times louder than the 10% threshold, so the detector was mostly reading its own dice. The float ordering catch is the rarer one — non-associative addition across DuckDB's parallel threads moving a distinct-count in the 15th decimal is exactly what no single-input unit test will ever surface.

The asymmetry at the end generalizes past data tools: every alerting system I've run eventually needed a canary that expects silence, otherwise thresholds get calibrated against vibes. Now that the four controls exist, do you keep them wired into CI as regression gates (fail the build if control-null screams), or run them as scheduled checks against live pipelines? The former catches code changes, the latter catches data-side surprises — I've ended up wanting both.

Collapse
 
ashg2099 profile image
ashg2099 •

You've got the HLL part exactly right — the estimator's own noise being ~3x the
threshold meant the detector was mostly reading its own dice, and 55/56 was
measuring the seeded faults rather than the detector.

On your question: I ended up wanting both too, and now run both, as two separate
workflows.

The controls are CI regression gates. Two profile runs over unchanged data, then
drift --fail-on warning, build fails if anything at all is raised. That runs on
every push — and as of this week on both DuckDB and Postgres. This is the one
that catches code changes: a refactor that quietly alters how a metric is
computed shows up in the same pull request that introduced it.

Separately, a nightly job injects a real fault into the demo warehouse, asserts
it's caught, asserts the root cause is named correctly, and publishes the
resulting report. That's the data-side canary.

One decision worth flagging: the nightly deliberately does not post to Slack. It
injects a fault every single night on purpose, so alerting on it would train me
to ignore the alerts — which is the exact failure mode the tool exists to
prevent. That result stays in CI.

For a live pipeline I'd split it the same way: run the silence check against
production data, but keep fault injection in CI against a fixture, since you
can't inject into real data. The two answer different questions and I don't
think either one substitutes for the other.

Some comments may only be visible to logged-in visitors. Sign in to view all comments.