DEV Community

Cover image for The 15-Line Test That Catches the #1 Killer of Operator Trust
Debashish Ghosal
Debashish Ghosal

Posted on AI-assisted

The 15-Line Test That Catches the #1 Killer of Operator Trust

The predecessor project had 56,869 anomalies across 100,000 traces. After fixing the noise, 11,294 remained. That's an 80% reduction — 45,575 of the original anomalies were false positives.

The operator saw 56,869 alerts. Investigated them. Found that four out of five were junk. And stopped trusting the system. The fix wasn't better detectors. It was removing and tightening the ones that fired too eagerly. A detector nobody believes is worse than no detector at all.

This is the failure mode that decides whether anyone uses anomaly detection in production. And the test that catches it before it ships is 15 lines of code that almost nobody writes.

Why per-detector tests miss this

Per-detector tests check each detector in isolation. In production, all 43 detectors run on the same trace at the same time. A clean trace can accidentally satisfy a detector's trigger conditions in ways that isolation testing never reveals.

A few that show up in practice:

  • An error rate of one bad call out of three is 33%, over a 30% threshold. The error was a transient blip. The trace is clean. The detector fires.
  • A clean trace has one-second gaps, well under a 30-second inactivity rule. But one slow API call takes 35 seconds. The trace is clean — the agent was legitimately waiting. The detector fires.
  • Three different tools look fine. Two search_kb calls in a row, with a threshold of 2, fires on an ordinary two-search workflow.

Each detector is correct in isolation. Each of those is a false positive. And per-detector tests can't see any of them, because they never run the detectors together.

This is from agentwatch, part of agentsec-ecosystem — a suite of open-source, harness-agnostic security tools for AI agents. 43 detectors, 226 scenarios, all green. The one test I'd keep if I could only keep one is the one that almost wasn't written.

The trace

The fix starts with a trace that is boring on purpose. Every value sits far from every threshold, so if anything fires there's no debate about whether it should have:

def _clean_trace():
    return (
        RunSummary(
            status="success",
            estimated_cost=0.12,        # threshold $5.00 — 42x margin
            duration_ms=4500,
            total_tool_calls=3,         # threshold 20 — 7x margin
            total_retries=0,            # threshold 5 — infinite margin
            total_interventions=0,      # threshold 3 — infinite margin
        ),
        [
            span("plan", status="ok", start=0),
            tool("search_kb", status="ok", start=1),
            tool("lookup_account", status="ok", start=2),
            tool("resolve", status="ok", start=3),   # all different — no loops
            output("Your password reset request has been completed successfully. "
                   "Here is a detailed summary of the steps taken and the final "
                   "outcome for your records and reference."),  # 200 chars — 4x margin
        ],
    )
Enter fullscreen mode Exit fullscreen mode

Distinct tools, no file-write tools, no network tools, no denied/error statuses, a comfortably long output. Nothing here is close to any rule's trigger. If something fires, it's wrong, and there's nothing to argue about.

The check

async def run_clean():
    summary, spans = _clean_trace()
    fired = []
    for det in create_all_detectors():
        res = await det.detect_async(summary, spans, pool=None)
        if res is not None:
            fired.append({
                "detector": det.anomaly_type,
                "severity": res.severity,
                "explanation": res.explanation,
            })
    return {"detectors": len(create_all_detectors()), "fired": fired, "ok": not fired}
Enter fullscreen mode Exit fullscreen mode

It comes back with an empty fired list across all 43 detectors. And when it doesn't, the list tells you which detector, at what severity, with what explanation — enough to diagnose and fix without further investigation.

tip: Run every detector together against a known-normal input and assert zero fires. It's the highest-leverage detector test there is, and it's usually the one missing. Fifteen lines, one trace, all detectors, zero expected fires.

What it catches

Five classes of bug that per-detector tests miss:

  1. Thresholds set too low. A detector that fires on 3 tool calls when the threshold should be 20. Per-detector tests check the threshold in isolation, but the clean trace catches misconfigured thresholds that fire on normal data.

  2. Buggy logic. A detector that fires when status == "success" because of a logic inversion (if status != "error" instead of if status == "error"). Per-detector tests might not test the "success" case for that detector.

  3. Misparse attributes. A detector that reads gen_ai.tool.name and fires when it's "resolve" because "resolve" matches a denylist pattern. Per-detector tests use specific tool names that might not trigger the denylist.

  4. Cross-detector cascades. Detector X's output triggers detector Y. anomaly_cluster checks how many distinct anomaly types fired on a trace. If 3 detectors fire falsely, anomaly_cluster fires too — a false positive cascade. Per-detector tests can't catch this because they run one detector at a time.

  5. State leakage. A previous scenario's state contaminates the clean trace. EmbeddingDriftDetector's baseline texts, set by an earlier scenario, cause it to compare the clean trace to a stale baseline and fire. Per-detector tests don't catch this because they don't run all detectors together after all scenarios.

The last two only exist when all detectors run together, which is exactly what production does and per-detector tests don't.

The predecessor data

The agent-exec-trace v2 report is the case study that made this test non-negotiable:

Detector Before fixes After fixes Reduction
premature_completion 35,930 0 Removed (fired on any error status)
argument_loop 5,768 0 Removed (missing args collapsed to empty)
redundant_tool_call 509 0 Removed (same args collapse)
pattern_loop 2,012 67 Deduped from loop
step_efficiency 1,797 184 Deduped from loop
wasted_tool_calls 1,451 2 Multi-tool gate added
Total 56,869 11,294 -80%

45,575 false positives. 80% of all anomalies were noise. The operator saw 56,869 alerts and stopped trusting the system. The fixes were not new detectors — they were removing or tightening existing detectors that fired too eagerly.

A clean-trace test would have caught premature_completion (fired on any error status, including normal errors) and argument_loop (fired when args were missing, which is normal for some tool types) before they reached production. Fifteen lines of code vs. 45,575 false positives.

premature_completion is the clearest example. It fired on any trace with an error status — but errors are normal in agent runs. A tool call fails, the agent retries, the run succeeds. The detector saw "error status" and fired, producing 35,930 anomalies on 100k traces, every one a false positive. The fix was to remove the detector entirely. A clean trace with one transient error would have caught it: the trace is normal, the detector fires, the test fails.

argument_loop is similar. It fired when tool arguments were missing — but missing args are normal for some tool types that don't use arguments. 5,768 false positives. The fix was removal. A clean trace that calls a no-argument tool would have caught it.

What I took away

False-positive noise is the #1 killer of operator trust — not missed detections. An alert system people ignore is worse than none.

Per-detector tests check logic; the clean-trace test checks integration. Both are necessary. Neither is sufficient alone.

The clean trace must be unambiguously normal. No edge cases, no borderline values. Every value should be far from every threshold. If a detector fires, there's no ambiguity — it's a false positive.

Run all detectors, not just the one you're testing. In production, they all run together. Cross-detector interactions and state leakage only appear when all detectors run on the same trace.

The fired list tells you exactly what to fix. Detector name, severity, explanation — enough to diagnose and fix the false positive without further investigation.

It's part of the 226-scenario harness in agentwatch v0.1.0, and it passes with zero fires across all 43 detectors.

The question worth asking of any anomaly system: how many false positives does it throw on a thousand clean inputs? If the answer isn't known, the clean-trace test is fifteen minutes away — and it's the cheapest way to find out whether the system earns trust or loses it.

The predecessor's 80% noise rate is the number to keep in mind. If four out of five alerts are junk, the operator stops looking. And a detector that isn't looked at is a detector that doesn't exist.

References

Top comments (0)