The predecessor project had 56,869 anomalies across 100,000 traces. After fixing the noise, 11,294 remained. That's an 80% reduction — 45,575 of the original anomalies were false positives.
The operator saw 56,869 alerts. Investigated them. Found that four out of five were junk. And stopped trusting the system. The fix wasn't better detectors. It was removing and tightening the ones that fired too eagerly. A detector nobody believes is worse than no detector at all.
This is the failure mode that decides whether anyone uses anomaly detection in production. And the test that catches it before it ships is 15 lines of code that almost nobody writes.
Why per-detector tests miss this
Per-detector tests check each detector in isolation. In production, all 43 detectors run on the same trace at the same time. A clean trace can accidentally satisfy a detector's trigger conditions in ways that isolation testing never reveals.
A few that show up in practice:
- An error rate of one bad call out of three is 33%, over a 30% threshold. The error was a transient blip. The trace is clean. The detector fires.
- A clean trace has one-second gaps, well under a 30-second inactivity rule. But one slow API call takes 35 seconds. The trace is clean — the agent was legitimately waiting. The detector fires.
- Three different tools look fine. Two
search_kbcalls in a row, with a threshold of 2, fires on an ordinary two-search workflow.
Each detector is correct in isolation. Each of those is a false positive. And per-detector tests can't see any of them, because they never run the detectors together.
This is from agentwatch, part of agentsec-ecosystem — a suite of open-source, harness-agnostic security tools for AI agents. 43 detectors, 226 scenarios, all green. The one test I'd keep if I could only keep one is the one that almost wasn't written.
The trace
The fix starts with a trace that is boring on purpose. Every value sits far from every threshold, so if anything fires there's no debate about whether it should have:
def _clean_trace():
return (
RunSummary(
status="success",
estimated_cost=0.12, # threshold $5.00 — 42x margin
duration_ms=4500,
total_tool_calls=3, # threshold 20 — 7x margin
total_retries=0, # threshold 5 — infinite margin
total_interventions=0, # threshold 3 — infinite margin
),
[
span("plan", status="ok", start=0),
tool("search_kb", status="ok", start=1),
tool("lookup_account", status="ok", start=2),
tool("resolve", status="ok", start=3), # all different — no loops
output("Your password reset request has been completed successfully. "
"Here is a detailed summary of the steps taken and the final "
"outcome for your records and reference."), # 200 chars — 4x margin
],
)
Distinct tools, no file-write tools, no network tools, no denied/error statuses, a comfortably long output. Nothing here is close to any rule's trigger. If something fires, it's wrong, and there's nothing to argue about.
The check
async def run_clean():
summary, spans = _clean_trace()
fired = []
for det in create_all_detectors():
res = await det.detect_async(summary, spans, pool=None)
if res is not None:
fired.append({
"detector": det.anomaly_type,
"severity": res.severity,
"explanation": res.explanation,
})
return {"detectors": len(create_all_detectors()), "fired": fired, "ok": not fired}
It comes back with an empty fired list across all 43 detectors. And when it doesn't, the list tells you which detector, at what severity, with what explanation — enough to diagnose and fix without further investigation.
tip: Run every detector together against a known-normal input and assert zero fires. It's the highest-leverage detector test there is, and it's usually the one missing. Fifteen lines, one trace, all detectors, zero expected fires.
What it catches
Five classes of bug that per-detector tests miss:
Thresholds set too low. A detector that fires on 3 tool calls when the threshold should be 20. Per-detector tests check the threshold in isolation, but the clean trace catches misconfigured thresholds that fire on normal data.
Buggy logic. A detector that fires when
status == "success"because of a logic inversion (if status != "error"instead ofif status == "error"). Per-detector tests might not test the "success" case for that detector.Misparse attributes. A detector that reads
gen_ai.tool.nameand fires when it's "resolve" because "resolve" matches a denylist pattern. Per-detector tests use specific tool names that might not trigger the denylist.Cross-detector cascades. Detector X's output triggers detector Y.
anomaly_clusterchecks how many distinct anomaly types fired on a trace. If 3 detectors fire falsely,anomaly_clusterfires too — a false positive cascade. Per-detector tests can't catch this because they run one detector at a time.State leakage. A previous scenario's state contaminates the clean trace.
EmbeddingDriftDetector's baseline texts, set by an earlier scenario, cause it to compare the clean trace to a stale baseline and fire. Per-detector tests don't catch this because they don't run all detectors together after all scenarios.
The last two only exist when all detectors run together, which is exactly what production does and per-detector tests don't.
The predecessor data
The agent-exec-trace v2 report is the case study that made this test non-negotiable:
| Detector | Before fixes | After fixes | Reduction |
|---|---|---|---|
premature_completion |
35,930 | 0 | Removed (fired on any error status) |
argument_loop |
5,768 | 0 | Removed (missing args collapsed to empty) |
redundant_tool_call |
509 | 0 | Removed (same args collapse) |
pattern_loop |
2,012 | 67 | Deduped from loop |
step_efficiency |
1,797 | 184 | Deduped from loop |
wasted_tool_calls |
1,451 | 2 | Multi-tool gate added |
| Total | 56,869 | 11,294 | -80% |
45,575 false positives. 80% of all anomalies were noise. The operator saw 56,869 alerts and stopped trusting the system. The fixes were not new detectors — they were removing or tightening existing detectors that fired too eagerly.
A clean-trace test would have caught premature_completion (fired on any error status, including normal errors) and argument_loop (fired when args were missing, which is normal for some tool types) before they reached production. Fifteen lines of code vs. 45,575 false positives.
premature_completion is the clearest example. It fired on any trace with an error status — but errors are normal in agent runs. A tool call fails, the agent retries, the run succeeds. The detector saw "error status" and fired, producing 35,930 anomalies on 100k traces, every one a false positive. The fix was to remove the detector entirely. A clean trace with one transient error would have caught it: the trace is normal, the detector fires, the test fails.
argument_loop is similar. It fired when tool arguments were missing — but missing args are normal for some tool types that don't use arguments. 5,768 false positives. The fix was removal. A clean trace that calls a no-argument tool would have caught it.
What I took away
False-positive noise is the #1 killer of operator trust — not missed detections. An alert system people ignore is worse than none.
Per-detector tests check logic; the clean-trace test checks integration. Both are necessary. Neither is sufficient alone.
The clean trace must be unambiguously normal. No edge cases, no borderline values. Every value should be far from every threshold. If a detector fires, there's no ambiguity — it's a false positive.
Run all detectors, not just the one you're testing. In production, they all run together. Cross-detector interactions and state leakage only appear when all detectors run on the same trace.
The fired list tells you exactly what to fix. Detector name, severity, explanation — enough to diagnose and fix the false positive without further investigation.
It's part of the 226-scenario harness in agentwatch v0.1.0, and it passes with zero fires across all 43 detectors.
The question worth asking of any anomaly system: how many false positives does it throw on a thousand clean inputs? If the answer isn't known, the clean-trace test is fifteen minutes away — and it's the cheapest way to find out whether the system earns trust or loses it.
The predecessor's 80% noise rate is the number to keep in mind. If four out of five alerts are junk, the operator stops looking. And a detector that isn't looked at is a detector that doesn't exist.
References
- Release: agentwatch v0.1.0 on GitHub · on PyPI
- Suite: agentsec-ecosystem — the full set of agent-security tools
- Field test report (clean-run check, FT-15)
- Harness:
scenario_validation.py - Predecessor noise data:
agent-exec-tracefield-test report v2
Top comments (0)