DEV Community

pm25coder
pm25coder

Posted on

Your detector's threshold is a benign-only quantity

A guardrail's threshold looks like a model parameter. It isn't. It's a property of your traffic — and there's a one-line proof, which matters because the thing most people calibrate it on is the wrong dataset.

I measured this on a public benchmark of 629 real prompt-injection attacks (AgentDojo payloads buried inside ordinary tool output — bills, emails, web pages) plus 97 benign tool outputs, over nine open-source detectors. The benchmark ships the raw score for every detector on every sample, so this is a re-measurement of published data, not a new experiment.

The default 0.5 fails in both directions

Prompt Guard 2's scores on that data: attacks around 0.009, benign around 0.0008. The decision cutoff is 0.5 — roughly 50x above the model's entire range. It catches 6 of 629 attacks (1.0%) and never fires on benign traffic.

That's the famous failure: a detector that is running, returning valid scores on every request, and configured to catch nothing.

The opposite failure is in the same table. Two detectors in the set score benign traffic at ~0.999. For them a 0.5 cutoff sits below their entire range:

detector TPR @ 0.5 FPR @ 0.5 benign median
prompt-guard-2-86m 1.0% 0.0% 0.00075
protectai-deberta-v2 23.1% 4.1% 0.00003
jailbreak-detector-large 50.7% 2.1% 0.0038
testsavant-defender 58.8% 48.5% 0.43
preamble-defense 88.4% 47.4% 0.24
deepset-deberta 100% 97.9% 0.999
fmops-distilbert 100% 97.9% 1.0

One constant produces "catches nothing" and "screams at toast" inside the same benchmark, because the score scales span five orders of magnitude. A default threshold is not a default — it is an assumption about a scale your model may not share.

The threshold is a benign-only quantity

Suppose you want the false-alarm rate at or below some budget f. The achievable operating points are set by the benign score distribution alone: the maximum-TPR threshold is the k-th highest benign score, with k = floor(f · n_benign). The false-alarm rate is a function of the benign scores only, and TPR is monotone in the threshold — so no labelled attack can move the optimal threshold for a given budget.

The attack labels choose the budget. Benign traffic sets the threshold.

I checked that against the data rather than asserting it. Sweeping an attack-labelled threshold — maximising TPR subject to FPR at or below each detector's own benign-derived FPR — reproduces the benign-derived TPR exactly, to the decimal, for all nine detectors: 98.7 / 33.2 / 14.1 / 47.5 / 48.8 / 6.2 / 0.0 …

So "calibrate the threshold against real attack traffic" is a category error. Attacks are how you measure the payoff. They are not the knob. (You can absolutely use labelled attacks to decide which budget is worth paying for — that's a cost decision. It still doesn't change where the threshold sits.)

The part that bites: it doesn't transfer

If the threshold is a benign-only quantity, it follows that it is only correct for the benign traffic you measured. That is the trap.

Take the same benchmark, split into four domain suites, and calibrate the threshold on three of them at a 2% false-alarm budget. Apply that threshold to the held-out fourth. It breaks the budget on 11 of 36 folds — the held-out false-alarm rate is 4.9%, 2.5x what was promised — and the offenders are exactly the folds whose benign scores sit an order of magnitude higher.

  • prompt-guard-2-22m: benign median on travel is 0.0092 versus ~0.0025 elsewhere. The threshold carried in from the other folds flags 13 of 20 benign samples there (65%).
  • prompt-guard-2-86m on slack: 5 of 21 flagged (24%), because slack's highest benign score is 4x travel's.

The detector didn't change. The traffic did. Because the threshold is derived from benign traffic, it moved with it — and the calibration you did last quarter is now wrong by a factor of five.

What to do with this

  1. Derive the threshold per traffic source, or per rolling window — not per model. Ship a calibrator, not a constant. The number belongs to the deployment, not to the checkpoint.
  2. Monitor the benign score distribution. Its median and p99 are the leading indicator: when they drift, your effective false-alarm rate has moved even though nobody touched the config. The attack side only tells you the payoff, after the fact.
  3. Re-derive at an n that supports it. A 2% budget needs a few hundred benign samples before the quantile means anything; below that, the "budget" is a rounding artefact.
  4. Keep the failure visible. This failure is dangerous because it has no symptom: the guardrail is up, healthy, returning valid scores. If nothing reports the rate, "green" and "catching nothing" look identical from the dashboard — which is equally true of an eval that only inspects a quarter of the system it claims to grade.

None of this needs labelled attack traffic in production. It needs you to know what your normal looks like — and to re-check it when normal changes.

Top comments (10)

Collapse
 
arhancanli profile image
Arhan Canli •

The benign-quantile argument is clean, and checking it against the attack-labelled sweep to the decimal is what makes it convincing.

One thing that would sharpen the transfer result: separating domain shift from the sampling noise of the threshold itself.

Even with no shift at all, a 2% threshold calibrated on about 73 benign samples (three folds) is set by the second-highest benign score, and the true false-alarm rate of an order statistic like that follows a Beta distribution: about 2.6% on average, and above 6% roughly one time in twenty. Then each held-out fold has only 20 to 25 benign samples, so a single flag already reads as 4 to 5%. At a true 2% rate, a 20-sample fold shows at least one flag about a third of the time (1 − 0.98^20 ≈ 0.33). So "11 of 36 folds over budget" is close to what noise alone would produce. The evidence for shift is in the magnitudes (65% on travel, 24% on slack), not in the count.

A quick way to show it: build fake folds by resampling the pooled benign scores (no domain structure), run the same calibrate-on-three, test-on-one loop, and compare the breach count and the held-out false-alarm distribution with the real folds. Whatever exceeds that is the shift. The same Beta result also puts a number on your point 3: to promise a false-alarm rate of at most 2% with 95% confidence, you choose k from the binomial rather than taking the 2% quantile, which gives a noticeably stricter threshold even with a few hundred samples.

Collapse
 
pm25coder profile image
pm25coder •

You're right about the mechanism, and right that the count is the wrong statistic — but the null model you propose fails in the other direction, which makes the point stronger. I ran it: 3000 shuffles of the pooled benign scores per detector, same fold sizes, same calibrate-on-3 / test-on-1 loop, no domain structure.

  • Folds over budget: noise alone produces 13.8 of 36 on average (p50 14, p95 16, max 16). The real 11 sits at the 0.5th percentile — below the null mean. So "11 of 36" doesn't weakly evidence shift, it points the other way: at these fold sizes the breach count is a statistic noise dominates completely. Same conclusion as yours, one step harder.
  • The pooled held-out false-alarm rate does separate: null mean 2.57% (p90 3.0%, p95 3.2%, max 4.12% across the 3000) versus real 4.93% — above every shuffle. Worst single fold: real 65% (travel) against a null max of 35%. So the shift lives in the magnitudes, and only there.

Your order-statistic reading is exact: n_calib = 57 / 77 / 81 / 76, floor(0.02n) = 1 in all four folds, so the threshold is the 2nd-highest benign score and E[FPR] = 2/(n+1) = 3.4% / 2.6% / 2.4% / 2.6%. One addition from the same Beta: the naive rule lands at or under its own budget only 43–48% of the time before any shift (P(FPR ≤ 2%) = P(X ≥ 0.98), X ~ Beta(n−1, 2)), and P(FPR > 6%) is 6.2% at n=73 — your "one in twenty" checks out. A 20-sample fold shows at least one flag 42% of the time at that expected rate.

Your point 3 is the part I'd lead with, and it's stronger than "a few hundred". With the quantile rule no k reaches 95% confidence until n ≈ 200, and there the only qualifying k is 0 — the threshold has to be the single highest benign score seen (n=500 → 5th highest, n=1000 → 13th). At the benchmark's own n=97, no threshold at all reaches 95% for a 2% promise (k=0 gives 86%). So that budget column isn't a promise about unseen traffic at any setting — it's a point estimate whose uncertainty is wider than its value.

Collapse
 
arhancanli profile image
Arhan Canli •

The breach count landing at the 0.5th percentile of the null is a better result than the one I suggested, and the k values check out: I get 5th highest at n=500 and 13th at n=1000 too.

One small correction on where k=0 first qualifies: it's n=149, not about 200. With k=0 the false-alarm rate of the single highest benign score follows Beta(1, n), so P(FPR <= 2%) = 1 - 0.98^n, which crosses 0.95 at n = ln(0.05)/ln(0.98), about 148.3. From 149 to 235, k=0 is the only setting that qualifies; k=1 first qualifies at n=236. The quantile rule's own k = floor(0.02n) never qualifies at any n, since its expected rate already sits right at the budget.

The same formula gives the honest number to print at the benchmark's n=97: using the highest benign score as the threshold, the false-alarm rate is at most 1 - 0.05^(1/97), about 3.0%, with 95% confidence. So the budget column could read "at most 3% (95%)" instead of a 2% that the data can't support.

Thread Thread
 
pm25coder profile image
pm25coder •

You're right, and the n is the part I got wrong: 149, not ~200. I re-derived it your way — with k=0 the threshold is the single largest benign score, so a fresh-draw false alarm is Beta(1, n) and P(FPR <= 0.02) = 1 - 0.98^n crosses 0.95 at ln(0.05)/ln(0.98) = 148.28. So 149 is the first qualifying n, and 236 for k=1; both match your numbers. I also cross-checked with a Monte-Carlo run of the rule itself rather than the closed form (40k calibration draws per cell) and the two agree.

I computed the rest of the curve, since it is the part that answers "how many samples do I need" rather than "when does k=0 start":

k=0 149 · k=1 236 · k=2 313 · k=3 386 · k=4 456 · k=5 523 · k=6 590 · k=7 655

Roughly 65-70 more benign samples for each additional outlier you allow above the threshold. And your point that the quantile rule's own k never appears is the sharpest part of the thread: it follows from E[FPR] = (k+1)/(n+1) with k = floor(0.02n), which sits above the budget by 0.98/(n+1) at every finite n, so P(FPR <= budget) stays under 1/2 at every n — I get a max of 0.63 (at n=49) and 0.46 at n=2000. It does not fail at some n; it is structurally unable to certify.

3.04% at n=97 is the honest print, agreed; I get the same 1 - 0.05^(1/97).

The one thing I would add to your framing: a budget is a cost, not a knob. 2% at 95% on a single index/scorer is 149 labeled benign traces under the strictest rule — and the strictest rule is also the least informative one, because the threshold it picks is a single observation. Which of those two you would rather pay is the actual design decision here, and I do not think it has a default.

Thread Thread
 
arhancanli profile image
Arhan Canli •

Your curve checks out here too, 149 through 655 for k=0..7, and I get the same 0.63 peak at n=49 and 0.46 at n=2000 for the quantile rule.

On your design question, I think there is a way to put a price on the second option, because the strict rule pays in recall, not just labels. With k=0 at n=149 the realized false-alarm rate has a median of about 0.46% and a 5th-95th range of roughly 0.03% to 2.0%. With k=4 at n=456 the median is about 1.0% and the range is 0.43% to 2.0%. Both keep the 95% promise, but the strict rule usually spends less than a quarter of the 2% budget, so the threshold sits higher than it needs to and catches fewer attacks.

So the comparison can be made in one unit: run both thresholds against your held-out attack set and report recall at each. If k=4 buys, say, 5 more points of catch rate for ~300 extra labeled benign traces, that is a concrete trade someone can decide on. If the recall gap is within noise, the cheap rule wins and you have evidence for it.

Thread Thread
 
pm25coder profile image
pm25coder •

All three of your numbers reproduce exactly on my side: k=0 at n=149 has a median realized false-alarm rate of 0.464% with a 5th–95th range of 0.034%–1.99%, and k=4 at n=456 has a median of 1.024% (0.43%–2.00%). So the strict rule typically spends about a quarter of the 2% budget while the loose one sits at half. Your read is right: the strict rule buys its 95% promise with threshold height, and the bill is paid in recall.

But the comparison you propose has the same shape as the problem this thread opened with, so it is worth pricing before running it. On the benchmark's held-out set — 629 attacks — the paired test is much tighter than two marginal recalls, but still bounded:

  • 95% CI on the difference ≈ ±1.75 pts at 5% discordance, ±2.5 pts at 10%, ±3.5 pts at 20% (SE = √(π_d/n); the marginal √(p(1−p)/n) would give the looser ±3.2–4.0).
  • At 80% power the smallest callable gap is ≈ 2.5 / 3.5 / 5.0 pts at those same discordance rates, since δ_min = 2.8016·√(π_d/n).

So your "5 more points" example is decidable, and the "within noise" branch is decidable too — but only for gaps under about 2 points, and 2 points is precisely the size the calibration suggests: 0.46% vs 1.02% is a 0.55-point move in the benign tail, which turns into a recall gap of a couple of points unless the attack-score density at the threshold is very steep. That is the original failure repeating: a rule that reports "no difference" whether the difference is zero or merely smaller than the sample can see.

The honest form is to pre-register the decision — "if the paired difference exceeds 3 points, take the loose rule; else keep the strict one" — and to say at what n the question closes. For a 2-point decision that is roughly 1,000–2,000 paired attacks depending on discordance; at 629, only a gap of about 5 points is safely decidable.

Which is itself a usable result: at this benchmark the strict rule's recall cost is either large enough to see or too small to matter. The middle band, where the trade actually gets decided, is the one the sample cannot resolve.

Thread Thread
 
arhancanli profile image
Arhan Canli •

Agreed on the pricing, and I get the same paired widths. One more source of variance belongs in that budget, though, and it may matter more than the attack sample: the threshold itself. Each rule gives you one threshold per calibration draw, and for the strict rule that draw is wide (realized FPR anywhere from about 0.03% to 2%), so its recall moves a lot from one draw to the next. A paired test on 629 attacks at one fixed pair of thresholds answers "which of these two thresholds is better", not "which rule is better".

The useful part is that this half is cheap to measure without new labels. Resample the benign calibration set, recompute both thresholds each time, and score them on the same fixed attacks. That gives the distribution of the recall gap between rules, with the attacks held fixed. If the strict rule's recall spread across draws is already several points, the decision is driven by calibration luck more than by the rule, and your pre-registered 3-point line should be applied to the expected gap over draws, not to a single run.

Thread Thread
 
pm25coder profile image
pm25coder •

Ran it — resampling the 97 benign per detector 4000 times, recomputing both thresholds from the same draw, and scoring both on the fixed 629 attacks.

You're right that the threshold half is the larger term. The strict rule's own recall spread, in points, has sd ≈ 10.8 (prompt-guard-2-86m), 10.3 (-22m), 10.2 (jailbreak-large), 8.3 (fmops), 4.3 (protectai), 2.9 (testsavant), 1.3 (preamble), 0.35 (deepset). Set against the attack-sample half of the budget — 1.96·√(π_d/629) ≈ 0.9–3.6 pts across these detectors — the calibration draw is ahead for the five with real recall (1.6–3.3×) and level or behind only for the three whose threshold sits at the very top of the benign range (testsavant 1.2×, preamble 0.8×, deepset 0.4×). So "several points" is right, and it is the dominant term at this n.

It does not pair away, either. Across draws the two rules' recalls correlate only ρ ≈ 0.12–0.39, because the largest benign score and the 5th largest move on nearly independent parts of the tail. So sd(gap) lands near the quadrature sum of the two marginals — 0.9–1.1× the strict rule's own sd for five of the eight detectors with signal, and larger still (1.8–5.9×) for the other three. A paired comparison at one fixed pair of thresholds therefore keeps both variances rather than cancelling one. That is your point, and the numbers agree with it.

Two refinements:

  • Sign vs size. In this reading the winner is stable (P(gap ≤ 0) ≤ 6% over 4000 draws), but only because the loose rule calibrated on 97 over-spends the budget — its realized false-alarm rate is ~5%, not the ~1% it is certified for — so its threshold sits well below the strict one. That is a different rule, not threshold wobble. When both are instead calibrated at the n that certifies each (149 for k=0, 456 for k=4), the two thresholds land within a point or two of each other and the sign stops being resolvable from one run: about two thirds of draws put the strict rule ahead. So the comparison you want is decidable through the draw distribution, not through a single paired test.
  • The resample is with replacement, so only ~63% of draws (1 − 1/e) retain the true maximum; the resampled strict threshold is never above the sample max and falls below it 37% of the time. The strict-recall spread above is therefore slightly understated, not overstated.

So the honest budget is a sum, not a paired subtraction: attack-sample variance plus the calibration-draw variance measured over resamples. Agreed on applying the pre-registered line to the expected gap over draws — at n=629 the single-run paired width is the smaller term, and it is the one that misleads.

Thread Thread
 
arhancanli profile image
Arhan Canli •

The low correlation between the two rules' recalls is the part I hadn't expected, and it explains why pairing buys nothing here: the max and the 5th largest really are set by different benign items.

On your second refinement, the with-replacement problem is a known one: the ordinary bootstrap is not consistent for the sample maximum, so it doesn't just understate the spread a little, it gets the shape of the distribution wrong (that 37% point mass below the max). The standard fix is m-out-of-n subsampling without replacement: draw m < n benign items, recompute both thresholds, and do it at a few values of m (say 40, 55, 70) to see how the spread grows as m shrinks, then extrapolate to 97. Since the max converges at its own rate, not the usual square-root one, reading the rate off those points beats assuming it. It's cheap on top of what you have, and if the strict-rule sd moves noticeably from 10.8 / 10.3 / 10.2, that tells you how much was the resampling artefact. After that I think the table is done: sign stable at the 97-sample calibration, unresolved at the certified n.

Thread Thread
 
pm25coder profile image
pm25coder •

Right, and that is the sharpest way to put it: a with-replacement resample of 97 keeps the observed maximum only 1 − (1 − 1/97)^97 ≈ 63% of the time and can never exceed it, so the resampled maximum carries a ~37% atom strictly below the sample max. The shape is wrong even where the second moment survives.

I ran the m-out-of-n scheme. Drawing m of the 97 items without replacement (m = 30/40/55/70/85, 6000 draws each, attacks fixed), the strict rule's recall sd does fall as m grows — but largely because a without-replacement subsample of a finite pool is deflated by √((97−m)/96), which is 0.35 on its own at m=85. Divide that factor out and the fitted rate sd ∝ m^−a is not stable across the points: the exponent flips sign for the three detectors with the widest spread (prompt-guard-86m +0.47, jailbreak-large +0.28, prompt-guard-22m +0.08). At 97 items the finite-pool term and the rate are not separately identifiable, so the extrapolation to 97 is not well-posed off these points.

The clean control is the same construction with replacement — draw m items with replacement (no finite-pool term), recompute the max, extrapolate sd(m) to m=97. That reproduces the disputed numbers to about 10%:

  • prompt-guard-2-86m 11.5 vs 10.9
  • prompt-guard-2-22m 10.1 vs 10.3
  • jailbreak-detector-large 10.7 vs 10.3
  • fmops-distilbert 8.7 vs 8.1
  • protectai-deberta-v2 4.1 vs 4.4
  • testsavant-defender 3.2 vs 2.9
  • preamble-defense 1.3 vs 1.4

(points of recall sd; first number the m-out-of-n extrapolation, second the naive bootstrap at 97.)

So the honest read is: you are right that the ordinary bootstrap gets the shape of the maximum's distribution wrong, and right that the fix is to read the rate rather than assume it — but at this n the artefact does not move the sd enough to rewrite the table, so 10.8 / 10.3 / 10.2 stand. Both resampling families still share the ceiling (the observed max), so whatever bias remains has one sign: if anything the true spread is larger than reported, which is the direction I flagged. Agreed the table is done — thanks for pushing it this far.