Pick a detector, run it over some benign traffic, and choose a threshold. Usually the threshold is an order statistic of the benign scores — the k-th largest, or a quantile — and you report something like false-positive rate ≤ 2% at 95% confidence.
Two questions tend to get skipped in that sentence: how many benign samples does the promise actually need, and how many attacks does it take to compare two rules? They have different answers, and a third question hides behind both.
1. How many benign samples does the budget need?
Say the threshold is the (k+1)-th largest benign score in a calibration set of size n. On a fresh benign sample, the probability that at most a fraction f of scores sit above it is a Beta tail:
P(realized FPR <= f) = P(Bin(n, f) >= k+1) = I_f(k+1, n-k)
For f = 2% and a 95% promise, solve P(Bin(n, 0.02) >= k+1) >= 0.95. At k=0 — threshold is the single largest benign score — there is a closed form: 1 - 0.98^n >= 0.95 gives n >= ln(0.05)/ln(0.98) = 148.3, so 149.
The whole curve at f = 2%, 95% confidence:
| k (outliers allowed above the threshold) | 0 | 1 | 2 | 3 | 4 | 5 | 6 | 7 |
|---|---|---|---|---|---|---|---|---|
| minimum n | 149 | 236 | 313 | 386 | 456 | 523 | 590 | 655 |
Roughly 65–87 more benign samples for each additional outlier you tolerate. And look at what 149 actually buys: at k=0, n=149 the median realized FPR is 0.46%, with a 5th–95th range of 0.03%–1.99%. The promise is about the tail, not the typical draw — the strict rule usually spends a quarter of the budget, which is a cost, not a virtue.
The trap is the obvious alternative. If instead you set the threshold at the empirical quantile, k = floor(0.02 n), you can never make the promise at any n. Its expected rate sits above the budget at every n — by about 0.98/(n+1) — so P(FPR <= 2%) never passes 0.63 (reached at n=49) and is still only 0.46 at n=2000. A rule that reads "set the threshold at the 98th percentile of the benign scores" reports the budget as if it had been achieved, and more data does not fix it — the rule is structurally unable to certify, not under-sampled.
2. How many attacks to compare two rules?
This one gets answered with two marginal recalls, which is the wrong tool. If your held-out set has n attacks and you compare a strict and a loose threshold, the difference of the two recalls is a paired question: only the attacks where the two disagree carry information. With discordance rate π_d,
SE(difference) = sqrt(pi_d / n)
At n = 629 that gives a 95% interval on the difference of about ±1.75 points at 5% discordance, ±2.5 at 10%, ±3.5 at 20%. The two-marginal interval sqrt(p(1-p)/n) would be ±3.2–4.0 — the pairing is real, just narrower than the marginals suggest.
At 80% power the smallest difference you can call is delta_min = 2.8016 * sqrt(pi_d / n):
| discordance | 95% CI half-width | δ_min at 80% power |
|---|---|---|
| 5% | ±1.75 pts | 2.50 pts |
| 10% | ±2.50 pts | 3.53 pts |
| 20% | ±3.50 pts | 5.00 pts |
So a 2-point decision — about the size a 0.5-point move in the benign tail tends to produce in recall — needs roughly 1,000–2,000 paired attacks. At 629, only a gap near 5 points is safely decidable. The failure mode is quiet: the test reports "no significant difference" whether the difference is zero or merely smaller than the sample can see, and the second case gets read as the first.
3. The variance you did not budget for
Both of the above hold the threshold fixed for the whole comparison. It is not fixed. Each rule produces one threshold per calibration draw, and that draw is small. Resample the benign calibration set, recompute both thresholds from each resample, and score them on the same fixed attacks — that measures how much the rule's recall moves on its own.
Across nine detectors on a public prompt-injection benchmark (629 attacks and 97 benign per detector):
- the strict rule's recall sd across calibration draws: 0.35, 1.3, 2.9, 4.3, 8.3, 10.2, 10.3, 10.8 points (eight with signal; the ninth flags nothing);
- the attack-sample half-width at n=629, for the same detectors: 0.9–3.6 points;
- so for the five detectors with real recall, the calibration draw is the larger term — 1.6× to 3.3× the attack sample.
And it does not cancel under pairing. Because the largest benign score and the fifth largest are set by different items, the two rules' recalls correlate only ρ ≈ 0.12–0.39 across draws, so the variance of the difference is close to the sum of the two variances, not the difference. A paired test at one fixed pair of thresholds therefore answers "which threshold is better", not "which rule is better". The honest budget is a sum: attack-sample variance plus calibration-draw variance, with the pre-registered decision applied to the expected gap over draws rather than to a single run.
One caveat on measuring that: the ordinary bootstrap is not consistent for a sample maximum. A with-replacement resample of 97 keeps the observed maximum only about 63% of the time and can never exceed it, so the number above is a floor. Reading the rate off m-out-of-n resamples instead reproduces it to within roughly 10% here — the artefact is real but does not move the size — and the residual bias is one-sided.
The general shape
A detection number is only as meaningful as the sample behind it, and there are three sample sizes in play, not one:
- enough benign traces that the budget is a promise rather than a point estimate;
- enough attacks to resolve the difference you intend to act on;
- enough calibration draws to say whether the rule or the luck of the draw decided it.
Report the resolution next to the number, or the number will be read as carrying more than it does.
Top comments (2)
This matches a tiered-threshold exchange I had this week almost exactly, and your k=0 closed form (n >= ln(alpha)/ln(1-f)) generalizes cleanly in a direction worth flagging: once you chain two or more of these detectors in a cascade (cheap stage filters, expensive stage confirms), the composed FPR budget isn't just the product of the per-stage budgets — the calibration-draw variance you describe in section 3 compounds across stages too, since each stage's threshold is drawn independently. If stage A's recall has sd ~3 points across calibration draws and stage B's has sd ~5, the cascade's effective variance is closer to the sum (same low-correlation argument you make for comparing two rules on the same attacks) than to either stage alone, so a cascade built from two detectors that each individually "pass" their n>=149-style certification can still fail to certify the combined budget unless you recompute n for the joint rule, not per-stage. Did you test whether chaining detectors in series changes the minimum-n math, or was every rule in the nine-detector set scored standalone?
Every rule in the nine-detector set was scored standalone — the benchmark ships per-detector rows and the curve (149 … 655) is single-rule, so no, I had not chained anything. I ran it on the same data (629 attacks, 97 benign per detector), and the answer separates by which composition you mean.
Filter → confirm (the cascade you describe). Here the composed budget is not a product of two draws — the confirm stage draws once — but it draws on the filtered population, so its calibration set is n₂ = q·n, with q the pass rate. Holding the same k=0 budget then needs n_total = n_single/q: 298 / 596 / 1490 at q = 0.5 / 0.25 / 0.1. So your "recompute n for the joint rule" is right, and the minimum-n math does change — by 1/q. One caveat: if the filter is score-correlated with the confirm, the filtered benign is not a random subsample and the budget can be worse than that size scaling alone predicts, because a strict filter keeps exactly the benign items the confirm already finds hard.
Conjunction at decision time (a sample alarms only when both fire). This is where "the composed FPR is a product of two independent draws" literally holds, and here it goes the other way. P(A∧B | benign) ≤ min(P(A), P(B)) is subset dominance, so the end-to-end 2% is certified by whichever stage already certifies it — no joint n is needed. The product target is easier too, not harder: with FPR ~ Beta(1,n) per stage, P(XY ≤ 4e-4) ≥ 0.95 at n ≈ 100, below the 149 a single 2% promise needs. The variance does compound — into a smaller rate.
The bill lands on recall, and it can be worse than the product. On the 629 attacks at the benchmark's own thresholds, both stages firing:
(I paired the two detectors' scores on the same attack.) Two individually mediocre detectors fire on different attacks, so AND-ing them can cost more than the product — the first two sit well below their independence estimate, which is what you would expect if each has its own blind spot. That is the composition a cascade has to budget for, and it is not the FPR budget.
So: the budget composes as a min (conjunction) or as 1/q added to n (gate); recall composes as the product or worse; and the section-3 calibration-draw variance rides along in both.