There are three ways a security guardrail can fail you.
Two of them are loud. One of them is a serial killer.
It's down. Crashed, misconfigured...
For further actions, you may consider blocking this person and/or reporting abuse
I'd extend the armed check all the way to the protected action. A detector can return BLOCK while an adapter logs the decision and still invokes the tool. A synthetic request against an isolated test resource should assert both that the policy rejected it and that no downstream write occurred; pair it with a benign request so a broken path that rejects everything cannot pass.
The shared-wrapper caveat also makes a useful next experiment: hold out wrapper families as well as domains, and report the false-positive count and denominator per split. With 97 benign examples overall, a couple of classifications materially move the rate. That would help separate threshold calibration from recognizing a recurring packaging pattern.
Yes — this is the version of the armed check that can't lie. A detector returning BLOCK proves nothing if the adapter still invokes the tool; the only honest assertion is at the protected resource: a synthetic request against an isolated test resource must show (a) the policy rejected it and (b) zero downstream write. Pair it with a benign request so a path that rejects everything can't pass. That's liveness and effectiveness in one probe.
And holding out wrapper families as well as domains is exactly the experiment I owe the post. With 97 benign examples a couple of flips move the rate materially, so per split I'd report the raw false-positive count and denominator, not a percentage — a rate over a small integer implies a precision it doesn't have. That's what separates "calibrated the threshold" from "memorized the packaging." Adding this. Thanks for sharpening it.
Thanks for taking this into the follow-up. One detail worth making explicit in the harness: when is the zero-write assertion evaluated? If an adapter queues work, an immediate read can pass before a delayed write arrives. I'd correlate the probe with a unique operation ID and drive the test queue to completion before checking the isolated resource. In a live probe, report the observation window rather than claiming the write can never happen.
The benign control should traverse that same adapter and queue. That makes a missing worker distinguishable from successful enforcement.
The distinction between liveness and effectiveness is the sharpest framing here. A health check proving the guardrail is running says nothing about whether it's stopping anything, and the Prompt Guard 2 numbers make that concrete: the attack scoring roughly 10x higher than benign text, both sitting far below 0.5, so every verdict reads "allow" while the ranking is nearly perfect. Moving the threshold to 0.003 and going from 1% to 99% shows how much one config value can hide.
I also appreciated the honesty clauses. Noting that the 99% may reflect AgentDojo's shared wrapper template rather than general detection, and that the real claim is the miscalibration, keeps the post from becoming a screenshot-able overclaim. Insisting on measuring at a false-alarm budget, with the "brick over the deny button scores 100%" point, was a useful reminder too.
To answer your closing question, the closest I've seen is an alert routed to a muted channel for months. Everything looked wired up, and nothing was actually being heard. Firing a known attack through the live system on a schedule seems like the only reliable fix. Thanks for open-sourcing the benchmark.
That's the purest version of it. The detection worked, the routing worked, and the last step, a human hearing it, was silently off. Every component would pass its own health check.
It also shows why the fix has to be end to end:
A scheduled known-bad event only counts if the assertion sits at the far end, with someone or something acknowledging it.
How did the muted channel finally get noticed: an incident, or someone cleaning up their notification settings?
@kartik-nvjk > How are you calibrating thresholds against real attack traffic rather than the default 0.5?
Honest answer first: I don't have production attack traffic. The benchmark calibrates on AgentDojo, tuning on 3 domains and measuring on the 4th. So what I can give you is the procedure, not a number to copy:
The weak point is step 3: known attacks are a proxy for real ones, and 97 benign samples is a small stick for measuring a 2% budget.
"Graded green while checking a quarter of it" is the same failure one layer up. Did Bedrock's eval tell you which three quarters it skipped, or did you have to find that out yourself?
@kartik-nvjk Short version: the threshold is a benign-only quantity — provable, and it changes what "calibrate against real attack traffic" should mean.
For a false-alarm budget f, the max-TPR threshold is the k-th highest benign score (k = floor(f·n)). Attack labels cannot move it: the FPR is a function of the benign scores alone, and TPR is monotone in the threshold. So the labelled attacks choose the budget; benign traffic sets the threshold. I swept an attack-labelled threshold over all nine detectors in this benchmark's own
bench/results/at_budget_2pct.json— at each detector's benign-derived FPR, the best attack-labelled TPR is the same number to the decimal (98.7 / 33.2 / 14.1 / 47.5 / 48.8 / 6.2 / 0.0 … , 9/9).Your ~50x reproduces exactly: at t=0.5 the same file gives prompt-guard-2-86m 6/629 = 1.0% — the headline number — with a false-alarm rate of 0.0%. 0.5 sits ~53x above its attack median and ~90x above its highest benign score.
And the reason a single default can't be shipped: the same 0.5 fails in both directions in one benchmark.
deepset-debertaandfmops-distilbertscore benign traffic at ~0.999, so 0.5 lands below their entire range — TPR 100% / FPR 97.9%. The constant that makes Prompt Guard catch nothing makes those two scream at toast.The part that bites once you do calibrate on real traffic: it doesn't transfer, because it's the benign distribution that isn't stationary across sources. Calibrating on three folds and applying the threshold to the fourth breaks the promised 2% budget on 11 of 36 folds (held-out FPR 4.9%, 2.5x the budget) — and the offenders are exactly the folds whose benign scores sit an order of magnitude higher. prompt-guard-2-22m: benign p50 on travel is 0.0092 vs ~0.0025 elsewhere, so the transferred threshold flags 13/20 benign. prompt-guard-2-86m on slack: 5/21 flagged (24%), because slack's highest benign score is 4x travel's.
So the shape I'd take away: derive the threshold per traffic source (or per rolling window), not per model, and watch the benign score distribution — its p50/p99 is the leading indicator; the attack side only tells you the payoff, after the fact. It's also the cheap version of your Bedrock point: an eval that reports on a quarter of the agent has the same defect as a threshold nobody asked what its traffic looks like.
That's the whole post in six words — stealing it.
On calibration, the counterintuitive part (which @pm25coder just proved above): you don't calibrate against attack traffic at all. The threshold is a benign-only quantity — for a false-alarm budget
f, it's the k-th highest benign score (k = floor(f·n)). Attacks only choose the budget you're willing to spend; your benign traffic sets the cutoff. So "calibrate against real attack traffic" is the wrong frame — calibrate against real benign traffic, per source, and let a known-attack canary tell you the payoff after the fact.Your Bedrock example is the same disease one layer up: an eval that grades green while covering a quarter of the agent is a threshold nobody asked what its traffic looks like.
Are you deriving one threshold globally, or per source? The benign distribution is the part that drifts — same detector, different source, different cutoff.
Ran
bench/at_budget.pyover your own saved scores (bench/results/at_budget_2pct.json) before writing this, because the port is what makes the rest checkable — all nine detectors reproduce both saved columns exactly (in-sample and cross-domain) once the suite rotation is ordered workspace/travel/banking/slack, so the numbers below are yours, not a re-derivation.Three things the pooled figure doesn't say.
1. The 2% budget is an in-sample property; the unseen column is a different quantity.
caught_at_budgetenforces the budget when picking (allowed = floor(budget * len(calib)), per fold on the calibration split), then measures on the held-out suite, where it can exceed it — and does, for eight of the nine (the regex baseline is the exception): 5.2% for Prompt Guard 2 86m, 13.4% for the 22m, 9.3% for testsavant. So "99% at a 2% false-alarm budget" splices an in-sample budget onto a cross-domain TPR. Both are honest numbers; the sentence that gets screenshotted has combined two.2. The pooled TPR hides a fold spread as large as the rate it reports. Per held-out suite, prompt-guard-2-22m is 100% on travel, 22.9% on workspace, 16.0% on banking, 1.0% on slack — pooled 35%, spread 99 points. jailbreak-detector-large runs 17.1% → 100% (pooled 51%). For six of the nine the spread is at least the pooled rate itself.
cross_domainalready has each fold incand accumulates it away (caught += c); keeping the list and printing min/max is two lines, and it changes what "the number to trust" means.3. For a control, the worst fold is the number to trust — not the mean. That's your own thesis one level down: a pooled average is a liveness statistic (it says the detector ran on four suites), while the min fold is closer to an effectiveness one. Prompt Guard 2 86m is the reassuring case — 97–100%, a 3.3-point spread — and most of the field isn't.
The benign-sample size and the rotating-canary points above are the other half of it; both are cheap.
This is exactly the kind of comment I was hoping this post would attract.
You’re right on the 2% point — I compressed two different quantities into one sentence. The budget is enforced during calibration, while the reported held-out/cross-domain catch rate is a different measurement. Putting those side by side without making that distinction explicit makes the headline number easier to screenshot than to interpret.
The fold spread point is even more important. A pooled TPR can look reassuring while one domain is basically a miss. I like the idea of keeping the per-fold values visible and reporting the worst fold alongside the pooled number.
That actually fits the thesis of the article better: the number isn't the measurement unless you preserve the conditions under which it was produced.
I'm going to fix this rather than defend the original wording. Thanks for actually running the code and checking it against the saved results.
This is a really good catch.
I was treating “0 false positives out of 97” as if it established a 2% false-alarm rate, when statistically it really only tells us what happened in that small sample.
Your point about the benign denominator is exactly the kind of thing that gets lost when a benchmark focuses on the attack-success column. At a threshold this close to the benign-score distribution, the uncertainty around the benign sample can matter more than another decimal place on TPR.
And I really like the held-out rotating-canary idea. A fixed canary can gradually become a memorization test instead of a security test.
So the revised measurement should probably report both:
effectiveness on the held-out pool + confidence around the benign false-alarm estimate.
That’s a much more useful production story than simply saying “99% caught.”
Great addition.
On the min-fold question — same re-run, and the answer changed twice while I was measuring it.
Headline pair: yes. But attach n and an interval to the min fold. Min-fold is by construction the smallest numerator, so it is also the widest interval. Your min folds, Wilson 95%, ordered: 97% [94–98], 26% [19–34], 17% [12–24], 1% [0–5], then five of the nine bottoming out at 0/140–240 → [0–3%]. A bare "17%" reads like a measurement; "17% [12–24], 140 attacks" reads like one.
The surprise: the min-fold leaderboard is mostly a tie. Walking the ranked min folds, 6 of the 8 adjacent pairs have overlapping 95% intervals — Prompt Guard 2 86m is cleanly separated at the top, and below third place the ordering is not resolvable at these n (four detectors are all 0% [0–3%]). So print the interval with the rank, or replace the rank with the interval below the top — a rank implies a separation the data doesn't have.
On readability: the per-suite row is four numbers — that is not what makes it unreadable. Put the 4-cell row under the headline and the full table in
results/*.json; what actually costs a reader is three quantities with different denominators (budget / false-alarm rate / TPR) sitting next to each other unlabelled. Label the denominators and you can afford the detail.One more, same re-run, and it cuts the same way: at this corpus size the 2% budget is degenerate.
allowed = floor(0.02 * len(calib)), and calib is 57/77/81/76 benign cases depending on which suite is held out → 1 in all four folds. So cross-domain, "2%" is not a rate at all; it is "exactly one benign case above the line". Reporting the budget as an allowed count until the benign corpus grows (your ~150+ fix) makes that visible from the header — and it is a third reason those two numbers shouldn't share a sentence.Taking all four together, they collapse into one correction I should have made myself: every headline in that benchmark is a percentage printed over a small integer, and a percentage over a small integer implies a resolution the count doesn't carry. That's the post's own thesis one level down — I stripped the conditions (the denominator, the n) that make a number a measurement, and kept the number.
Point by point, because each fix is different:
Min-fold gets an interval and an n, always —
x% [lo–hi], Nper cell. The interval isn't decoration here; it's the thing that stops a min-fold (smallest numerator, widest interval by construction) from being read as a point.The ranking below the top isn't real, so it goes. This is the one that stings and it's the most important. If 6 of 8 adjacent pairs overlap at 95%, the leaderboard is asserting an order the data can't resolve — and a rank is itself "green by construction": it looks like information and encodes none below Prompt Guard 2 86m. So the honest output is a partial order — one detector cleanly separated at the top, and an unranked pack of "indistinguishable at this n" (four literally 0% [0–3%]). Intervals, not positions, below first place.
The unreadability was unit-collision, not count. You're right — four numbers is fine; three different kinds of rate (budget / false-alarm / TPR) sitting unlabelled side by side is the cost. Label the denominators and the detail pays for itself: 4-cell row under the headline, full table in
results/*.json.The 2% budget is an integer wearing a percentage.
floor(0.02 × {57,77,81,76}) = 1in every fold — so "2%" is "exactly one benign over the line," and printing it as a rate invents precision the calibration split doesn't have. It reports as an allowed count until the benign corpus clears the size where the percentage means anything.The question that follows is the one I actually don't know: below the top, is the pack resolvable at any feasible n, or is "these are indistinguishable" the real scientific finding? At the min-fold widths you listed, separating two detectors ~10 points apart needs the attack corpus to grow a lot — and if ranking the pack is just noise-mining, the benchmark's honest output is "one detector clears the bar, the rest are tied," full stop. Did your re-run give you any feel for whether more attacks would ever break the tie, or is the pack genuinely level?
Keeping the false-positive column in, and the honesty clause about the shared AgentDojo template, is what makes the 1% → 99% result believable.
Two things I'd add to the checklist.
The false-alarm budget needs enough benign samples to back it. If "wrongly flags at most 2%" is checked on something like the 97 benign outputs, the data can't really say 2%: with 0 flags in 97 the 95% upper bound on the false-alarm rate is about 3.7%, and with 1 flag it's about 5.6%. To show "at most 2%" with zero flags you need roughly 150 benign samples, and more if a few flags are allowed. At a threshold like 0.003, sitting that close to the benign scores, this is the number most likely to surprise someone in production.
On check 4, the known attack fired through the live guardrail: if it's always the same one, it can keep passing after the detector has drifted, for the same reason you flag with the 99%. It may be recognising that one string. Drawing the canary from a held-out pool of attacks that were never used to set the threshold, and rotating it, turns the armed check into a small ongoing measurement instead of a single fixed probe.
@rudratosh thanks, and a small mix-up: your replies to pm25coder and me seem to have swapped places. The at_budget.py run and the three points about the splice, the fold spread and min-fold are pm25coder's; the benign sample size and the rotating canary were mine. Worth crediting them in the changelog under the right name.
On your canary question: I'd do both, at two speeds. Resample one canary from the held-out pool on every run, as you lean towards, and track the catch rate as a running count with its interval, so drift shows up as a trend. Then on every deploy or threshold change, run the whole held-out pool, which gives a proper effectiveness number at the moment it's most likely to break. And retire any canary that ever gets used to tune anything, or the pool slowly turns back into the fixed probe.
Absolutely — and thanks for the correction on attribution. I’ll make sure the changelog credits the
at_budget.pyanalysis and the benign-sample/canary suggestions separately.I like the two-speed approach.
For normal runs, rotate a canary from the held-out pool and track the catch rate over time. Then, whenever the detector is deployed or the threshold changes, run the entire held-out pool to get the real effectiveness measurement.
The “retire any canary that gets used for tuning” rule is especially important. Otherwise the canary quietly stops being a test and becomes another training signal.
That gives me a much better definition of what the live check should be:
the canary detects drift; the held-out pool measures effectiveness.
That distinction wasn't explicit enough in my original post. Really useful comment.
You report the 1%-to-99% swing on Prompt Guard 2 came down to one threshold number, 0.5 to 0.003, on the same weights. The detail worth underlining is where 0.5 came from: it's the demo default that most classifier wrappers inherit, chosen on a score distribution nobody's real traffic matches. We've seen the same shape in production — scores sitting at 0.009 versus 0.0008 look safely separated until attack phrasing shifts and the gap closes, which is why a one-time sweep doesn't hold. The armed check that survives it is re-running the calibration on fresh traffic, not just the canary.
This is the sharper version of my point, and I'm going to steal the framing: a one-time sweep calibrates to a distribution that's already expiring.
You nailed why 0.5 is so dangerous — it's not a value anyone chose for your traffic, it's the demo default that rides in with the wrapper and never gets questioned because the light stays green.
The bit you added that I under-weighted: the 0.009 vs 0.0008 gap isn't a fixed property of the model, it's a property of the current attack phrasing. Shift the phrasing and the gap closes, and a threshold you swept last quarter is now sitting in the dead zone again — silently, with every health check still passing.
So the real control isn't "sweep once, set threshold, done." It's:
Out of curiosity — in your production case, what cadence did you land on for re-calibration, and was it time-based or triggered by the separation metric closing?
This connects directly to something we were working through on your provenance/taint thread: a single global operating point is the same "green by construction" failure mode, just one level up. If you pick one threshold calibrated to one false-alarm budget across all traffic, you're implicitly saying every request deserves the same tolerance for a miss — but a buried injection aimed at
search_weband one aimed atsend_paymentare not equally expensive to let through. The fix in both cases is the same shape: stop gating on one axis. Tier the operating point the way you'd tier the risk score — a stricter (lower) threshold, i.e. a smaller false-alarm budget you're willing to pay, for calls that reach destructive/high-trust tools, a looser one for calls that reach read-only ones. Otherwise you can be "well-calibrated" on your own benchmark and still be arming the wrong gate for the one call that actually matters, for the same reason your own post's checklist flags: a single calibration check doesn't tell you it's calibrated correctly for every downstream consequence, only for the traffic mix you averaged over.Agreed, and that's a fair hit on my own checklist. One threshold averaged over all traffic is calibrated for the mix, not for the call that matters. A miss on
search_weband a miss onsend_paymentare priced the same, which is obviously wrong.So the operating-point line should really read:
The practical cost is sample size. I only had 97 benign outputs for one global threshold. Split that across tiers and each tier's 2% budget is being measured with a handful of examples.
How do you get enough benign volume per tier to calibrate the strict ones, since those are usually the rarest calls?
"Green, and catching nothing" is the best description of the worst failure mode I work with, and it shows up long before AI: the guard's logic is correct, its dashboard is healthy, and the number it compares against froze two weeks ago.
I build a risk guard for MetaTrader 5, and the bug that cost me the most was exactly this shape. A snapshot that a different code path was supposed to refresh quietly stopped refreshing. No exception, no error log, no red metric - every number in the system agreed with every other number, and they were all wrong together. A guard reading a plausible frozen value behaves identically to a guard that works, right up until the day you need it.
Three things I now treat as non-negotiable, which I think transfer directly to guardrails:
Never let "no data" collapse into "zero" or "clear". The moment a failed read is represented as a benign value, the guard has no way to distinguish "nothing happened" from "I cannot see anything" - and only one of those means you're safe. If a measurement cannot be taken, that is an incident, not a pass.
Assert on the effective value, from the same object production reads. Tests that assert against a freshly computed copy prove the formula, not the pipeline. The frozen-input bug is invisible to any test that doesn't touch the real input path.
Test by destroying the input, not by unit-testing the logic. Cut the data source, point it at stale data, and watch what the guard does. If it keeps reporting green, you've found the same hole - and no amount of passing logic tests would have found it.
One more that's specific to anything doing time-based checks: measure both time-to-detect and time-to-effect. A guard that detects in 200ms and acts in 15 minutes protects no one, while looking perfectly healthy the whole time.
(I maintain a free, open-source MT5 risk guard - disclosure, so my bias is visible: xuks124.github.io/vigildesk/free.html - and wrote up the frozen-value pattern here: xuks124.github.io/vigildesk/blog/s...)
Great title by the way - it says more than most guardrail posts manage in the whole article.
Great post. The liveness vs. effectiveness distinction is the cleanest way I've seen to explain why this class of failure survives so long. A health check that returns "healthy" for a guardrail catching 1% and one catching 99% isn't measuring the thing you care about, and the bouncer-with-closed-eyes image makes that stick.
I also liked that you separated discrimination from operating point. Prompt Guard 2 ranking attacks roughly 10x above benign text while sitting entirely under the default 0.5 cutoff is a good reminder that AUC-style thinking and deployment thinking are different exercises. A model can be good at the first and useless at the second, and most dashboards only surface neither.
Your honesty clause on the 99% is the right call. With every AgentDojo attack sharing one wrapper template, a threshold tuned that finely could be keying on the template rather than on injection in general. The durable finding is the miscalibration, not the headline number, and I'm glad you said so before someone screenshotted it wrong.
On the "armed check," one practical addition: run a small canary suite through the live guardrail on a schedule, not just at deploy time, and alert when the block rate on known attacks drops below a floor. That catches regressions from model swaps, threshold edits, or a config that quietly reverts. It's also worth varying the canaries, since a fixed set can be memorized by a threshold tuned to it, which is the same template problem you flagged.
Two questions:
The closing list of "green by construction" examples, from log-only WAF rules to muted alert channels and skipped tests, is the part I'll be repeating to my team. Thanks for open-sourcing the benchmark.
Thanks, and the scheduled canary with a block-rate floor is the right upgrade to my "armed check". Deploy-time only catches the config you shipped, not the one that quietly reverted. Rotating the canaries matters for the reason you gave: a fixed set becomes one more template to overfit.
Your two questions, answered straight:
I can't claim stability. I tuned on 3 domains and measured on the held-out 4th, which is one split, not a full rotation with a per-domain threshold table. A threshold that small, set with 97 benign samples, should be assumed to drift. I read that as an argument for per-deployment calibration, same as you.
Not tested, and it's the biggest hole. Every AgentDojo attack shares one wrapper template, so the benchmark measures buried attacks, not adaptive ones. I'd expect paraphrase to hurt and splitting across outputs to hurt more, since each fragment is scored alone.
If you were adding one of those two to the benchmark first, which would you pick?