DEV Community

Rudratosh Shastri
Rudratosh Shastri

Posted on

Your AI guardrail is green. It's also catching nothing.

There are three ways a security guardrail can fail you.

Two of them are loud. One of them is a serial killer.

  1. It's down. Crashed, misconfigured, not deployed. You find out fast — errors, alerts, a red dashboard, an on-call page at 3am. Painful, but honest.
  2. It's weak. It runs, it catches some attacks, it misses others. You can measure it, argue about it, improve it. Also honest.
  3. It's running perfectly, passing every health check, returning valid scores on every request — and configured to catch nothing. No error. No alert. No red anything. Every dashboard it touches is green by construction.

The third one is the problem. Because it has no symptom. Your monitoring says "guardrail: healthy ✅" and it's telling the truth. The guardrail is healthy. It's just not a guardrail.

I ran into a textbook case of #3 while benchmarking prompt-injection detectors, and it's worth dissecting, because once you see the shape of it you'll start finding it everywhere.

The setup: the smoke alarm of the AI stack

Everyone shipping an AI agent bolts on a prompt-injection detector — a little classifier that reads the text flowing into the model and screams if it smells an attack. It's the smoke alarm of the agent world. You install it, you see the green light, you move on.

So I did the obvious experiment: I took 10 of these smoke alarms and set 629 real fires.

Specifically — I took 629 real injection attacks from AgentDojo, buried each one inside ordinary tool output (a bill, an email, a web page — the way an agent firewall actually sees them, not the clean lab version), and ran 10 open-source detectors over the lot, plus 97 benign tool outputs to catch the ones that just alarm at everything.

Most of the results were the boring kind of bad (the "weak" and "screams at toast" failures). But one detector produced the interesting kind of bad.

Exhibit A: Meta's Prompt Guard 2, catching 1%

Meta's Prompt Guard 2 — the model everyone name-drops — caught 6 of 629 buried attacks. About 1%.

Now, your first instinct is "the model is bad." Reasonable instinct. Wrong.

Here's what the model was actually doing on each request:

Attack Benign
Prompt Guard 2 score ~0.009 ~0.0008
The 0.5 decision cutoff 0.5 0.5
Verdict ✅ ALLOW ✅ ALLOW

Look at the top row. The attack scores roughly 10x higher than normal text — the model's ranking is nearly perfect. As a discriminator, it works.

Now look at the cutoff. Use the model the obvious way — the standard 0.5 decision boundary you'd apply to any binary classifier out of the box (block if P(malicious) ≥ 0.5) — and every score it produces, the attack at 0.009 and grandma at 0.0008 alike, sits far below the line. So the verdict on every single request is identical: "nah, we're good." ✅

(To be precise, because this is the sentence people will poke at: 0.5 isn't some evil value Meta hard-coded — it's the threshold you get by treating a probability classifier the normal way. The failure isn't a bad config someone shipped; it's that **the obvious, default way to use this model is off by ~50x for buried attacks* — and nothing about the running system would ever tell you.)*

The detector was producing different scores. The control just wasn't doing anything with the difference.

It's a smoke alarm with the sensitivity dial turned all the way to "only trigger for a literal supernova." The sensor is fine. The wiring is fine. It will faithfully detect the heat death of the universe. Your kitchen fire? Not so much.

That is not a weak guardrail. That is a perfectly functional discriminator thresholded into a no-op. It loads, it returns scores, it passes health checks, it logs clean — and it catches 1% of attacks. Green by construction.

The plot twist that makes it worse

Here's the part that should genuinely unsettle you.

I re-ran Prompt Guard 2 with the threshold moved from 0.5 down to 0.003 — into the range where its scores actually live — and calibrated so it wrongly flags at most 2% of normal traffic, measured on an AgentDojo domain it was never tuned on.

It went from catching 1% to catching 99% (621/629).

Same model. Same weights. Same requests. One number. The difference between "protects you from basically nothing" and "catches almost everything" was a config value the model shipped with — off by roughly two orders of magnitude for this use case.

The uncomfortable takeaway isn't "the model is good" or "the model is bad." It's that the shipped default was wrong by ~100x, and nothing about the running system would ever tell you. Every health check passes at 1% exactly as it does at 99%.

(Honesty clause, because this is the part everyone screenshots wrong: don't read "99%" as "Prompt Guard 2 solves prompt injection." Every AgentDojo attack uses the same wrapper template, so a threshold tuned that finely may be recognizing the template, not attacks in general. The load-bearing claim is the miscalibration — "the default is wrong by 100x" — not the 99%. Tune on **your* traffic, then trust a number.)*

Why "green by construction" is the dangerous class

Compare the three failure modes by how you'd find out:

Failure Dashboard How you learn about it
Guardrail down 🔴 red Instantly — errors, alerts, paging
Guardrail weak 🟡 measurable Eventually — an incident, a red-team, a metric
Guardrail green & useless 🟢 green Never — until the breach, and even then you'll swear it was on

The first two fail toward visibility. The third fails toward silence. It's the security equivalent of a bouncer who's clocked in, in uniform, standing at the door, checking every ID — with his eyes closed. Every metric you have says "door: staffed." The metric you don't have says "staffed ≠ guarding."

And "it passed the health check" is doing a lot of unearned work in most teams' heads. This is the gap between a liveness check (is the thing running?) and an effectiveness check (is it actually stopping what it's supposed to stop?). A health check proves liveness. It says nothing about effectiveness — and almost every dashboard measures the first while quietly implying it measured the second.

How to actually tell if yours is armed

This isn't a "shame on Meta" post — the model does what it was trained to do, and defaults have to be conservative to avoid blocking everyone. It's a "shame on us if we trust the green light" post. Concretely:

  1. A score that ranks well with a bad cutoff is worth zero. Separately measure discrimination (does it rank attacks above benign?) and operating point (is the threshold where the scores actually live?). A model can be great at the first and useless at the second — which is exactly this case.
  2. Never trust a vendor's default threshold on your data. It was picked for a distribution that isn't yours. Sweep it. Find where your attack scores and your benign scores actually sit.
  3. Measure at a false-alarm budget, not in the abstract. "Catches 99%" is meaningless without "…while wrongly blocking X% of normal traffic." A detector that blocks 98% of legit calls isn't a control; it's a way to make everyone route around your agent. (Two detectors in my benchmark did exactly this and still "scored" 100% on attacks. A brick over the deny button scores 100% too.)
  4. Health check ≠ armed check. Add a test that fires a known attack through the live guardrail and asserts it's blocked. If your monitoring can't tell the difference between "catching 99%" and "catching 1%," your monitoring is green by construction too.
  5. Red-team the config, not just the model. The vulnerability here wasn't in the weights. It was in a single threshold. Your attack surface includes your YAML.

Or, as a checklist you can staple to any guardrail:

Liveness check:     Is the detector running at all?
Effectiveness check: Does it BLOCK a known attack, right now, in prod?
Calibration check:   Are attack scores actually separated from benign scores?
Operating-point check: At our tolerable false-alarm rate, what % of attacks do we catch?
Regression check:    Is all of the above still true after the next model/config change?
Enter fullscreen mode Exit fullscreen mode

Only the first line is on most dashboards. The other four are the difference between a control and a decoration.

The bigger, more uncomfortable question

Prompt Guard 2 at 1% is a clean, measurable example because I had 629 attacks to throw at it. But the shape generalizes way past prompt injection:

  • The WAF rule set that's deployed but in "log-only" mode.
  • The alert that's been firing into a muted channel since Q1.
  • The MFA that's enforced… except for the legacy API path.
  • The test suite that's green because half of it is skip.

All green. All healthy. All protecting you from a supernova and nothing smaller.

So the question I'd actually sit with: how many of your green dashboards are green because the thing works — and how many are green by construction? The scary answer is that, by definition, you can't tell the difference by looking at the dashboard. You have to fire a real attack at it and watch.

The whole benchmark — 10 detectors, 629 attacks, the threshold sweep, the cross-domain calibration — is open and reproducible here, if you want to see which alarms are real and which are decorative:

github.com/rudratoshs/buried-injections ⭐

The full thing — 10 detectors, 629 attacks, the threshold sweep, the cross-domain calibration — is open and reproducible. If you find a flaw in the methodology, I genuinely want to know — security benchmarks are more useful when people try to break them. And if it made you want to go check whether your own guardrail is armed or just green, a ⭐ helps it reach the next person about to trust a green light. 🙏


Be honest in the comments: have you ever found a security control that was "on," passing every check, and doing absolutely nothing? What was it, and how did you finally catch it — an incident, a red-team, or dumb luck? 👇

I write about AI, security, and the honest ways things break — benchmarks with the false-positive column left in. Follow me here if that's your lane. 👋

Top comments (26)

Collapse
 
naveen_alavilli profile image
Naveen Alavilli •

I'd extend the armed check all the way to the protected action. A detector can return BLOCK while an adapter logs the decision and still invokes the tool. A synthetic request against an isolated test resource should assert both that the policy rejected it and that no downstream write occurred; pair it with a benign request so a broken path that rejects everything cannot pass.

The shared-wrapper caveat also makes a useful next experiment: hold out wrapper families as well as domains, and report the false-positive count and denominator per split. With 97 benign examples overall, a couple of classifications materially move the rate. That would help separate threshold calibration from recognizing a recurring packaging pattern.

Collapse
 
rudratosh profile image
Rudratosh Shastri •

assert both that the policy rejected it and that no downstream write occurred

Yes — this is the version of the armed check that can't lie. A detector returning BLOCK proves nothing if the adapter still invokes the tool; the only honest assertion is at the protected resource: a synthetic request against an isolated test resource must show (a) the policy rejected it and (b) zero downstream write. Pair it with a benign request so a path that rejects everything can't pass. That's liveness and effectiveness in one probe.

And holding out wrapper families as well as domains is exactly the experiment I owe the post. With 97 benign examples a couple of flips move the rate materially, so per split I'd report the raw false-positive count and denominator, not a percentage — a rate over a small integer implies a precision it doesn't have. That's what separates "calibrated the threshold" from "memorized the packaging." Adding this. Thanks for sharpening it.

Collapse
 
naveen_alavilli profile image
Naveen Alavilli •

Thanks for taking this into the follow-up. One detail worth making explicit in the harness: when is the zero-write assertion evaluated? If an adapter queues work, an immediate read can pass before a delayed write arrives. I'd correlate the probe with a unique operation ID and drive the test queue to completion before checking the isolated resource. In a live probe, report the observation window rather than claiming the write can never happen.

The benign control should traverse that same adapter and queue. That makes a missing worker distinguishable from successful enforcement.

Collapse
 
iamnaomi profile image
N A O M I •

The distinction between liveness and effectiveness is the sharpest framing here. A health check proving the guardrail is running says nothing about whether it's stopping anything, and the Prompt Guard 2 numbers make that concrete: the attack scoring roughly 10x higher than benign text, both sitting far below 0.5, so every verdict reads "allow" while the ranking is nearly perfect. Moving the threshold to 0.003 and going from 1% to 99% shows how much one config value can hide.

I also appreciated the honesty clauses. Noting that the 99% may reflect AgentDojo's shared wrapper template rather than general detection, and that the real claim is the miscalibration, keeps the post from becoming a screenshot-able overclaim. Insisting on measuring at a false-alarm budget, with the "brick over the deny button scores 100%" point, was a useful reminder too.

To answer your closing question, the closest I've seen is an alert routed to a muted channel for months. Everything looked wired up, and nothing was actually being heard. Firing a known attack through the live system on a schedule seems like the only reliable fix. Thanks for open-sourcing the benchmark.

Collapse
 
rudratosh profile image
Rudratosh Shastri •

an alert routed to a muted channel for months. Everything looked wired up, and nothing was actually being heard.

That's the purest version of it. The detection worked, the routing worked, and the last step, a human hearing it, was silently off. Every component would pass its own health check.

It also shows why the fix has to be end to end:

  • Component check: did the alert fire? Yes.
  • Delivery check: did it reach the channel? Yes.
  • Outcome check: did anyone react? Nothing was measuring that.

A scheduled known-bad event only counts if the assertion sits at the far end, with someone or something acknowledging it.

How did the muted channel finally get noticed: an incident, or someone cleaning up their notification settings?

Collapse
 
rudratosh profile image
Rudratosh Shastri •

@kartik-nvjk > How are you calibrating thresholds against real attack traffic rather than the default 0.5?

Honest answer first: I don't have production attack traffic. The benchmark calibrates on AgentDojo, tuning on 3 domains and measuring on the 4th. So what I can give you is the procedure, not a number to copy:

  1. Start from benign traffic, not attacks. Score a sample of your own real tool outputs and look at where the scores actually sit.
  2. Pick the false-alarm budget first (I used 2%), then set the threshold at that percentile of the benign scores.
  3. Measure the catch rate at that threshold on attacks buried in your kind of tool output, on data the threshold never saw.
  4. Re-run it on every model or config change, because the score distribution moves.

The weak point is step 3: known attacks are a proxy for real ones, and 97 benign samples is a small stick for measuring a 2% budget.

"Graded green while checking a quarter of it" is the same failure one layer up. Did Bedrock's eval tell you which three quarters it skipped, or did you have to find that out yourself?

Collapse
 
pm25coder profile image
pm25coder •

@kartik-nvjk Short version: the threshold is a benign-only quantity — provable, and it changes what "calibrate against real attack traffic" should mean.

For a false-alarm budget f, the max-TPR threshold is the k-th highest benign score (k = floor(f·n)). Attack labels cannot move it: the FPR is a function of the benign scores alone, and TPR is monotone in the threshold. So the labelled attacks choose the budget; benign traffic sets the threshold. I swept an attack-labelled threshold over all nine detectors in this benchmark's own bench/results/at_budget_2pct.json — at each detector's benign-derived FPR, the best attack-labelled TPR is the same number to the decimal (98.7 / 33.2 / 14.1 / 47.5 / 48.8 / 6.2 / 0.0 … , 9/9).

Your ~50x reproduces exactly: at t=0.5 the same file gives prompt-guard-2-86m 6/629 = 1.0% — the headline number — with a false-alarm rate of 0.0%. 0.5 sits ~53x above its attack median and ~90x above its highest benign score.

And the reason a single default can't be shipped: the same 0.5 fails in both directions in one benchmark. deepset-deberta and fmops-distilbert score benign traffic at ~0.999, so 0.5 lands below their entire range — TPR 100% / FPR 97.9%. The constant that makes Prompt Guard catch nothing makes those two scream at toast.

The part that bites once you do calibrate on real traffic: it doesn't transfer, because it's the benign distribution that isn't stationary across sources. Calibrating on three folds and applying the threshold to the fourth breaks the promised 2% budget on 11 of 36 folds (held-out FPR 4.9%, 2.5x the budget) — and the offenders are exactly the folds whose benign scores sit an order of magnitude higher. prompt-guard-2-22m: benign p50 on travel is 0.0092 vs ~0.0025 elsewhere, so the transferred threshold flags 13/20 benign. prompt-guard-2-86m on slack: 5/21 flagged (24%), because slack's highest benign score is 4x travel's.

So the shape I'd take away: derive the threshold per traffic source (or per rolling window), not per model, and watch the benign score distribution — its p50/p99 is the leading indicator; the attack side only tells you the payoff, after the fact. It's also the cheap version of your Bedrock point: an eval that reports on a quarter of the agent has the same defect as a threshold nobody asked what its traffic looks like.

Collapse
 
rudratosh profile image
Rudratosh Shastri •

~1% is indistinguishable from off

That's the whole post in six words — stealing it.

On calibration, the counterintuitive part (which @pm25coder just proved above): you don't calibrate against attack traffic at all. The threshold is a benign-only quantity — for a false-alarm budget f, it's the k-th highest benign score (k = floor(f·n)). Attacks only choose the budget you're willing to spend; your benign traffic sets the cutoff. So "calibrate against real attack traffic" is the wrong frame — calibrate against real benign traffic, per source, and let a known-attack canary tell you the payoff after the fact.

Your Bedrock example is the same disease one layer up: an eval that grades green while covering a quarter of the agent is a threshold nobody asked what its traffic looks like.

Are you deriving one threshold globally, or per source? The benign distribution is the part that drifts — same detector, different source, different cutoff.

Collapse
 
pm25coder profile image
pm25coder •

Ran bench/at_budget.py over your own saved scores (bench/results/at_budget_2pct.json) before writing this, because the port is what makes the rest checkable — all nine detectors reproduce both saved columns exactly (in-sample and cross-domain) once the suite rotation is ordered workspace/travel/banking/slack, so the numbers below are yours, not a re-derivation.

Three things the pooled figure doesn't say.

1. The 2% budget is an in-sample property; the unseen column is a different quantity. caught_at_budget enforces the budget when picking (allowed = floor(budget * len(calib)), per fold on the calibration split), then measures on the held-out suite, where it can exceed it — and does, for eight of the nine (the regex baseline is the exception): 5.2% for Prompt Guard 2 86m, 13.4% for the 22m, 9.3% for testsavant. So "99% at a 2% false-alarm budget" splices an in-sample budget onto a cross-domain TPR. Both are honest numbers; the sentence that gets screenshotted has combined two.

2. The pooled TPR hides a fold spread as large as the rate it reports. Per held-out suite, prompt-guard-2-22m is 100% on travel, 22.9% on workspace, 16.0% on banking, 1.0% on slack — pooled 35%, spread 99 points. jailbreak-detector-large runs 17.1% → 100% (pooled 51%). For six of the nine the spread is at least the pooled rate itself. cross_domain already has each fold in c and accumulates it away (caught += c); keeping the list and printing min/max is two lines, and it changes what "the number to trust" means.

3. For a control, the worst fold is the number to trust — not the mean. That's your own thesis one level down: a pooled average is a liveness statistic (it says the detector ran on four suites), while the min fold is closer to an effectiveness one. Prompt Guard 2 86m is the reassuring case — 97–100%, a 3.3-point spread — and most of the field isn't.

The benign-sample size and the rotating-canary points above are the other half of it; both are cheap.

Collapse
 
rudratosh profile image
Rudratosh Shastri •

This is exactly the kind of comment I was hoping this post would attract.

You’re right on the 2% point — I compressed two different quantities into one sentence. The budget is enforced during calibration, while the reported held-out/cross-domain catch rate is a different measurement. Putting those side by side without making that distinction explicit makes the headline number easier to screenshot than to interpret.

The fold spread point is even more important. A pooled TPR can look reassuring while one domain is basically a miss. I like the idea of keeping the per-fold values visible and reporting the worst fold alongside the pooled number.

That actually fits the thesis of the article better: the number isn't the measurement unless you preserve the conditions under which it was produced.

I'm going to fix this rather than defend the original wording. Thanks for actually running the code and checking it against the saved results.

Collapse
 
rudratosh profile image
Rudratosh Shastri •

This is a really good catch.

I was treating “0 false positives out of 97” as if it established a 2% false-alarm rate, when statistically it really only tells us what happened in that small sample.

Your point about the benign denominator is exactly the kind of thing that gets lost when a benchmark focuses on the attack-success column. At a threshold this close to the benign-score distribution, the uncertainty around the benign sample can matter more than another decimal place on TPR.

And I really like the held-out rotating-canary idea. A fixed canary can gradually become a memorization test instead of a security test.

So the revised measurement should probably report both:

effectiveness on the held-out pool + confidence around the benign false-alarm estimate.

That’s a much more useful production story than simply saying “99% caught.”

Great addition.

Thread Thread
 
pm25coder profile image
pm25coder •

On the min-fold question — same re-run, and the answer changed twice while I was measuring it.

Headline pair: yes. But attach n and an interval to the min fold. Min-fold is by construction the smallest numerator, so it is also the widest interval. Your min folds, Wilson 95%, ordered: 97% [94–98], 26% [19–34], 17% [12–24], 1% [0–5], then five of the nine bottoming out at 0/140–240 → [0–3%]. A bare "17%" reads like a measurement; "17% [12–24], 140 attacks" reads like one.

The surprise: the min-fold leaderboard is mostly a tie. Walking the ranked min folds, 6 of the 8 adjacent pairs have overlapping 95% intervals — Prompt Guard 2 86m is cleanly separated at the top, and below third place the ordering is not resolvable at these n (four detectors are all 0% [0–3%]). So print the interval with the rank, or replace the rank with the interval below the top — a rank implies a separation the data doesn't have.

On readability: the per-suite row is four numbers — that is not what makes it unreadable. Put the 4-cell row under the headline and the full table in results/*.json; what actually costs a reader is three quantities with different denominators (budget / false-alarm rate / TPR) sitting next to each other unlabelled. Label the denominators and you can afford the detail.

One more, same re-run, and it cuts the same way: at this corpus size the 2% budget is degenerate. allowed = floor(0.02 * len(calib)), and calib is 57/77/81/76 benign cases depending on which suite is held out → 1 in all four folds. So cross-domain, "2%" is not a rate at all; it is "exactly one benign case above the line". Reporting the budget as an allowed count until the benign corpus grows (your ~150+ fix) makes that visible from the header — and it is a third reason those two numbers shouldn't share a sentence.

Thread Thread
 
rudratosh profile image
Rudratosh Shastri •

a bare "17%" reads like a measurement; "17% [12–24], 140 attacks" reads like one

Taking all four together, they collapse into one correction I should have made myself: every headline in that benchmark is a percentage printed over a small integer, and a percentage over a small integer implies a resolution the count doesn't carry. That's the post's own thesis one level down — I stripped the conditions (the denominator, the n) that make a number a measurement, and kept the number.

Point by point, because each fix is different:

Min-fold gets an interval and an n, always — x% [lo–hi], N per cell. The interval isn't decoration here; it's the thing that stops a min-fold (smallest numerator, widest interval by construction) from being read as a point.

The ranking below the top isn't real, so it goes. This is the one that stings and it's the most important. If 6 of 8 adjacent pairs overlap at 95%, the leaderboard is asserting an order the data can't resolve — and a rank is itself "green by construction": it looks like information and encodes none below Prompt Guard 2 86m. So the honest output is a partial order — one detector cleanly separated at the top, and an unranked pack of "indistinguishable at this n" (four literally 0% [0–3%]). Intervals, not positions, below first place.

The unreadability was unit-collision, not count. You're right — four numbers is fine; three different kinds of rate (budget / false-alarm / TPR) sitting unlabelled side by side is the cost. Label the denominators and the detail pays for itself: 4-cell row under the headline, full table in results/*.json.

The 2% budget is an integer wearing a percentage. floor(0.02 × {57,77,81,76}) = 1 in every fold — so "2%" is "exactly one benign over the line," and printing it as a rate invents precision the calibration split doesn't have. It reports as an allowed count until the benign corpus clears the size where the percentage means anything.

The question that follows is the one I actually don't know: below the top, is the pack resolvable at any feasible n, or is "these are indistinguishable" the real scientific finding? At the min-fold widths you listed, separating two detectors ~10 points apart needs the attack corpus to grow a lot — and if ranking the pack is just noise-mining, the benchmark's honest output is "one detector clears the bar, the rest are tied," full stop. Did your re-run give you any feel for whether more attacks would ever break the tie, or is the pack genuinely level?

Collapse
 
arhancanli profile image
Arhan Canli •

Keeping the false-positive column in, and the honesty clause about the shared AgentDojo template, is what makes the 1% → 99% result believable.

Two things I'd add to the checklist.

The false-alarm budget needs enough benign samples to back it. If "wrongly flags at most 2%" is checked on something like the 97 benign outputs, the data can't really say 2%: with 0 flags in 97 the 95% upper bound on the false-alarm rate is about 3.7%, and with 1 flag it's about 5.6%. To show "at most 2%" with zero flags you need roughly 150 benign samples, and more if a few flags are allowed. At a threshold like 0.003, sitting that close to the benign scores, this is the number most likely to surprise someone in production.

On check 4, the known attack fired through the live guardrail: if it's always the same one, it can keep passing after the detector has drifted, for the same reason you flag with the 99%. It may be recognising that one string. Drawing the canary from a held-out pool of attacks that were never used to set the threshold, and rotating it, turns the armed check into a small ongoing measurement instead of a single fixed probe.

Collapse
 
arhancanli profile image
Arhan Canli •

@rudratosh thanks, and a small mix-up: your replies to pm25coder and me seem to have swapped places. The at_budget.py run and the three points about the splice, the fold spread and min-fold are pm25coder's; the benign sample size and the rotating canary were mine. Worth crediting them in the changelog under the right name.

On your canary question: I'd do both, at two speeds. Resample one canary from the held-out pool on every run, as you lean towards, and track the catch rate as a running count with its interval, so drift shows up as a trend. Then on every deploy or threshold change, run the whole held-out pool, which gives a proper effectiveness number at the moment it's most likely to break. And retire any canary that ever gets used to tune anything, or the pool slowly turns back into the fixed probe.

Collapse
 
rudratosh profile image
Rudratosh Shastri •

Absolutely — and thanks for the correction on attribution. I’ll make sure the changelog credits the at_budget.py analysis and the benign-sample/canary suggestions separately.

I like the two-speed approach.

For normal runs, rotate a canary from the held-out pool and track the catch rate over time. Then, whenever the detector is deployed or the threshold changes, run the entire held-out pool to get the real effectiveness measurement.

The “retire any canary that gets used for tuning” rule is especially important. Otherwise the canary quietly stops being a test and becomes another training signal.

That gives me a much better definition of what the live check should be:

the canary detects drift; the held-out pool measures effectiveness.

That distinction wasn't explicit enough in my original post. Really useful comment.

Collapse
 
aifrontierpost profile image
AI Frontier Post •

You report the 1%-to-99% swing on Prompt Guard 2 came down to one threshold number, 0.5 to 0.003, on the same weights. The detail worth underlining is where 0.5 came from: it's the demo default that most classifier wrappers inherit, chosen on a score distribution nobody's real traffic matches. We've seen the same shape in production — scores sitting at 0.009 versus 0.0008 look safely separated until attack phrasing shifts and the gap closes, which is why a one-time sweep doesn't hold. The armed check that survives it is re-running the calibration on fresh traffic, not just the canary.

Collapse
 
rudratosh profile image
Rudratosh Shastri •

The armed check that survives it is re-running the calibration on fresh traffic, not just the canary.

This is the sharper version of my point, and I'm going to steal the framing: a one-time sweep calibrates to a distribution that's already expiring.

You nailed why 0.5 is so dangerous — it's not a value anyone chose for your traffic, it's the demo default that rides in with the wrapper and never gets questioned because the light stays green.

The bit you added that I under-weighted: the 0.009 vs 0.0008 gap isn't a fixed property of the model, it's a property of the current attack phrasing. Shift the phrasing and the gap closes, and a threshold you swept last quarter is now sitting in the dead zone again — silently, with every health check still passing.

So the real control isn't "sweep once, set threshold, done." It's:

  • Calibration as a recurring job, re-run on fresh traffic, not a one-time setup step
  • A live known-attack canary that asserts blocked, so drift trips an alarm instead of hiding
  • Watching the score gap itself as a metric — if attack/benign separation shrinks, that's your early warning before catch-rate craters

Out of curiosity — in your production case, what cadence did you land on for re-calibration, and was it time-based or triggered by the separation metric closing?

Collapse
 
hamid_ahmadian_3570449f72 profile image
Hamid Ahmadian •

This connects directly to something we were working through on your provenance/taint thread: a single global operating point is the same "green by construction" failure mode, just one level up. If you pick one threshold calibrated to one false-alarm budget across all traffic, you're implicitly saying every request deserves the same tolerance for a miss — but a buried injection aimed at search_web and one aimed at send_payment are not equally expensive to let through. The fix in both cases is the same shape: stop gating on one axis. Tier the operating point the way you'd tier the risk score — a stricter (lower) threshold, i.e. a smaller false-alarm budget you're willing to pay, for calls that reach destructive/high-trust tools, a looser one for calls that reach read-only ones. Otherwise you can be "well-calibrated" on your own benchmark and still be arming the wrong gate for the one call that actually matters, for the same reason your own post's checklist flags: a single calibration check doesn't tell you it's calibrated correctly for every downstream consequence, only for the traffic mix you averaged over.

Collapse
 
rudratosh profile image
Rudratosh Shastri •

a single global operating point is the same "green by construction" failure mode, just one level up

Agreed, and that's a fair hit on my own checklist. One threshold averaged over all traffic is calibrated for the mix, not for the call that matters. A miss on search_web and a miss on send_payment are priced the same, which is obviously wrong.

So the operating-point line should really read:

  • Per tier, not global: a tighter gate for calls that reach destructive or high-trust tools, a looser one for read-only calls
  • Reported per tier: catch rate and false-alarm rate for each, not one blended number

The practical cost is sample size. I only had 97 benign outputs for one global threshold. Split that across tiers and each tier's 2% budget is being measured with a handful of examples.

How do you get enough benign volume per tier to calibrate the strict ones, since those are usually the rarest calls?

Collapse
 
xuks124 profile image
xuks124 •

"Green, and catching nothing" is the best description of the worst failure mode I work with, and it shows up long before AI: the guard's logic is correct, its dashboard is healthy, and the number it compares against froze two weeks ago.

I build a risk guard for MetaTrader 5, and the bug that cost me the most was exactly this shape. A snapshot that a different code path was supposed to refresh quietly stopped refreshing. No exception, no error log, no red metric - every number in the system agreed with every other number, and they were all wrong together. A guard reading a plausible frozen value behaves identically to a guard that works, right up until the day you need it.

Three things I now treat as non-negotiable, which I think transfer directly to guardrails:

  1. Never let "no data" collapse into "zero" or "clear". The moment a failed read is represented as a benign value, the guard has no way to distinguish "nothing happened" from "I cannot see anything" - and only one of those means you're safe. If a measurement cannot be taken, that is an incident, not a pass.

  2. Assert on the effective value, from the same object production reads. Tests that assert against a freshly computed copy prove the formula, not the pipeline. The frozen-input bug is invisible to any test that doesn't touch the real input path.

  3. Test by destroying the input, not by unit-testing the logic. Cut the data source, point it at stale data, and watch what the guard does. If it keeps reporting green, you've found the same hole - and no amount of passing logic tests would have found it.

One more that's specific to anything doing time-based checks: measure both time-to-detect and time-to-effect. A guard that detects in 200ms and acts in 15 minutes protects no one, while looking perfectly healthy the whole time.

(I maintain a free, open-source MT5 risk guard - disclosure, so my bias is visible: xuks124.github.io/vigildesk/free.html - and wrote up the frozen-value pattern here: xuks124.github.io/vigildesk/blog/s...)

Great title by the way - it says more than most guardrail posts manage in the whole article.

Collapse
 
normalnorma profile image
Norma •

Great post. The liveness vs. effectiveness distinction is the cleanest way I've seen to explain why this class of failure survives so long. A health check that returns "healthy" for a guardrail catching 1% and one catching 99% isn't measuring the thing you care about, and the bouncer-with-closed-eyes image makes that stick.

I also liked that you separated discrimination from operating point. Prompt Guard 2 ranking attacks roughly 10x above benign text while sitting entirely under the default 0.5 cutoff is a good reminder that AUC-style thinking and deployment thinking are different exercises. A model can be good at the first and useless at the second, and most dashboards only surface neither.

Your honesty clause on the 99% is the right call. With every AgentDojo attack sharing one wrapper template, a threshold tuned that finely could be keying on the template rather than on injection in general. The durable finding is the miscalibration, not the headline number, and I'm glad you said so before someone screenshotted it wrong.

On the "armed check," one practical addition: run a small canary suite through the live guardrail on a schedule, not just at deploy time, and alert when the block rate on known attacks drops below a floor. That catches regressions from model swaps, threshold edits, or a config that quietly reverts. It's also worth varying the canaries, since a fixed set can be memorized by a threshold tuned to it, which is the same template problem you flagged.

Two questions:

  1. When you swept the threshold, how stable was the 0.003 operating point across the other AgentDojo domains? If it drifts a lot between domains, that argues for per-deployment calibration over any shipped default.
  2. Did you look at how the detectors behave when the injection is paraphrased or split across tool outputs? Buried-in-template attacks are one thing, but adaptive attackers will rewrite around whatever the detector keys on.

The closing list of "green by construction" examples, from log-only WAF rules to muted alert channels and skipped tests, is the part I'll be repeating to my team. Thanks for open-sourcing the benchmark.

Collapse
 
rudratosh profile image
Rudratosh Shastri •

Thanks, and the scheduled canary with a block-rate floor is the right upgrade to my "armed check". Deploy-time only catches the config you shipped, not the one that quietly reverted. Rotating the canaries matters for the reason you gave: a fixed set becomes one more template to overfit.

Your two questions, answered straight:

how stable was the 0.003 operating point across the other AgentDojo domains?

I can't claim stability. I tuned on 3 domains and measured on the held-out 4th, which is one split, not a full rotation with a per-domain threshold table. A threshold that small, set with 97 benign samples, should be assumed to drift. I read that as an argument for per-deployment calibration, same as you.

how the detectors behave when the injection is paraphrased or split across tool outputs?

Not tested, and it's the biggest hole. Every AgentDojo attack shares one wrapper template, so the benchmark measures buried attacks, not adaptive ones. I'd expect paraphrase to hurt and splitting across outputs to hurt more, since each fragment is scored alone.

If you were adding one of those two to the benchmark first, which would you pick?

Some comments may only be visible to logged-in visitors. Sign in to view all comments.