Everyone shipping an AI agent right now bolts on a "prompt-injection detector" — a little classifier that reads the text flowing through the agent ...
For further actions, you may consider blocking this person and/or reporting abuse
The flashcards-versus-open-book framing should worry anyone running a text classifier as their firewall: scoring an attack on its own says almost nothing about catching it buried in a Jira ticket. And the 0.003 threshold learning the AgentDojo wrapper rather than attacks in general is the scarier footnote. I land where you do, that provenance of the instruction plus what the tool call is allowed to do beats any detector. How do you track instruction origin once tool outputs nest a few calls deep?
The short version: don't ask the model to track it. The harness tracks it, and it's deliberately pessimistic.
The cost is over-tainting. Three calls deep, nearly everything is tainted, which is why it only works when paired with per-tool policy: a tainted read is fine, a tainted payment is not.
Are your nested chains mostly read-then-read, or do they end in a write?
The live harness that measures whether attacks succeed is where this belongs, because text-level catch rates decouple from actual blast radius. If an injection slips past Prompt Guard but the model treats it as body text and calls no sensitive tool, the vulnerability never materialized. If one slips through and touches the shell tool, a 9% catch rate means nothing.A threshold-tuner or CI classifier check still treats prompt injection as a text classification task. Running the test against the execution graph (recording whether tainted context induced an unauthorized tool call or mutated a downstream argument) gives you a deterministic pass/fail that doesn't drift when attackers change their prompt wrappers.
The reframe from "catch rate" to "catch rate at a fixed false-positive budget" is the right move — a detector that blocks 98% of legitimate traffic isn't a security control, it's a way to make everyone route around your agent. One thing I'd push on further: a single global threshold (2% FPR across the board) still treats every tool call as equally risky. In practice the false-positive budget you can tolerate on a read-only search_web call is very different from what you can tolerate on send_payment or delete_file — you'd happily eat a much higher FPR on the latter, because the cost of a missed attack vs. a blocked legitimate call isn't symmetric across tools the way a single global cutoff assumes. Have you tried per-tool-tier thresholds (tight budget for destructive tools, loose budget for read-only ones) instead of one number for everything? Also curious whether jailbreak-detector-large's 51%/2% held up on the held-out domain or regressed the way the others did once you cross-validated.
Fold in provenance too, not just destructiveness. Tiering purely by tool destructiveness still assumes the risk lives entirely in what the tool can do, but the attack surface you're measuring (prompt injection) is about where the instruction came from, not which tool eventually executes it. A
send_paymentcall built from a hardcoded, developer-authored amount is a different risk profile than the same call built from a value that was just scraped out of a webpage the agent fetched two turns ago — same tool, same destructiveness tier, very different trust in the input. If you only tier by tool, an attacker who can't get a destructive call approved directly will instead aim for a normally-loose read-only tool and use its output as the injection vector into a later, more sensitive call — the tight budget never triggers because the entry point looked "safe." Folding in provenance means the budget tightens the moment untrusted-origin data enters the chain, independent of which tool it eventually reaches, which closes exactly that laundering path. Practically: tag args by origin (user-typed vs. tool-output vs. web-fetched) and take the min of the tool-tier budget and the provenance-tier budget, so neither dimension alone can create a permissive combination. Great benchmark, by the way — the per-tool-tier addition you mentioned above is going straight onto my own reading list.This is the correction that turns the whole thing from a text problem into a data-flow problem, and it's the one I'd build on. Tiering by tool destructiveness alone still assumes the risk lives in what the tool can do — but prompt injection is defined by where the argument came from.
So the risk score is two-dimensional: destructiveness of the tool × provenance of its arguments. The false-positive budget lives in that grid, not on one axis:
send_payment, all args developer/user-authored → barely needs the detector at all.send_payment, any arg derived from fetched/untrusted content → strictest budget, or a hard gate.search_web, tainted args → who cares.And provenance has the property text scoring never will: it doesn't drift when attackers reword. A taint label is structural — "this value descended from untrusted content" — not a probability that moves when the wrapper changes. That's the deterministic pass/fail @reidmarlow was pointing at, one step earlier: not "did the model call a sensitive tool," but "did untrusted data reach a sensitive argument."
Which usefully collapses the detector's job: the text classifier stops being the gate and becomes a tiebreaker for the residual — the calls where tainted data legitimately must flow into a sensitive tool. "Pay the invoice amount I just read from this PDF" is the whole point of the agent, and it's tainted-data-into-
send_paymentby construction.So the question I'm stuck on, and I think you've thought about it: how do you keep the provenance gate from becoming a block-everything wall on exactly the useful flows — where tainted data is supposed to reach a sensitive tool? A second human-approval bound to that specific (value, source, destination)? Or is there a way to launder taint you'd actually trust?
Good question, and I don't think a single global rule survives contact with real traffic. The move that's worked for us is to stop treating "human approval" as the only release valve and instead let specific (source, destination) pairs earn a narrower, re-verifiable path:
So less "launder the taint" and more "make the destination able to independently confirm the tainted claim before it's allowed to act on it." Curious whether you've tried the cross-check-against-structured-extraction angle, or if your sensitive tools don't have an independent source to check against.
That's the version I'd build. Re-verification beats attestation because a human okay is itself something an attacker can farm, and it doesn't scale to the routine 95%.
To your question: no, I haven't built the structured-extraction cross-check into anything I've measured. One caveat I'd add before trusting it:
Binding per
invoice_idrather than clearing the whole flow is the part I'd copy straight away.For tools with no independent source at all, like a free-text email send, do you fall back to human approval every time, or is there a cheaper check?
The default is the finding, and it is worse than the headline number makes it sound. Prompt Guard 2 at 1% was running. It was loading, returning scores, passing health checks, and configured to catch nothing. So this is not the guardrail-is-not-running failure or the guardrail-is-weak failure. It is a third one: running, correct, and thresholded into a no-op. That version has no symptom at all, because every dashboard it touches is green by construction.
Where I would tighten the claim is the 99%. A threshold tuned on AgentDojo is a number about AgentDojo, and AgentDojo is a public corpus that detector authors can read. "The shipped default is wrong by two orders of magnitude" is fully supported by your run and is the important sentence. "One config change made it 99%" is supported against the set you tuned on. Those are different strengths of claim and the first one does not need the second.
The payments lens on the whole leaderboard: fraud screening has never been good enough to be the boundary, and card rails did not solve that by improving the classifier. The score routes, it does not decide. What makes a 1% catch rate on novel attacks survivable is that the effect is bounded and reversible, with a chargeback window and a liability shift behind it. So the number I would put next to catch rate and false-positive budget is what one miss costs and whether you can unwind it. An agent with a payment tool and no reversal path is the case where the detector has to be the boundary, and none of the ten in your table can be.
Your repo is the thing I want to see forked into a per-deployment harness rather than a leaderboard, for the reason in your own conclusion: the leaderboard tells me which alarm to buy and the harness tells me whether mine is armed.
Agreed — that's the detail that makes the approval a real gate instead of a token. If the signature covers (amount, currency, destination) but not the source that justified it, a re-plan can keep the money identical and swap the justification, and your old "yes" silently authorizes a new transaction. Binding the source id means the approval is scoped to why it was granted, not just what moves.
The sharp version: the approval token should commit to the full provenance of the decision, so any change to the inputs that produced it invalidates it. A yes is an answer to a specific question, not a standing permission.
Would you bind the source's id, or a hash of its actual content? An id survives the content being edited underneath it — which is its own small version of the label-stripping problem from the post: same reference, different meaning, and the old approval still matches.