DEV Community

Vishal Habib
Vishal Habib

Posted on Edited on AI-assisted

"0 of 18 got through" isn't a launch. Here's the number that is.

Every AI guardrail launch deck has a slide like this: "Blind test: 0 of 18 unauthorized actions got through." I've written that slide. Three times, for three of my own open-source checkers.

It passes the gate. It doesn't show what everyone in the room hears, which is "it doesn't let things through."

What "0 of N" actually tells you

If a checker lets through some fraction p of bad cases, the chance of seeing 0 misses in N tries is (1 − p)^N. Ask which values of p would make "0 of N" a not-too-surprising result (more than 5% likely), and you get the one-sided 95% upper bound:

p_max = 1 − 0.05^(1/N)

For my three checkers:

Checker Headline True miss rate could still be up to
agent-handoff-check 0 of 18 unauthorized calls through 15%
retirement-answer-check 0 of 25 wrong facts through 11%
listing-claim-check 0 of 11 high-harm claims through 24%, above its own 10% gate

None of these results were wrong. They just prove much less than they sound like they do.

The listing checker makes the point twice. After it passed, I had an agent attack it with the code open, and all 22 in-scope attacks got through. A clean blind run tells you the checker handles what a spec-reader imagines, not what an adversary writes.

The number that is a launch

Flip it around. To show the miss rate is under a target t with 95% confidence and no misses at all, you need:

N ≥ ln(0.05) / ln(1 − t)

  • Under 10%: 29 in a row
  • Under 5%: 59
  • Under 1%: 299

That's roughly the old "rule of three": with zero failures, the 95% bound is about 3/N. And it's the minimum. Every miss you do see pushes it up.

Two consequences I didn't appreciate until I ran the numbers:

  1. A zero-tolerance gate can't be passed by any sample. "0 unauthorized actions allowed" has to become a number, like "under 1% at 95% confidence", before evidence can ever meet it. Choosing that number is a product decision, not a stats one.
  2. A weekly point estimate isn't a rollback rule. One of my own PRDs said "roll back if recall drops below 95% over a week." A week with 19 of 20 caught passes that rule, and it's consistent with true recall as low as 78%.

What I changed

Each of the three READMEs now says, right under the headline result, how much it proves and how many more cases it would take. It makes the result look smaller. It's also the first question a careful reviewer would ask, so it's better answered up front.

I also tried to build a tool that turns a shadow-mode log into READY / NOT YET / INVALID. The math held up in every test I set in advance. The log handling didn't: two red teams got 16 of 25 and then 18 of 20 forged logs marked READY. No tool that just reads a log can beat someone who writes a fake one. That needs a tamper-evident log the checker writes itself. With no real shadow traffic yet to justify that, I parked it.

The slide I'd put up now

Blind test: 0 of 18 unauthorized actions got through. On its own, that bounds the miss rate at 15%. Our launch bar is under 1%, which needs 299 clean cases. Shadow mode collects them; we exit when the bound clears.

It's a less exciting slide. It's also one a risk reviewer will sign.


The three checkers, with every blind set and red-team run: agent-handoff-check · listing-claim-check · retirement-answer-check. Everything else I build: github.com/vishalhabib99.

Top comments (5)

Collapse
 
arhancanli profile image
Arhan Canli •

The slide at the end is the right one, and there's one more line worth writing down before shadow mode starts: what happens after the first miss. Even a genuinely good checker will usually see one before it reaches 299. At a true miss rate of 0.5%, only about 22% of 299-case runs come back clean; at 0.3%, about 41%.

The tempting move is to fix the miss and restart the count from zero, and that quietly breaks the bound, because you end up keeping the runs that finish clean and throwing away the ones that don't. Put the continuation in the plan instead. For the same under-1%-at-95% bar, one miss needs 473 cases, two need 628, three need 773 (exact binomial). If a miss makes you change the checker, that is a new checker and a new run, which is fine, as long as the failed run stays in the report.

Your listing-checker result fits the same frame: 0 of 11 blind and 22 of 22 adversarial aren't in conflict, they're bounds on two different populations. Each bound only covers the distribution its cases were drawn from, so the slide could say which one it's about.

Collapse
 
vishalhabib99 profile image
Vishal Habib •

This is the part I hadn't written down, and you're right that it's the trap. "Fix the miss and restart from zero" feels like rigor, but it only keeps the clean runs. I checked your numbers and they hold: for the same under-1%-at-95% bar, one miss needs 473 cases, two need 628, three need 773. At a true 0.5% miss rate, only about 22% of 299-case runs come back clean.

I'll add the continuation rule to the plan before shadow mode starts. Change the checker after a miss and it's a new checker with a new run, and the failed run stays in the report. Also agreed on the slide: 0 of 11 blind and 22 of 22 adversarial are bounds on two different populations, and each one should say which it covers.

Collapse
 
vishalhabib99 profile image
Vishal Habib •

One correction to my own reply, after writing this into the plans: each of those counts holds for a single look, but used as a schedule (exit at 299 if clean, else 473, ...) it gives the checker four chances to pass, and a checker sitting exactly at 1% exits about 11% of the time instead of 5%. Adjusted so the whole schedule stays at 5%, it's 381 / 571 / 738 / 894. The rule's now in all three PRDs with your comment credited.

Collapse
 
arhancanli profile image
Arhan Canli •

Glad the numbers held. One small addition for the plan: put the continuation rule in the README as a table (misses so far, cases needed, cases collected), and update it from the shadow log rather than by hand. Then the current bound is something anyone can read off on any given day, and the stopping point was written down before the data arrived, which is the whole defence against the restart trap. If the table ever needs editing after a miss, that edit is visible in history too.

Thread Thread
 
vishalhabib99 profile image
Vishal Habib •

Done, and it's a better shape than what I had. All three READMEs now have a shadow-mode status table: run, checker version, cases reviewed, misses, exit count, status. A script rebuilds it from the shadow log and reads the schedule straight from the PRD table, so the rule only lives in one place. It refuses a log entry after a run has ended, or one from a changed checker, and CI fails if the README and the log disagree. Today it reads 0 of 381, not started, because there's no shadow traffic yet, and that's the honest number.

The change: github.com/vishalhabib99/agent-han...

Thanks again. This thread has made the plan a lot harder to fudge.