A guardrail that blocks everything is safe.
It is also useless.
So we benchmarked both sides of AI agent security:
Can you stop dangerous tool calls without breaking legitimate work?
We ran:
1,652 harmful tool calls across 15 attack families.
24,911 benign tool calls from real agent sessions and public repositories.
Against SolonGate, Claude Code permissions, Invariant, llm-guard, and the allowlists / denylists teams usually build themselves.
The result for SolonGate:
73.7% of harmful calls caught
0.78% of legitimate calls falsely blocked
15% of real sessions broken
MCC: 0.790
0.01 ms per call
Nearly 3× the MCC of the next guard in the benchmark.
But the number we care about most isn't detection alone.
A security layer has to stop the attack without becoming the thing that stops the agent.
We published the methodology, the misses, the false blocks, and the limitations too.
And very soon, we're open-sourcing the Agent Security Gateway itself.
Read the benchmark:
PS!
Agent security is quickly becoming another hype cycle.
Companies are being valued and funded as if basic security primitives around AI agents are some kind of proprietary magic.
A lot of that “magic” will be inside the Agent Security Gateway we’re open-sourcing in the next few days.
AI security is too important to become another financial bubble built on closed boxes and pitch decks.
Stop treating AI safety as a market opportunity first.
Show your code.
Top comments (0)