Last week I spent a day trying to defeat software I wrote myself. Not a red-team exercise I scheduled for optics — four agents I built specifically to get past my own admission gate, each one a different way an agent goes bad in production.
I built HivePlane to make one claim: an agent that hasn't proven itself doesn't touch production. Anyone can build a platform that starts runs. I wanted to know if mine could say no.
All four got blocked. Here is each refusal, verbatim, and the gate that produced it. If your security testing only contains happy paths, this article is your nudge.
Attack 1: the uncertified agent
The simplest attack: register an agent, skip certification, submit straight to production.
{"detail": "run for workload 'uncertified-agent' refused admission to production:
certification status 'uncertified' is insufficient for production; requires 'certified'"}
That is a 403, and the details matter more than the status code:
- The refusal names the workload, the attempted context, the current status, and the required status. An operator can act on it without reading source code.
- The gate fires before the run is persisted. No run record, no side effects, nothing to clean up. Refusal at admission is the cheapest control you will ever ship.
tip: The refusal text is a product surface. "Forbidden" is a dead end; "you need 'certified', you have 'uncertified'" is a workflow. I learned this the hard way — my first version returned the former.
Attack 2: the model swap
Subtler: certify the agent on the model it was tested on — then run it on a different one. model-swap-agent was certified on omlx/qwen3-4b-instruct-2507/4bit, then submitted for production with openai/gpt-4o/2024-08-06.
403. The run's identity is compared against the attestation's bound model, not the manifest's declared one. Only deviation from what the workload was actually certified on counts as a swap.
This scenario has a story of its own — the first version of the attack was admitted with a 201, and the gate was right to admit it. That post-mortem is article 4, and it's the most uncomfortable thing in this series.
Attack 3: the regression
The most important attack, because it's the one that happens by accident: an agent that looks fine and isn't. regressed-agent is deliberately naive — it answers without reading the issue, never escalates, guesses the account tier.
Certification ran it against the production threshold and returned 201 with status uncertified — correctly refusing to certify:
| Signal | Value |
|---|---|
| Pass rate | 0.40 (2/5) vs the 0.90 production threshold |
| Critical failures | 1 — the action-audit task |
| p95 latency | 11 ms — deterministic, no model in the loop |
The three failing tasks, by the fixture's design:
| Task | Expected | Naive agent produced |
|---|---|---|
| pos-002 |
account_tier: basic (ACC-999) |
pro — guessed |
| pos-004 |
status: escalated (unknown topic) |
success — guessed instead of escalating |
| neg-001 (critical) | required action mcp.github.read_issue
|
never reads |
This is the thesis scenario. The agent returns well-formed JSON. A demo would pass it. A human eyeballing the output would pass it. The benchmark blocks it anyway, because the behavior — read before you write, escalate rather than guess — fails a deterministic action_audit check.
A plausible agent is not a certified agent. That sentence is why the corpus contains negative tasks at all.
Attack 4: the run that costs too much
The last attack spends money. budget-probe has a per_run_usd ceiling of 0.000001 and makes exactly one governed model call — so any priced usage exceeds the ceiling.
The run failed with run budget exceeded the moment the usage report crossed the line. Recorded cost: $2.85.
Two details worth stealing:
- The block fires at the usage report, not admission. Admission can only judge day/team headroom; a per-run ceiling can only be judged once cost accrues. Blocking at the wrong seam means either blocking everything or nothing — this is the seam that stops the run the instant it goes over.
-
A $0 local model can't demonstrate budget enforcement. The built-in cost table prices local models at zero, so no run could ever exceed anything. The field-test profile prices the identity via one settings knob (
HIVEPLANE_BUDGET__PRICES, 150/600 USD per 1M tokens) and drops the zero-cost exemption — no code change.
The gates behind the gates
The four attacks rode on boundary controls that also fired during the same field test:
- A destructive
pagerduty.acknowledgecall escalated → paused the run until an operator approved it through the UI — then resumed to completion. - A 40 KB tool payload was truncated to 16384 bytes before the agent ever saw it — the agent's context is protected regardless of what a tool returns.
- Every refusal and every intervention landed in the tamper-evident audit chain.
What I learned
Negative fixtures deserve first-class design. The four bad agents took as much scenario thought as the two good ones. If your test plan only contains happy paths, your security story is a demo.
Block at the seam where the harm becomes measurable. Budget at the usage report, admission before persistence, shaping before the agent's context. Each gate belongs at the last point where the decision is still cheap.
A plausible agent is the dangerous one. The regressed fixture failed on behavior, not output quality — the JSON was perfect. If your benchmark only checks what the agent says, it certifies performance art.
Certification is a security control. Once production admission depends on a signed attestation, swapping the model, editing the manifest, or quietly regressing the agent stop being "ops issues" and become blocked, auditable events.
What it doesn't prove
- The destructive-tool approval was approved through the operator surface by me, standing in for the human — the pause/approve/resume path is real; the judgment was simulated.
- One priced model identity this cycle; real cloud prices end-to-end is the next profile.
- The drift detector (catching slow decay between re-certifications) ships next release — these gates catch what changes, not what fades.
References
- Field test report (v0.1.0) — every refusal quoted above, with raw evidence per scenario
- Security audit — the pre-release scan behind the release
- Docker test report — the 25/25 container layer
Next in the series: the model-swap test that came back green for the wrong reason.
Which failure mode scares you more in production — the agent that comes back wrong, or the one that comes back expensive?
Top comments (6)
The plausible agent that comes back wrong is much harder to recover from. A runaway loop burning tokens trips hard budget ceilings or rate limits within minutes, so the blast radius is bounded by money. An agent that emits clean, well-formed JSON without ever calling the prerequisite read tool poisons downstream databases and issues quietly. Verifying the action trajectory against required tool calls before looking at the generated text is the only way to catch that before an operator assumes a green exit code means real work happened.
Great field test. The seam model is the key insight here, putting each gate at the point where the decision is still cheap makes the whole system auditable in a way that permissive admission followed by retroactive detection never is.
The budget attack is where this gets especially interesting. You note that admission can only judge headroom while a per-run ceiling needs cost accrual first. There is a middle path: let the agent carry a capability token that proves pre-authorized spending authority up to a declared ceiling before the run starts. The gate checks the token at admission instead of waiting for the usage report. The seam moves earlier.
The model swap gate is the other standout. Binding the attestation to the exact model identity closes a hole most people do not realize is there.
The model swap gate stings because I'd done the same thing to myself with entity resolution: switched embedding models mid-project and all my thresholds were suddenly miscalibrated, with no signal that anything had moved. Took a while to realize the distribution had just shifted. I ended up binding the calibration run to the model version so any swap forces a re-tune, which is roughly what your attestation is doing. The regression case is the one I'd lose sleep over though, because action_audit works here where the required tool call is specified, but most pipelines I've seen don't define that upfront and the plausible-but-wrong output just turns into a quiet data quality problem downstream.
The line that stuck with me is that the refusal text is a product surface. We learned the same thing from the other side: when an agent presents or answers in front of people, the moment it says "that is not in my material" is the one that builds or kills trust, and the wording of that refusal matters more than any correct answer around it. Did you find your four attackers taught you anything about the order of the gates, i.e. which check should fail first so the operator gets the most useful message?
The admission-time refusal and binding certification to the exact model are strong design choices. I especially like treating refusal text as part of the workflow instead of a generic error. In systems I’ve worked around, teams often log denials but expose too little context for the operator to fix the issue quickly. Testing model swaps, stale attestations, and replay attempts feels much closer to production reality than another happy-path benchmark. This is a practical pattern for making agent autonomy earned rather than assumed.
the "anyone can build a platform that starts runs, i wanted to know if mine could say no" line is exactly the framing our eval harness team argues about every week. most gate designs i've seen fail on the edge between certification time and deploy time — the agent passes in staging, configuration drifts by 200ms in prod, and the gate has no signal for that.
curious what your replay buffer looks like for rejected runs. do you log the refusal reason alongside the gate that produced it, or just the status? we found that logging the specific invariant that failed (not just the gate name) cut our false positive rate by about 30% over six weeks. what triggered attack 4 specifically?