Here's a pattern I keep seeing.
A team wires up an AI agent that can do real things — send emails, run commands, query and modify the database, call external APIs. The demo is magical. It reads a request, figures out the steps, takes them, reports success. Everyone's impressed, and it ships.
Then comes the first incident. It emails the wrong list. It runs a destructive command against the wrong environment. It reads a web page that quietly tells it to do something nobody asked for, and it obliges. Suddenly the magical demo is a very un-magical cleanup, and someone's asking how this was allowed to happen.
Here's the thing: the demo was never the hard part. Getting an agent to do something is genuinely easy now. The hard part — the part that almost always gets skipped in the rush — is everything that keeps the agent from hurting you when it inevitably does the wrong thing. And it will do the wrong thing, because it's a probabilistic system acting in an unpredictable world.
So this is the checklist I think belongs before you let an agent touch anything that matters. Not the fun part. The part that separates an agent you can actually deploy from a liability with a good demo.
1. Least privilege — the agent can only reach what it strictly needs
This is the highest-leverage guardrail, and it's the one most worth getting right first, because it makes entire categories of disaster simply impossible.
An agent that cannot reach production cannot wipe production. An agent that cannot send money cannot be talked into sending money. An agent with no access to a system can't be the cause of an incident in that system, no matter how confused or compromised it gets. So the question to ask before anything else isn't "what should this agent be able to do?" — it's "what is the least it needs to do its job?" Then give it exactly that and nothing more.
Almost every catastrophic agent story, when you trace it back, turns out to be a permissions decision someone made without quite noticing — handing over credentials that could reach production at all, or a tool scope broader than the task required. The destructive action got the headline, but the over-broad grant was the actual mistake, made quietly, long before. Least privilege is the guardrail that catches the error at the point where it's cheapest to prevent.
2. Human approval for consequential actions — gated by blast radius
Some actions you can let an agent take freely. Some you absolutely should not let it take unattended. The skill is in drawing that line well.
The irreversible or high-impact ones — send, spend, delete, export, deploy — should pause for a human. But two things make or break this guardrail. First, don't gate everything. An agent that asks permission forty times a session trains the human to click "approve" without reading, and now your approval step is theater — the click records attendance, not consent. Reserve the gate for actions that actually warrant it.
Second, the approval has to be meaningful. "Approve this action? y/n" on something the human can't evaluate is a rubber stamp. Show the blast radius: what this affects, what it will cost, what it's about to change, what the relevant history is. Give the approver something concrete enough to actually reject. An approval made blind isn't a control; it's a liability with a signature on it.
3. Treat everything the agent reads as untrusted
Anything your agent ingests — web pages, emails, documents, a teammate's file, the output of a tool, a code comment — can carry instructions. This is prompt injection, and it's not a solved problem: you cannot reliably detect malicious instructions hidden in natural language, because the space of ways to phrase them is endless and attackers adapt to whatever filter you deploy.
So don't put your faith in detecting the bad input. Put the control on the action instead. The set of consequential things an agent can do (send, spend, delete, export) is small and you can enumerate it — so gate those few calls with hard checks that don't depend on the model's judgment about whether the content it just read was trustworthy. The practical rule of thumb: an agent becomes dangerous when it has private data access, exposure to untrusted content, and the ability to act externally, all at once. Remove any one leg of that triad and a successful injection has far less it can reach.
4. A reviewer that can actually say "no" — and has been tested saying it
For consequential actions, a second check — a validation step, or a separate "judge" agent whose job is to review the proposed action before it executes — adds real protection. A system that has to convince an independent reviewer is much harder to push into a bad outcome than one that just acts on its own first impulse.
But here's the part people skip: a reviewer that always approves is not a control. It feels like safety, it generates a reassuring log, and it stops exactly nothing. Before you trust a reviewer, hand it a deliberately bad action — one you know should be refused — and confirm that it actually blocks it. A guardrail you've never watched engage is not a guardrail; it's a hope with good UI. Test the "no," or you don't have one.
Better still, make that test permanent: have a deliberately-failing probe run on a schedule and emit a receipt like any other check, so "the reviewer has been tested saying no" becomes a standing, provable property of the system instead of a memory of one afternoon. (Credit: @slabb.)
5. An independent audit trail — because the witness can't be the suspect
When something goes wrong, you'll go to the logs to reconstruct what happened. And here's the trap: if the logs were written by the agent, they were written by the very thing that misbehaved. A confused or compromised agent doesn't produce a broken log that tips you off — it produces a clean one, a tidy record of a bad decision, which is worse, because a clean log makes you stop looking.
So the record has to be produced somewhere the agent doesn't control — the infrastructure or supervisor layer around it, not the agent's own self-report. Seal it so that tampering leaves a visible gap rather than a silent edit. And capture not just what the agent did but what it believed at the time — which environment it thought it was in, which target it thought it was acting on — because the action is usually defensible given a wrong belief, and the belief is the part that actually explains the incident.
A quick, disclosed note on what this looks like in practice: I work on xenition.com, an AI workspace, and this exact cluster — approval gates, an independent audit log, a second reviewer — is something we had to build in deliberately rather than bolt on afterward. In our setup an agent's consequential actions pass through an approval gate, every step is written to an audit log the agent itself doesn't author, and a separate judge-agent reviews an action before it runs. I'm not holding it up as the answer — plenty of stacks assemble these pieces differently, and the right shape depends on your system. I mention it only as a concrete illustration that none of this is theoretical: the approval gate, the independent record, and the second reviewer are things you can actually ship today, and increasingly things you should.
6. Blast-radius limits — assume it will go wrong, and bound how bad
The guardrails above reduce how often things go wrong. This one accepts that something eventually will, and makes sure no single mistake is catastrophic.
Put hard limits around the agent: spending caps, rate limits, quotas on how many actions it can take before it has to check in. Prefer reversible actions by default — soft deletes over hard ones, staged rollouts over all-at-once, drafts over sends. The mindset shift is the important part: stop trying to guarantee the agent never errs (you can't), and start guaranteeing that when it does, the damage is small, contained, and recoverable. A mistake that costs you a reversible change and a shrug is a completely different thing from one that costs you a database.
7. Observability that measures the effect, not the report
Finally, watch the right thing. A dashboard built to answer "did the agent report success?" is, by construction, a dashboard built to trust the witness — and we just covered why the witness can't be trusted. The agent will happily report a valid-looking success while pointing at the wrong target, and your green dashboard will tell you everything is fine.
So measure the world, not the agent's account of it. Did the invariant hold? Does the total still balance? Is the referenced record actually there? Is the system still in a consistent state? These checks are more annoying to write than "did it return 200," and they are the only ones that catch the failure where the report looks perfect and the reality is wrong. Measure the effect, not the self-report.
The takeaway
Capability is the easy part. It's also the exciting part, which is exactly why it gets all the attention and all the demo time — and why the guardrails, which are none of those things, get left for "later," which often means "after the incident."
But an agent's real value in production was never what it can do. It's what it can do safely, repeatably, and recoverably — what it can do without becoming the thing you spend next quarter cleaning up after. The seven items above aren't the glamorous part of building with AI. They're the part that decides whether the glamorous part survives contact with the real world.
Build the brakes before you build the engine. Or at the very least, before you take it out on the highway.
If you've shipped an agent to production, I'm genuinely curious: which of these did you have in place before your first incident — and which one did you add right after it taught you the hard way? Most of us learned at least one of these the expensive way. Which was yours?
Top comments (7)
Which did I have before the first incident: none — I built them after, which is why I recognize where several of these came from. Sections 4, 5 and 7 were ground out in public this week: the tested-no reviewer is the negative-control receipt from the sunnydachs thread, "seal it so tampering leaves a visible gap" is the witness conversation from your last piece — whose lineage edit, promised Monday, is still pending — and "measure the world, not the report" is the process_ok=true finding: the frameworks' health channel missed five out of five wrong-count runs because the health signal and the failure live in different artifacts.
Two items I'd add to the checklist, from the same week's threads: (1) above "seal it," cross-stream reconciliation — one sealed chain vouches for who claimed what and when; two sealed chains written by different parties about the same event can disagree, and disagreement between tamper-evident sources is the only evidence about the claims themselves. (2) make the negative control permanent — the deliberately-failed probe emits a receipt like every other check, so "the reviewer has been tested saying no" becomes a property of the chain instead of a memory of one afternoon. The checklist is the right shape. Its provenance should travel with it.
You're right, including the uncomfortable part: the lineage edit I promised Monday is still pending, and a piece arguing "provenance should travel with the artifact" has no business shipping with its own unattributed. Fixing it — crediting sections 4, 5, 7 to the threads they came from.
Both additions go in. Cross-stream reconciliation above "seal it" (disagreement between two sealed sources is the only evidence that reaches the claims themselves), and the permanent negative control — a deliberately-failed probe emitting a receipt turns "tested once" into a standing property of the chain. Memory decays; a receipt doesn't.
One thing I’d add is that these safeguards shouldn’t be treated as independent checklist items. Their real value comes from the failure modes they cover together. Least privilege limits what the agent can reach, approval controls what it can execute, and independent verification checks whether the resulting state matches the intended outcome. If one layer fails, the next layer should still have enough information to catch the mistake. That suggests testing the guardrails as failure chains, not just testing each control in isolation. For example, deliberately bypassing an approval path or giving the agent misleading context should still leave the system with a separate boundary capable of preventing or detecting the consequential action.
Testing them as failure chains rather than isolated checks is the move most people miss, and it's the right one — a checklist invites you to tick each control green on its own, but the real question is whether layer N+1 still catches the mistake when layer N fails. Defense-in-depth only means anything if the layers are independent, so the test isn't "does approval work?" — it's "if I bypass approval, or feed the agent misleading context, does a separate boundary still prevent or detect the consequential action?" Deliberately breaking one layer and watching whether the next one holds is how you find out if you have real depth or just a row of controls that all fail together. Going in the revision — the failure-chain framing is the thing that turns a checklist into an architecture.
"Forty approvals a session trains the human to click without reading" has a corollary: the prompt has to carry enough to decide with. "Agent wants to run a command" is unapprovable and gets clicked through regardless. The command, the target, and what changes have to be visible.
Least privilege also has a second axis , reachability. An agent that can't route to production beats one told not to use production credentials. Instructions are preferences; boundaries are controls.
The item missing from most checklists is reversibility. If the worst action is undoable, every other guardrail gets cheaper to get slightly wrong.
The approval corollary is the sharper version of my point — gating isn't enough if the prompt can't carry the decision. "Agent wants to run a command → approve?" is unapprovable, so it gets clicked through, which means a gate with no blast-radius detail is just slower rubber-stamping; the command, the target, and the diff have to be in the prompt or the human is approving a shape, not a decision.
And "instructions are preferences; boundaries are controls" is the cleanest statement of the least-privilege point anyone's made — told not to use production is a sticky note the agent can ignore or be injected past; can't route to production is physics. The reachability axis beats the permission axis every time, because one is enforced and the other is hoped.
Reversibility being the missing item is right, and it's the one that makes the whole checklist cheaper: if the worst action is undoable, every other guardrail is allowed to fail a little, because no single miss is terminal. It's the slack that makes imperfect controls survivable. Going in the revision — that's the item I shouldn't have left out.
The independent audit trail is the part most teams skip. If the agent writes its own log, a clean record of a bad decision is worse than a missing one, because the reviewer stops looking. Proof of data means the record is produced outside the agent, with the inputs and the decision captured, not only the happy-path result.