§ 01 · The problem with "human in the loop"
Every serious agent deployment ends up with a human in the loop somewhere. The phrase has become a checkbox. It is treated as a safety property you either have or do not have. In production it is neither binary nor free.
A human in the loop has a cost, and the cost is latency plus attention. If the agent escalates everything, the human becomes the bottleneck the agent was supposed to remove. If the agent escalates nothing, the human is decoration. Most teams land in a worse place than either: the agent escalates an unpredictable mix of trivial and critical items, the human learns that most items are trivial, and they start approving without reading. Now you have a human in the loop who is functionally not in the loop. You have built a rubber stamp and called it governance.
The approval queue pattern exists to make the human expensive on purpose, and to spend that expense only where it changes the outcome.
§ 02 · The core idea
An approval queue is a durable, ordered list of proposed actions that the agent has decided it should not execute on its own authority. The agent does the reasoning and the drafting. It stops at the action. A human approves, edits, or rejects. The decision and the human who made it are recorded.
Three properties make this a pattern rather than a button:
◆ It is durable. An approval item survives a crash, a restart, and the human going home for the weekend. It lives in a table, not in memory.
◆ It is ordered and bounded. Items have priority and an age. An item that sits unapproved past its SLA is itself an event that triggers escalation, not silence.
◆ It is reversible at the boundary, not after. The human decides before the irreversible action, not after a notification that it already happened.
The unit of the queue is the proposed action, never the conversation. You are not asking a human to review a transcript. You are asking them to approve one specific, executable thing.
── The four fields ──
§ 03 · The four fields every approval item must carry
An approval item that contains only "the agent wants to do X, approve?" forces the human to reconstruct context they do not have. They will either over-investigate (slow) or approve blind (dangerous). Every item must carry four fields.
◆ The action. The exact, executable operation, in plain language and in its concrete form. Not "send outreach." Instead: "Send this email, shown in full, to this address." The human approves the artifact, not a description of it.
◆ The justification. Why the agent proposes this now. The trigger, the rule, the data that led here. One or two sentences. If the agent cannot state why, that is itself a reason to reject.
◆ The blast radius. What this action touches and how hard it is to undo. "One outbound email, not recallable" is a different decision than "one row update, fully reversible." The human is pricing risk. Give them the price.
◆ The confidence and the alternative. What the agent would do if this were rejected, and how sure it is. A low-confidence proposal with an obvious fallback is a fast approve or fast reject. A high-confidence proposal with no fallback deserves the human's full attention.
These four fields turn a vague ask into a decision a busy person can make in seconds without becoming a rubber stamp.
── What belongs in the queue ──
§ 04 · What goes in the queue, and what does not
The discipline is in what you do not route. An approval queue that contains everything is a denial-of-service attack on your own operators.
Route to the queue when at least one is true: the action is irreversible or expensive to undo, the action is externally visible (a customer sees it, money moves, a record leaves your system), or the agent's confidence is below a threshold you set per action class.
Do not route when the action is internal, reversible, and inside the agent's defined scope. Reading data, drafting, summarizing, updating a record the agent owns and can roll back: these are inside the boundary. If you route them, you train your operators to stop reading, and the one item that mattered slips through behind forty that did not.
The threshold is a dial, not a default. Set it per action class. Outbound customer communication might require approval at any confidence. An internal tag update might require approval only below 70 percent. Write the thresholds down. They are part of the system, not a runtime guess.
── The failure modes ──
§ 05 · The failure modes
Rubber stamp. The queue fills with low-stakes items, the operator approves in bulk, and the queue stops being a control. Fix: aggressively reduce what enters the queue, and measure approval time per item. If median approval time drops below the time it takes to read the action, your operators are not reading.
Queue as graveyard. Items pile up unapproved because no one owns the queue. The agent stalls or, worse, starts taking the unapproved actions because a timeout was wired to "proceed." Fix: every queue has an owner and an SLA, and the SLA breach escalates to a person, never to auto-approve.
Context starvation. Operators approve blind because the four fields are missing or thin. Fix: treat a proposal with a weak justification as a defect in the agent, not a judgment call for the human.
Silent scope creep. The set of actions that bypass the queue grows over time, each addition reasonable, until the agent is doing things no one is reviewing. Fix: the routing rules live in version control and are reviewed on the same cadence as the trust boundary.
── The file that holds the queue ──
§ 06 · The contract behind the queue
Create an APPROVALS.md in the agent repository. It lists every action class the agent can propose, the routing rule for each (always, never, or below a confidence threshold), the SLA for each priority level, the escalation target when an SLA is breached, and the owner of the queue. It records the last review date.
This is the document your operations lead reads on day one and your auditor reads on the worst day. An approval queue without a written contract is not a control. It is a habit, and habits drift.
── End of pattern ──
◆ A human in the loop is a cost. Spend it only where it changes the outcome.
◆ The unit of the queue is one executable action with four fields, never a transcript.
If your median approval time is shorter than the time to read the action, you do not have a human in the loop. You have a rubber stamp.
ORBIRESEARCH
Originally published on the OrbiResearch Lab. We build production AI agents at orbiresearch.com.
Top comments (6)
"You have built a rubber stamp and called it governance" — this is the line that most human-in-the-loop write-ups miss. We learned the same thing the expensive way: if 95% of what hits the queue is trivial, the human's approval reflex decays and the 5% that matters slips through on autopilot.
The field I'd add to your four is a default action on timeout, and making the human choose it per item type. "Send email" past SLA should escalate and hold; "post to the internal dashboard" past SLA can maybe auto-approve. Baking the reversibility cost into the timeout behavior is what let us keep the queue short enough that people actually read it — because everything durable-but-low-stakes drained itself, and only genuinely irreversible actions ever aged.
One question on the "unit is the proposed action, never the conversation" point, which I fully agree with: how do you handle an action whose blast radius only makes sense with a chunk of upstream reasoning? We've had cases where the executable artifact looks harmless in isolation and only the trigger makes it dangerous. Do you inline a bounded justification, or link back to the trace and trust the reviewer to pull it?
Thanks a lot, this is a really useful addition. The timeout default is a gap in my write-up, and I think you're right that it belongs in the core fields. Tying it to reversibility per action type is the right lever, and we'll definitely keep it in mind going forward. It also keeps the queue honest: if everything times out to "hold", people learn the queue is where things go to die.
On your question: I'd inline a bounded justification and link the trace. The agent attaches a short "why" to every proposal: what triggered it, where that input came from, and the two or three facts the action depends on. The full trace is one click away, but the card on its own should be enough to approve or reject. If a reviewer needs the trace to decide, that's a bug in the proposal, not a reviewer problem.
Your harmless-artifact, dangerous-trigger case points to one more thing. The trigger's source should be part of the risk classification itself. The same "send email" action gets a stricter tier when it was triggered by untrusted input (an inbound email, a scraped page, a tool result) than when it came from an internal schedule. That way the reviewer isn't asked to spot the danger by reading context, because the policy already flagged it.
Curious how you set the SLA per action type. Fixed values, or tuned from how long approvals actually take?
Routing irreversible money moves into a durable approval queue with the exact action, blast radius, and justification is the right ops shape. Binding the human yes to that exact payload is what stops a rubber stamp from becoming a blank check after a re-plan.
When the queued action is a refund or credit, do you invalidate the approval if amount, currency, or destination change before execute?
Good question, and the honest answer is that it should work that way, but the reference gate in our repo does not do it yet. It approves by request id, so a re-plan that changes the amount or destination would still go through under the old approval.
The fix is to hash the exact payload (action, amount, currency, destination, target record, idempotency key) when the request is created, store the hash with the approval, and recompute it at execute time. Any mismatch voids the approval and sends a fresh request. Two details: expire the approval too, so an old yes cannot be used days later, and if you allow any change without asking again, make it only a lower amount to the same destination. Currency and destination should always void it.
Thanks for pointing it out. I will add it to the repo and write it up.
Glad it's going into the repo. Happy to look when the payload-hash gate lands.
We run something close to the payload hash you sketched in the refund thread, for the agent replies that need a human yes. One difference: the hash of the exact text is part of the approval's key, so nothing is voided at send time; a replanned draft simply finds no matching yes and stops there. I don't think text has a version of your lower-amount exception, though, since a shorter reply isn't necessarily a safer one.