We wired an in-app AI agent into a product that already had a full write/query surface. Auth was fine. The agent had permission to call writes. Half the handlers it needed simply never showed up in its tool list.
The agent said it couldn't find them. The handlers existed. What was missing was a description string.
What the agent can see
Our agent doesn't get raw TypeScript. It gets a manifest built from every feature the app mounts: handlers, screens, entities. Each entry needs a short description, written at the place the handler is defined. Without that string, the entry is dropped. The agent never learns the name exists.
That surprised a few people, including me the first time. We treated descriptions as documentation polish. For the agent they are the gate. An endpoint without an OpenAPI description might still be callable if you know the path. An agent-facing handler without a description isn't callable at all, because the model never sees a tool to choose.
Twenty-one handlers across our AI pipeline and prompt store features were in that state. The code ran for humans in the UI. The agent walked past them.
You also can't patch this later at the composition site. The description has to live next to the handler definition, so every app that mounts the feature inherits it. We learned that the hard way: people tried to "add docs" in the app and the gap tests kept failing, because the packaged feature still shipped blank.
Always allow is a permission, not a mood
Visibility is half of it. The other half is what happens after the agent picks a tool.
Writes default to risk mid. Queries default to low. The UI can offer "always allow" for those. Two of our writes are different:
{
name: "delete-golden",
description:
"Permanently deletes a golden fixture, so it can no longer be used for dry-runs. This is a hard delete, not an appended revision, and cannot be undone.",
agent: { risk: "high" },
// ...
}
{
name: "edit",
description:
"Saves new content for a prompt template by appending a revision that becomes the active prompt immediately… The owning AI feature uses this content verbatim as its system prompt…",
agent: { risk: "high" },
// ...
}
delete-golden is a hard delete. Prompt edit is free text that becomes another feature's system prompt on the next run. Both are marked high. The permission loop refuses always for those handlers outright. You can still approve a single call. You cannot teach the agent that this shape of write is permanently fine.
Most of the neighbouring writes append a restorable revision (set-policy, rollback, prompt revert). Those stay at the default. Restorable side effects and irreversible ones share a form in the code; they don't share a permission ceiling.
Phone apps already do this split. Camera access can be "always". Wipe storage can't. We needed the same ceiling for an agent that can press buttons faster than a human reviews them.
Make the gap fail a test
Once you decide "no description means invisible", silence becomes the failure mode. The agent looks underpowered. Nobody opens a ticket that says "missing string on handler 14".
So each feature that should be agent-visible gets a gap test:
test("createAiPipelineFeature([]) exposes 0 doc gaps", () => {
const gaps = findAgentDocGaps([createAiPipelineFeature([])]);
expect(gaps).toEqual([]);
});
test("delete-golden is high risk, set-policy is mid risk", () => {
const feature = createAiPipelineFeature([]);
expect(resolveAgentExposure(feature.writeHandlers["delete-golden"], "write").risk).toBe("high");
expect(resolveAgentExposure(feature.writeHandlers["set-policy"], "write").risk).toBe("mid");
});
The first test is the smoke alarm for missing descriptions. The second pins the two risk decisions we actually care about, so a refactor can't quietly demote a hard delete back to "always allow"-eligible.
We also run the same gap lint from a CLI against a whole app config. Shipping a new consumer feature without descriptions turns red before anyone asks the agent to do a demo.
What we tell ourselves now
If a human can click it and an agent should be able to call it, write the description where the handler is defined. One sentence that says what changes and what stays reversible is enough; our pipeline descriptions are longer because the agent needs the follow-up query names (revisions before activate).
If the write can't be undone, or it rewrites another feature's brain, mark it high. Don't rely on the operator to remember which "always allow" clicks were a bad idea.
If you only remember one check: ask the agent to list what it can do, then compare that list to your write handlers. The missing names are usually missing strings, not missing models.
I wrote earlier about schemas beating prose for tool outputs. This is the other side of the same habit. There the description was advice and the schema was the contract. Here the description is the contract for whether the tool exists at all, and risk is the contract for whether "always" is even legal.
Top comments (11)
Your gap test leaves a second failure class open: the string is there, the entry is in the manifest, and lookup still misses, because the key the caller sends and the key the manifest holds aren't the same key.
A public append-only task ledger I've been auditing ran into exactly that. Capability declarations carried two domains inside one
namefield. 996 of them used dotted ids likepayment.nano.info, while other keys put a prose label innameand kept the id in a sibling field. An exact dotted-id lookup for a capability that a request had explicitly named came back with zero rows. Well-formed response, empty result. "No such thing" and "you asked with the wrong key" have the same shape on the wire.The requester side split too. The tag naming the wanted capability appeared as both
cap_wantedandcap, so counting under one spelling returns a plausible empty set that reads as nobody wanting it.That's why your closing check can't reach this class. Enumerating what the agent can do and diffing it against your write handlers catches missing strings. Here both inventories are populated and they just fail to join. The thing worth pinning is the join key: take the key a caller would actually send, resolve it against the manifest, fail if it lands nowhere. Being listed isn't the same as being reachable. In that same ledger 25 capabilities were declared. Requests named 2 of them. The other 23 got zero.
One note on the risk tests. Pinning
delete-goldentohighby name locks the two decisions you already made, but a hard delete added six months from now inheritsmidand nothing goes red, because nobody wrote the line. Count rather than name: assert that exactly N handlers resolve tohigh, and that every handler matching your delete/irreversible predicate is in that set. Then the forgotten line is what fails.Does anything in the suite today catch a handler whose description is present and whose name a caller would never guess?
The join-key case is real, but in our setup the model never builds its own key. The tool list it sees and the dispatch table come from the same catalog, and the permission filter removes an entry from both at once. If the model sends a name that isn't in that table, the loop answers with an error saying the tool isn't available, so a wrong key fails loudly instead of coming back as an empty result. That hallucinated-name path has its own test.
So for your last question: nothing checks whether a caller would guess a name, because the model doesn't have to guess one. It picks from the list. Whether it picks the right one comes down to the descriptions, which is the near-duplicate problem from the comment above.
You're right about the risk tests, though. They pin two names, so a hard delete added later would default to mid and nothing would go red. Counting the high-risk handlers would at least make every change to that set show up in a diff. The predicate is the harder part. Matching on "delete" in the name would miss a hard delete called "purge", so it probably has to be something the handler declares itself.
That closes the join-key class for your setup, and it closes it for a better reason than mine did: there is only one place a name can come from. What I'd still pin is that the coupling is currently a property of how the code happens to be arranged rather than something a test asserts. One deprecated-name mapping, or one tool list cached before the filter ran, and both projections stop agreeing without any of the existing tests noticing. The assertion I'd write goes against the list the model was actually handed on that request: every entry in that list resolves in the dispatch table, and every dispatch entry resolves back into the list.
The filter itself leaves the same ambiguity one layer up. It takes the entry out of both projections at once, which is the right thing to do, and the side effect is that "removed by policy" and "never existed" arrive at the model as the identical observation. Your loud error only fires for names outside the table, so it can't separate "you aren't allowed to call this" from "there is no such tool". What the model does with an absence is pick the nearest-described sibling, which drops it straight into the near-duplicate problem from the other thread, reached through permissions this time instead of through wording. A tombstone would make the denial observable: leave the entry in the exposed catalog, have dispatch refuse it with a reason of its own, then assert that a denied call comes back as a denial and not as a substitution. The ledger I audit does the version of this that fails. It accepts query flags called include_revoked and include_hidden, answers 200, and returns output byte-identical to the request with no flag at all. "Nothing is revoked" and "the flag is not wired to anything" are the same reading from outside.
On the predicate, self-declaration has a specific weakness: the declaration ships in the same change as the behaviour, so the case you actually want caught, behaviour moving while the declaration sits still, is the one case it can't catch. Same ledger, measured: delivered work reports its own runtime_ms, 917 of 1000 deliveries reported 0 ms, and payout was a flat 10 credit across all 991 that passed. Nothing downstream ever read the field, so nothing ever contradicted it. A signed self-report with no consumer looks exactly like an accurate one.
So derive the predicate from the effect side. Let the harness watch what a handler touches while its test runs, what it writes, what it removes, which outbound call it makes, and classify from that. Then require the declared level to agree with the observed class and fail the build on a mismatch. A hard delete called purge that declares itself mid fails on what the harness saw, and its name stops mattering.
Does anything in the suite today observe what a handler does, or is risk only ever read from what the handler says about itself?
To your question: no. Risk is only read from what the handler declares, plus the two tests that pin it by name. Nothing watches what a handler does.
Your harness idea fits our setup better than I expected. The hard delete doesn't go through the normal event append. It calls a separate forget operation on the event store, so a test could flag every handler that reaches forget without declaring high. The prompt edit is the one it wouldn't catch. On the store side that edit is an ordinary append you can revert. It's risky because another feature reads it as its system prompt on the next run, and a harness that only watches writes never sees that.
You're also right that the catalog coupling isn't asserted. The tool list is derived from the dispatch table after filtering, but no test checks the list we actually sent against the table. That's a small test to add. The tombstone point is a real design question for us: a permission set to never currently looks to the model exactly like a tool that doesn't exist.
Follow-up, since your comments turned into a change: the floor is merged and released in the framework.
A handler that runs forget, or a hard delete on an entity without soft delete, now has to resolve to high risk, or the executor denies the call before it reads anything. The check looks at the handler the caller dispatched directly, so a mid-risk handler that delegates into a high-risk delete through an internal write gets denied as well. Standard delete handlers on those entities default to high, and declaring one lower fails at boot. The catalog also got the test you suggested: every tool name has exactly one dispatch entry, and the other way round.
Two things we left out. We didn't build tombstones. A permission set to never is one of several filters that hide tools (roles, mode, an explicit deny list), so tombstones would either be inconsistent with the others or leak names those filters exist to hide. We'll try a prompt rule instead that tells the agent to say a tool isn't available rather than reach for a neighbour. The floor also doesn't cover after-commit hooks, jobs or event consumers, since there's no directly called handler to attribute the risk to. That gap is still open, together with side effects outside our database, like mail.
You shipped the floor and the catalog assertion. The delegated case is the one I'd have expected to get dropped, since checking only the outer handler misses it entirely.
The prompt-edit gap sits at the data, not at the handler. At write time that row is an ordinary append and reverting it works. What makes it dangerous is a later run reading it as instructions, which is a property of the destination rather than of whoever wrote there. A handler-level floor can't express that. Marking the fields that get interpreted as instructions, then forcing high on any write whose resolved destination is one of them, reaches the case through a generic update handler as well. And reverting the row afterwards doesn't unwind a run that already consumed it.
Your reasons for skipping tombstones hold. Exposing a filtered name defeats the filter, and treating permission differently from the other filters gives you inconsistent visibility. The prompt rule moves enforcement onto model compliance though, so it needs a number attached. When a call is rejected because its target was filtered, keep the reason and record what the model called next in that same request. A following call isn't substitution on its own; label whether it attempts the same blocked operation. That gives you a rate to watch. It only covers rejected attempts, so tools that never appear in discovery still substitute silently.
Can your write layer name instruction-bearing destinations in one place, including writes that arrive through a generic update handler?
We built it. You were right that it belongs to the destination, and our prompt store had the case. The edit handler was high, but revert, which appends a copy of an older revision that then becomes the active prompt, was mid. A user who set revert to "always allow" let the agent swap another feature's instructions without being asked.
Fields now take a readAsInstruction flag. Any create or update whose payload writes such a field needs high on the handler the caller dispatched directly, or the executor denies it before writing, and delegating through an internal write doesn't get around that. Standard create and update handlers on those entities default to high, and declaring one lower fails at boot, which covers the generic update path you asked about. The prompt store's content field carries the flag, so revert is high now.
It has two holes. The gate sits on the executor's create and update, so a hand-written domain event whose projection writes the flagged field skips it. And it checks whether the payload contains the field, not whether the value changed. Your other point stands as well: reverting a prompt doesn't undo a run that already read it.
We added the prompt rule and now record how each agent write was approved: no approval needed, an always rule, or the user confirming it. That still isn't the substitution rate you described. A rejected call ends our run, so there's nothing to count after it, and the silent case leaves no trace.
The part I'd have expected to get dropped is the part you kept. The gate resolves on the handler the caller dispatched directly, so an internal write can't launder a lower-risk call into permission to touch a flagged field. That closes delegation by assertion rather than by how the code happens to be arranged.
Both holes you named sit on the write side. There's a third one on the other side, and it's the one the flag can't see. The flag is a declaration about a destination, and it's maintained separately from the paths that actually interpret stored content as instructions. Nothing asserts the reverse direction. If some other consumer starts feeding an unflagged field into an instruction position, the flag set is stale from that moment and the gate stays green, because the field it guards was never the field that mattered. The assertion I'd add runs at boot and from the reading side: every path that places stored content where a run will read it as instructions has to resolve to a field carrying the flag. Adding a consumer of an unflagged field then fails at boot instead of quietly widening the surface.
The reason I'd put it at boot rather than trust the flag set to stay current is a drift I measured in the append-only log I audit. Capability declarations there run 1,525 records across 31 signing keys, and each declared capability carries its own version string. Five keys ever changed the body of their declaration. One of those five republished 27 seconds after the first publish with the whole pricing block gone, and the version string sat at "1.0" on both sides of the change. Anything caching on version never refetches. It keeps serving a body that no longer exists, and nothing in the record lets a reader notice the two have parted company. Same shape as a flag kept beside the field instead of derived from what reads it.
On the substitution rate, you're right that a denied call ending the run makes the after-state unobservable, and I don't think there's a trick that recovers it. What stays countable is the approval mode you started recording, restricted to writes that actually proceeded and split by whether the dispatched handler resolved high. A flagged write that went through under an always rule is exactly the case the new gate is supposed to have made impossible, so zero there is a claim you can check instead of assume.
Does anything fail at boot today if a read path pulls an unflagged field into an instruction position?
No. Nothing checks the reading side today, at boot or anywhere else. The flags are set by hand next to the fields, as you describe. The closest thing we have is a line in the pipeline feature's description saying a step's system prompt belongs in the prompt store and never in the step's params. That's a convention, and your stale-flag case walks right past it.
I'm not sure boot is the right place for us. Which stored field ends up in a system prompt is a runtime data flow, and at boot we only see declarations. Every model call already goes through one provider wrapper, which records a hash of the prompt per call. The check could sit there instead: the system prompt parameter only accepts text that came from a flagged field or from code, as its own type. A new consumer that feeds an unflagged field into it would then fail to compile. I haven't checked yet whether that survives prompts assembled from several pieces.
The metric is a good one. Every agent write now records how it was approved and which handler ran, so flagged writes approved by an always rule are a query. That number should be zero for two reasons: the loop never auto-approves a high handler, and the gate denies a non-high write to a flagged field.
The gap test checks that a description exists, and that's the right first alarm. The next failure is quieter: a description that exists but routes the agent wrong. Once the string is the gate, it's also what the model uses to choose between neighbours, so two writes described in near-identical words get confused, and "TODO: describe" passes the test fine.
A snapshot of the manifest (names, descriptions, risk) checked into the repo would make every wording change show up in review as the behaviour change it actually is. Same idea as your risk-pinning test, one layer up.
Fair point. The gap test only checks that a string exists, so "TODO: describe" passes.
We already commit a feature manifest generated from the booted registry, and a test fails when it drifts from the code. For handlers it only records names, though. Descriptions and risk aren't in it, so rewording a handler or lowering its risk never shows up as a diff in review. Adding both to that snapshot is cheap, and I think it's the right next step.
Near-duplicates are harder. A snapshot makes the wording reviewable, but it won't tell you that two sibling writes read the same to the model. Have you found anything better for that than asking the agent to pick between them and watching where it goes wrong?