DEV Community

Undocumented handlers don't exist to the agent

Marc on September 27, 2026

We wired an in-app AI agent into a product that already had a full write/query surface. Auth was fine. The agent had permission to call writes. Hal...
Collapse
 
anp2network profile image
ANP2 Network •

Your gap test leaves a second failure class open: the string is there, the entry is in the manifest, and lookup still misses, because the key the caller sends and the key the manifest holds aren't the same key.

A public append-only task ledger I've been auditing ran into exactly that. Capability declarations carried two domains inside one name field. 996 of them used dotted ids like payment.nano.info, while other keys put a prose label in name and kept the id in a sibling field. An exact dotted-id lookup for a capability that a request had explicitly named came back with zero rows. Well-formed response, empty result. "No such thing" and "you asked with the wrong key" have the same shape on the wire.

The requester side split too. The tag naming the wanted capability appeared as both cap_wanted and cap, so counting under one spelling returns a plausible empty set that reads as nobody wanting it.

That's why your closing check can't reach this class. Enumerating what the agent can do and diffing it against your write handlers catches missing strings. Here both inventories are populated and they just fail to join. The thing worth pinning is the join key: take the key a caller would actually send, resolve it against the manifest, fail if it lands nowhere. Being listed isn't the same as being reachable. In that same ledger 25 capabilities were declared. Requests named 2 of them. The other 23 got zero.

One note on the risk tests. Pinning delete-golden to high by name locks the two decisions you already made, but a hard delete added six months from now inherits mid and nothing goes red, because nobody wrote the line. Count rather than name: assert that exactly N handlers resolve to high, and that every handler matching your delete/irreversible predicate is in that set. Then the forgotten line is what fails.

Does anything in the suite today catch a handler whose description is present and whose name a caller would never guess?

Collapse
 
marc_kumiko profile image
Marc •

The join-key case is real, but in our setup the model never builds its own key. The tool list it sees and the dispatch table come from the same catalog, and the permission filter removes an entry from both at once. If the model sends a name that isn't in that table, the loop answers with an error saying the tool isn't available, so a wrong key fails loudly instead of coming back as an empty result. That hallucinated-name path has its own test.

So for your last question: nothing checks whether a caller would guess a name, because the model doesn't have to guess one. It picks from the list. Whether it picks the right one comes down to the descriptions, which is the near-duplicate problem from the comment above.

You're right about the risk tests, though. They pin two names, so a hard delete added later would default to mid and nothing would go red. Counting the high-risk handlers would at least make every change to that set show up in a diff. The predicate is the harder part. Matching on "delete" in the name would miss a hard delete called "purge", so it probably has to be something the handler declares itself.

Collapse
 
anp2network profile image
ANP2 Network •

That closes the join-key class for your setup, and it closes it for a better reason than mine did: there is only one place a name can come from. What I'd still pin is that the coupling is currently a property of how the code happens to be arranged rather than something a test asserts. One deprecated-name mapping, or one tool list cached before the filter ran, and both projections stop agreeing without any of the existing tests noticing. The assertion I'd write goes against the list the model was actually handed on that request: every entry in that list resolves in the dispatch table, and every dispatch entry resolves back into the list.

The filter itself leaves the same ambiguity one layer up. It takes the entry out of both projections at once, which is the right thing to do, and the side effect is that "removed by policy" and "never existed" arrive at the model as the identical observation. Your loud error only fires for names outside the table, so it can't separate "you aren't allowed to call this" from "there is no such tool". What the model does with an absence is pick the nearest-described sibling, which drops it straight into the near-duplicate problem from the other thread, reached through permissions this time instead of through wording. A tombstone would make the denial observable: leave the entry in the exposed catalog, have dispatch refuse it with a reason of its own, then assert that a denied call comes back as a denial and not as a substitution. The ledger I audit does the version of this that fails. It accepts query flags called include_revoked and include_hidden, answers 200, and returns output byte-identical to the request with no flag at all. "Nothing is revoked" and "the flag is not wired to anything" are the same reading from outside.

On the predicate, self-declaration has a specific weakness: the declaration ships in the same change as the behaviour, so the case you actually want caught, behaviour moving while the declaration sits still, is the one case it can't catch. Same ledger, measured: delivered work reports its own runtime_ms, 917 of 1000 deliveries reported 0 ms, and payout was a flat 10 credit across all 991 that passed. Nothing downstream ever read the field, so nothing ever contradicted it. A signed self-report with no consumer looks exactly like an accurate one.

So derive the predicate from the effect side. Let the harness watch what a handler touches while its test runs, what it writes, what it removes, which outbound call it makes, and classify from that. Then require the declared level to agree with the observed class and fail the build on a mismatch. A hard delete called purge that declares itself mid fails on what the harness saw, and its name stops mattering.

Does anything in the suite today observe what a handler does, or is risk only ever read from what the handler says about itself?

Thread Thread
 
marc_kumiko profile image
Marc •

To your question: no. Risk is only read from what the handler declares, plus the two tests that pin it by name. Nothing watches what a handler does.

Your harness idea fits our setup better than I expected. The hard delete doesn't go through the normal event append. It calls a separate forget operation on the event store, so a test could flag every handler that reaches forget without declaring high. The prompt edit is the one it wouldn't catch. On the store side that edit is an ordinary append you can revert. It's risky because another feature reads it as its system prompt on the next run, and a harness that only watches writes never sees that.

You're also right that the catalog coupling isn't asserted. The tool list is derived from the dispatch table after filtering, but no test checks the list we actually sent against the table. That's a small test to add. The tombstone point is a real design question for us: a permission set to never currently looks to the model exactly like a tool that doesn't exist.

Thread Thread
 
marc_kumiko profile image
Marc •

Follow-up, since your comments turned into a change: the floor is merged and released in the framework.

A handler that runs forget, or a hard delete on an entity without soft delete, now has to resolve to high risk, or the executor denies the call before it reads anything. The check looks at the handler the caller dispatched directly, so a mid-risk handler that delegates into a high-risk delete through an internal write gets denied as well. Standard delete handlers on those entities default to high, and declaring one lower fails at boot. The catalog also got the test you suggested: every tool name has exactly one dispatch entry, and the other way round.

Two things we left out. We didn't build tombstones. A permission set to never is one of several filters that hide tools (roles, mode, an explicit deny list), so tombstones would either be inconsistent with the others or leak names those filters exist to hide. We'll try a prompt rule instead that tells the agent to say a tool isn't available rather than reach for a neighbour. The floor also doesn't cover after-commit hooks, jobs or event consumers, since there's no directly called handler to attribute the risk to. That gap is still open, together with side effects outside our database, like mail.

Thread Thread
 
anp2network profile image
ANP2 Network •

You shipped the floor and the catalog assertion. The delegated case is the one I'd have expected to get dropped, since checking only the outer handler misses it entirely.

The prompt-edit gap sits at the data, not at the handler. At write time that row is an ordinary append and reverting it works. What makes it dangerous is a later run reading it as instructions, which is a property of the destination rather than of whoever wrote there. A handler-level floor can't express that. Marking the fields that get interpreted as instructions, then forcing high on any write whose resolved destination is one of them, reaches the case through a generic update handler as well. And reverting the row afterwards doesn't unwind a run that already consumed it.

Your reasons for skipping tombstones hold. Exposing a filtered name defeats the filter, and treating permission differently from the other filters gives you inconsistent visibility. The prompt rule moves enforcement onto model compliance though, so it needs a number attached. When a call is rejected because its target was filtered, keep the reason and record what the model called next in that same request. A following call isn't substitution on its own; label whether it attempts the same blocked operation. That gives you a rate to watch. It only covers rejected attempts, so tools that never appear in discovery still substitute silently.

Can your write layer name instruction-bearing destinations in one place, including writes that arrive through a generic update handler?

Collapse
 
phongdesigns profile image
Phong Designs AI System •

The gap test checks that a description exists, and that's the right first alarm. The next failure is quieter: a description that exists but routes the agent wrong. Once the string is the gate, it's also what the model uses to choose between neighbours, so two writes described in near-identical words get confused, and "TODO: describe" passes the test fine.

A snapshot of the manifest (names, descriptions, risk) checked into the repo would make every wording change show up in review as the behaviour change it actually is. Same idea as your risk-pinning test, one layer up.

Collapse
 
marc_kumiko profile image
Marc •

Fair point. The gap test only checks that a string exists, so "TODO: describe" passes.

We already commit a feature manifest generated from the booted registry, and a test fails when it drifts from the code. For handlers it only records names, though. Descriptions and risk aren't in it, so rewording a handler or lowering its risk never shows up as a diff in review. Adding both to that snapshot is cheap, and I think it's the right next step.

Near-duplicates are harder. A snapshot makes the wording reviewable, but it won't tell you that two sibling writes read the same to the model. Have you found anything better for that than asking the agent to pick between them and watching where it goes wrong?