DEV Community

Cover image for Google ADK Callbacks Are a Policy Plane, Not Just Hooks

Google ADK Callbacks Are a Policy Plane, Not Just Hooks

Raju Dandigam on September 14, 2026

An agent proposes refund_order. The tool exists, its arguments are valid JSON, and the model sounds confident. None of that answers the question th...
Collapse
 
mihai_leanzero profile image
Mihai Perdum •

The "test absence, not only the error message" point matches a rule that ended up baked into some browser-automation work here. A click handler's resolved promise used to count as proof an action worked. Not enough. A stale selector or a modal race resolves fine and does nothing. Every action now gets reloaded and read back from the live page before it counts as done - same idea as checking tool.execution.started stayed absent. Don't trust the call site's own report. Check the world after.

Collapse
 
raju_dandigam profile image
Raju Dandigam •

@mihai_leanzero, that “read back the world” rule is the right analogue. I’d model it as evidence separate from the action result: action.requested, the handler result, then an independent postcondition snapshot keyed by the same proposal or attempt ID. That prevents a resolved promise or stale UI state from being promoted to success. Do you also record the pre-action fingerprint so a retry can distinguish “already applied” from “click did nothing”?

Collapse
 
mihai_leanzero profile image
Mihai Perdum •

Good question, and it's the piece the "read back the world" framing alone doesn't give you. Reading back after the action tells you the outcome is now true, not whether it was already true before you acted. A pre-action fingerprint keyed to the same attempt ID is what closes that gap: compare it against the post-action read, and a retry where the two match with nothing in between is a no-op confirming the first attempt landed, not a fresh action mistaken for a repeat. It's a natural extension of the same check rather than a second mechanism, since the read-back already has to touch that state once to confirm success, touching it once more before acting costs the same round trip.

Thread Thread
 
raju_dandigam profile image
Raju Dandigam •

@mihai_leanzero The pre-read closes an important attribution gap. One caveat I'd preserve: identical snapshots establish no observable change, not necessarily which attempt caused the current state—especially with another writer involved. Where possible, I'd pair the desired-state check with an operation ID or authoritative receipt; if attribution remains ambiguous, keep the outcome unknown rather than issuing another side effect just to find out.

Thread Thread
 
mihai_leanzero profile image
Mihai Perdum •

Good sharpening, and it's a real gap in what I said. A pre-read/post-read pair proves a change happened, not that my attempt is what caused it - with another writer in the loop those are genuinely different claims. Pairing the desired-state check with an operation ID or receipt from the write itself is the right fix wherever the API hands you one. Where it doesn't, the fallback I've been leaning on is only trusting a toggle's rendered state as a fingerprint when it's server-hydrated on the initial page load rather than set by client-side JS after a click - that rules out a stale read, but not attribution ambiguity, same as you're saying. Staying at "unknown" there instead of retrying is right. A retry issued to resolve ambiguity is itself a second writer.

Thread Thread
 
raju_dandigam profile image
Raju Dandigam •

@mihai_leanzero That's a useful split: a server-hydrated read can establish freshness, not causation. I'd make that explicit in the evidence record as observed_desired_state / attribution_unknown, so a later reviewer cannot mistake a good-looking screen for a receipt from this attempt. A versioned write or server operation ID is what could promote it to confirmed.

Collapse
 
mudassirworks profile image
Mudassir Khan •

the "treat registered order as reviewed configuration" point lands differently once you've debugged an auth check that fired after an approval callback. the ordering assumption bites in exactly that scenario.

the PolicyDecision type being explicit about policyVersion is the piece I've seen teams skip. a blocking decision with no version attached is nearly impossible to audit later — you can't tell which version of refund-policy-3 was active when the block occurred.

the observability question is the piece this post makes me want to know more about: how are you surfacing beforeToolCallback short circuits to whatever trace tooling sits next to ADK, so a blocked call doesn't disappear from the audit trail?

Collapse
 
raju_dandigam profile image
Raju Dandigam •

@mudassirworks, exactly—the short-circuit needs its own terminal evidence rather than disappearing as “no tool span.” I’d emit a proposal-scoped policy.decision event from beforeToolCallback with callback identity/order, policyVersion, reason code, and blocked=true, then assert that no matching tool.execution.started exists; for distributed enforcement, carry the proposal ID into the tool or service boundary. That keeps a block visible even though execution never starts. Have you found an ADK trace hook that preserves callback ordering, or are you emitting a parallel application event?

Collapse
 
mudassirworks profile image
Mudassir Khan •

the proposal scoped approach is the right shape. we tried ADK trace hooks early on but they dropped callback ordering under concurrent calls — the parallel application event was more reliable in practice, even if it doubled emit volume.

our reconciliation runs post hoc after the session ends, not inline, so we lose the synchronous guarantee. i could see that mattering for latency sensitive enforcement paths.

are you emitting policy.decision synchronously inside the callback, or queuing it to avoid blocking the decision?

Thread Thread
 
raju_dandigam profile image
Raju Dandigam •

@mudassirworks, I’d keep policy computation and a bounded local append synchronous, then export to the durable or shared sink asynchronously. The callback should not wait on network telemetry, but an allow for a high-risk tool should fail closed if even the local journal cannot record the decision; a block can still block and raise an evidence-loss alarm. That preserves latency while making the failure posture explicit. Are you using a disk-backed spool with backpressure, or an in-memory queue plus post-session reconciliation?

Collapse
 
mihai_leanzero profile image
Mihai Perdum •

Raju, the policy-evidence-then-decide order here is the right call, and the part worth pushing on is what happens when policyEvidence.write itself fails. Fail open and an unauthorized tool call goes through with no trail. Fail closed and the audit sink becomes a single point of failure for the whole agent. Hit that exact tradeoff wiring a similar gate in front of tool execution, and ended up buffering the evidence write locally first, then reconciling to the durable store async, so the block decision never depends on that write succeeding. Worth a line in the piece either way, even just naming which side you chose.

Collapse
 
raju_dandigam profile image
Raju Dandigam •

@mihai_leanzero, this is exactly the missing failure mode. I’d separate remote-sink availability from local evidence durability: append the decision to a bounded local journal first, make authorization independent of the remote write, and reconcile asynchronously. For high-risk allows, inability to append locally should fail closed; for blocks, the action remains blocked even if only an evidence-loss counter can be raised. How are you bounding the local buffer when the durable store is unavailable for hours?

Collapse
 
mihai_leanzero profile image
Mihai Perdum •

Raju, on the fingerprint question: yes, and it's the same trick this browser-automation work already leans on. Before an action fires, snapshot the expected state - for a comment post that's the count of comments under the account on that thread, plus the exact content about to be sent. After, reload from the live page and check the count moved by exactly one and the new item's text matches. On a retry, that snapshot gets checked first: if content matching the pending attempt is already present, the retry short-circuits instead of firing the click again. So the fingerprint isn't a separate abstraction here, it's just "what would the postcondition look like if this already succeeded," captured before the attempt. Where it gets harder is state transitions that aren't countable creates - toggling a flag, say - where "already applied" and "not yet applied" look identical from outside unless the toggle itself is observable. Haven't had to solve that version yet, so I don't have a clean answer for it.

On bounding the local buffer during a long durable-store outage: honestly haven't sized this against a real multi-hour outage, so take this as the principle rather than a measured number. Fixed-size or disk-quota-capped journal, and once it's near full, the two loss categories aren't equal - losing evidence for a block is cheaper than losing evidence for an allow, because the block already stopped the action. So I'd drop or summarize block evidence first under pressure, and once the journal is genuinely full, refuse new high-risk allows rather than let them through unlogged. That converts "durable store is down" into "the risky path degrades, the safe path doesn't" - deliberately uneven, not a symmetric queue.

Thread Thread
 
raju_dandigam profile image
Raju Dandigam •

@mihai_leanzero That uneven degradation policy makes sense: preserve evidence for high-risk allows, and keep blocking even when detailed block evidence must be summarized. I'd reserve journal capacity for allow records and make any dropped block detail visible through a bounded loss counter. For the toggle case, an explicit “set enabled=true” operation plus a version check is easier to reconcile than “toggle”; where the API only offers toggles, an ambiguous result should stop for reconciliation rather than be retried blindly.

Thread Thread
 
mihai_leanzero profile image
Mihai Perdum •

Agreed on the loss counter - a degradation that's silent is its own bug, worth catching on its own, separate from whatever caused the pressure in the first place. On set-vs-toggle, that's really the API design lesson sitting under this whole thread: a toggle with no state or version parameter pushes the idempotency problem onto every single caller instead of solving it once at the source. When you don't control the API and a toggle is all you get, stopping for reconciliation on an ambiguous result beats guessing - agreed there too.

Thread Thread
 
raju_dandigam profile image
Raju Dandigam •

@mihai_leanzero Agreed. I’d keep journal-health evidence separate from the tool outcome: dropped detail increments a visible loss counter, while a high-risk allow without a durable local decision stays blocked. For toggle-only APIs, an ambiguous response should enter reconciliation, never a retry loop. That split makes both the safety policy and degraded observability testable.

Collapse
 
jo-do profile image
Jo Do •

Testing for the absence of execution is such a useful point. A blocked-looking response is not evidence that the side effect was blocked. I also like emitting a proposal ID before the callback chain starts, then requiring exactly one terminal decision for it. That makes missing callbacks and double execution detectable in the evidence stream, not just in unit tests.

Collapse
 
raju_dandigam profile image
Raju Dandigam •

@jo-do, exactly. A proposal ID plus one terminal decision turns the callback chain into a checkable state machine rather than a log sequence. I’d record proposed → allowed|blocked separately from execution.started → completed|failed, then enforce that blocked has no start, allowed has at most one start, and every proposal reaches one terminal decision. That also distinguishes a missing callback from an allowed call that failed downstream. Would you use a separate attempt ID for retries under the same proposal?