DEV Community

Cover image for Google ADK Callbacks Are a Policy Plane, Not Just Hooks
Raju Dandigam
Raju Dandigam

Posted on

Google ADK Callbacks Are a Policy Plane, Not Just Hooks

An agent proposes refund_order. The tool exists, its arguments are valid JSON, and the model sounds confident. None of that answers the question that matters: is this action allowed now?

Google's Agent Development Kit gives us interception points before and after agents, models, and tools. The tempting implementation is to add a log statement to each callback. The more useful design is to treat those callbacks as a small policy plane.

That does not mean placing every business rule in a callback. It means using the boundary to make four decisions explicit:

  • may this operation proceed?
  • should its inputs be normalized?
  • which budget or authorization applies?
  • what evidence should be recorded?

A callback can change control flow

ADK callbacks are not passive listeners. A beforeModelCallback can return a response and skip the model call. A beforeToolCallback can return a tool result and skip the tool. After-callbacks can replace results. That makes callback return values part of application behavior, not merely observability. The ADK callback documentation explains these interception semantics for TypeScript and the other supported SDKs.

Callback order is part of that behavior. ADK can run callback lists and plugin callbacks in sequence, with earlier callbacks able to modify inputs seen by later ones or short-circuit the remaining chain. Treat the registered order as reviewed configuration: authentication and tenant resolution should not accidentally run after a callback that can approve or execute an operation.

Start with a small decision vocabulary:

type PolicyDecision = {
  outcome: "allow" | "block";
  reasonCode: string;
  policyVersion: string;
};

function authorizeTool(
  toolName: string,
  args: Record<string, unknown>,
): PolicyDecision {
  if (toolName === "refund_order" && args.approved !== true) {
    return {
      outcome: "block",
      reasonCode: "APPROVAL_REQUIRED",
      policyVersion: "refund-policy-3",
    };
  }

  return {
    outcome: "allow",
    reasonCode: "POLICY_SATISFIED",
    policyVersion: "refund-policy-3",
  };
}
Enter fullscreen mode Exit fullscreen mode

This decision is deterministic and reviewable. It does not ask the model whether its own proposed action is safe.

Keep the callback thin

The callback should delegate to policy code rather than growing into an untestable collection of if statements. The following is deliberately schematic because exact tool callback types can change between ADK releases:

async function beforeTool({ tool, args }: {
  tool: { name: string };
  args: Record<string, unknown>;
}) {
  const decision = authorizeTool(tool.name, args);

  await policyEvidence.write({
    tool: tool.name,
    outcome: decision.outcome,
    reasonCode: decision.reasonCode,
    policyVersion: decision.policyVersion,
  });

  if (decision.outcome === "block") {
    return {
      error: "Tool execution blocked by policy",
      reasonCode: decision.reasonCode,
    };
  }

  return undefined; // continue with normal tool execution
}
Enter fullscreen mode Exit fullscreen mode

Register that function through the agent's beforeToolCallback. Use beforeModelCallback for model-input limits or serving an approved cache entry, and afterToolCallback for result normalization and outcome evidence.

The boundary becomes easier to review when responsibilities stay distinct:

Boundary Suitable responsibility
Before model Input limits, budget checks, approved cache lookup
Before tool Authorization, argument validation, freshness checks
After tool Result normalization, outcome classification
After agent Final evidence and response-level validation

Test absence, not only the error message

A guardrail test is incomplete if it checks only the returned text. The callback might return “blocked” after the real tool already ran.

it("does not execute a refund without approval", async () => {
  const refund = vi.fn();
  const decision = authorizeTool("refund_order", {
    orderId: "synthetic-123",
    approved: false,
  });

  if (decision.outcome === "allow") await refund();

  expect(decision.reasonCode).toBe("APPROVAL_REQUIRED");
  expect(refund).not.toHaveBeenCalled();
});
Enter fullscreen mode Exit fullscreen mode

At integration level, preserve evidence that the policy callback ran and the protected tool did not. That is stronger than matching a model-generated refusal sentence.

Also test the installed callback chain, not only the pure policy function. A unit test can prove authorizeTool() works while one runner quietly omits the plugin. A useful integration assertion records the callback name, policy version, decision, proposal ID, and whether execution began. For a blocked proposal, tool.execution.started must be absent.

Callbacks are not the entire security architecture

A callback runs inside the application process. It can contain a bug, be omitted from another agent, or be registered in the wrong order. High-risk authorization should also be enforced at the tool or service boundary. The application still needs least-privilege credentials, idempotency, audit records, and tests.

Google explicitly recommends ADK Plugins for modular security guardrails instead of scattering individual callbacks across agents. That is the right graduation path when the same policy must apply consistently to several agents.

Plugins still run inside your application. They improve reuse and consistency; they do not replace authorization at the resource server. Think of the callback or plugin as the orchestration decision point and the downstream service as the final enforcement point.

Make policy visible

Callbacks are most valuable when they expose a control point that already exists conceptually. Give each decision a stable reason code and policy version. Record bounded facts, not raw prompts or hidden reasoning. Test both the allowed path and the absence of a forbidden side effect.

Used that way, an ADK callback is more than a hook. It is where an agent proposal becomes an application decision.

References

Top comments (18)

Collapse
 
mihai_leanzero profile image
Mihai Perdum •

The "test absence, not only the error message" point matches a rule that ended up baked into some browser-automation work here. A click handler's resolved promise used to count as proof an action worked. Not enough. A stale selector or a modal race resolves fine and does nothing. Every action now gets reloaded and read back from the live page before it counts as done - same idea as checking tool.execution.started stayed absent. Don't trust the call site's own report. Check the world after.

Collapse
 
raju_dandigam profile image
Raju Dandigam •

@mihai_leanzero, that “read back the world” rule is the right analogue. I’d model it as evidence separate from the action result: action.requested, the handler result, then an independent postcondition snapshot keyed by the same proposal or attempt ID. That prevents a resolved promise or stale UI state from being promoted to success. Do you also record the pre-action fingerprint so a retry can distinguish “already applied” from “click did nothing”?

Collapse
 
mihai_leanzero profile image
Mihai Perdum •

Good question, and it's the piece the "read back the world" framing alone doesn't give you. Reading back after the action tells you the outcome is now true, not whether it was already true before you acted. A pre-action fingerprint keyed to the same attempt ID is what closes that gap: compare it against the post-action read, and a retry where the two match with nothing in between is a no-op confirming the first attempt landed, not a fresh action mistaken for a repeat. It's a natural extension of the same check rather than a second mechanism, since the read-back already has to touch that state once to confirm success, touching it once more before acting costs the same round trip.

Thread Thread
 
raju_dandigam profile image
Raju Dandigam •

@mihai_leanzero The pre-read closes an important attribution gap. One caveat I'd preserve: identical snapshots establish no observable change, not necessarily which attempt caused the current state—especially with another writer involved. Where possible, I'd pair the desired-state check with an operation ID or authoritative receipt; if attribution remains ambiguous, keep the outcome unknown rather than issuing another side effect just to find out.

Thread Thread
 
mihai_leanzero profile image
Mihai Perdum •

Good sharpening, and it's a real gap in what I said. A pre-read/post-read pair proves a change happened, not that my attempt is what caused it - with another writer in the loop those are genuinely different claims. Pairing the desired-state check with an operation ID or receipt from the write itself is the right fix wherever the API hands you one. Where it doesn't, the fallback I've been leaning on is only trusting a toggle's rendered state as a fingerprint when it's server-hydrated on the initial page load rather than set by client-side JS after a click - that rules out a stale read, but not attribution ambiguity, same as you're saying. Staying at "unknown" there instead of retrying is right. A retry issued to resolve ambiguity is itself a second writer.

Thread Thread
 
raju_dandigam profile image
Raju Dandigam •

@mihai_leanzero That's a useful split: a server-hydrated read can establish freshness, not causation. I'd make that explicit in the evidence record as observed_desired_state / attribution_unknown, so a later reviewer cannot mistake a good-looking screen for a receipt from this attempt. A versioned write or server operation ID is what could promote it to confirmed.

Collapse
 
mudassirworks profile image
Mudassir Khan •

the "treat registered order as reviewed configuration" point lands differently once you've debugged an auth check that fired after an approval callback. the ordering assumption bites in exactly that scenario.

the PolicyDecision type being explicit about policyVersion is the piece I've seen teams skip. a blocking decision with no version attached is nearly impossible to audit later — you can't tell which version of refund-policy-3 was active when the block occurred.

the observability question is the piece this post makes me want to know more about: how are you surfacing beforeToolCallback short circuits to whatever trace tooling sits next to ADK, so a blocked call doesn't disappear from the audit trail?

Collapse
 
raju_dandigam profile image
Raju Dandigam •

@mudassirworks, exactly—the short-circuit needs its own terminal evidence rather than disappearing as “no tool span.” I’d emit a proposal-scoped policy.decision event from beforeToolCallback with callback identity/order, policyVersion, reason code, and blocked=true, then assert that no matching tool.execution.started exists; for distributed enforcement, carry the proposal ID into the tool or service boundary. That keeps a block visible even though execution never starts. Have you found an ADK trace hook that preserves callback ordering, or are you emitting a parallel application event?

Collapse
 
mudassirworks profile image
Mudassir Khan •

the proposal scoped approach is the right shape. we tried ADK trace hooks early on but they dropped callback ordering under concurrent calls — the parallel application event was more reliable in practice, even if it doubled emit volume.

our reconciliation runs post hoc after the session ends, not inline, so we lose the synchronous guarantee. i could see that mattering for latency sensitive enforcement paths.

are you emitting policy.decision synchronously inside the callback, or queuing it to avoid blocking the decision?

Thread Thread
 
raju_dandigam profile image
Raju Dandigam •

@mudassirworks, I’d keep policy computation and a bounded local append synchronous, then export to the durable or shared sink asynchronously. The callback should not wait on network telemetry, but an allow for a high-risk tool should fail closed if even the local journal cannot record the decision; a block can still block and raise an evidence-loss alarm. That preserves latency while making the failure posture explicit. Are you using a disk-backed spool with backpressure, or an in-memory queue plus post-session reconciliation?

Collapse
 
mihai_leanzero profile image
Mihai Perdum •

Raju, the policy-evidence-then-decide order here is the right call, and the part worth pushing on is what happens when policyEvidence.write itself fails. Fail open and an unauthorized tool call goes through with no trail. Fail closed and the audit sink becomes a single point of failure for the whole agent. Hit that exact tradeoff wiring a similar gate in front of tool execution, and ended up buffering the evidence write locally first, then reconciling to the durable store async, so the block decision never depends on that write succeeding. Worth a line in the piece either way, even just naming which side you chose.

Collapse
 
raju_dandigam profile image
Raju Dandigam •

@mihai_leanzero, this is exactly the missing failure mode. I’d separate remote-sink availability from local evidence durability: append the decision to a bounded local journal first, make authorization independent of the remote write, and reconcile asynchronously. For high-risk allows, inability to append locally should fail closed; for blocks, the action remains blocked even if only an evidence-loss counter can be raised. How are you bounding the local buffer when the durable store is unavailable for hours?

Collapse
 
mihai_leanzero profile image
Mihai Perdum •

Raju, on the fingerprint question: yes, and it's the same trick this browser-automation work already leans on. Before an action fires, snapshot the expected state - for a comment post that's the count of comments under the account on that thread, plus the exact content about to be sent. After, reload from the live page and check the count moved by exactly one and the new item's text matches. On a retry, that snapshot gets checked first: if content matching the pending attempt is already present, the retry short-circuits instead of firing the click again. So the fingerprint isn't a separate abstraction here, it's just "what would the postcondition look like if this already succeeded," captured before the attempt. Where it gets harder is state transitions that aren't countable creates - toggling a flag, say - where "already applied" and "not yet applied" look identical from outside unless the toggle itself is observable. Haven't had to solve that version yet, so I don't have a clean answer for it.

On bounding the local buffer during a long durable-store outage: honestly haven't sized this against a real multi-hour outage, so take this as the principle rather than a measured number. Fixed-size or disk-quota-capped journal, and once it's near full, the two loss categories aren't equal - losing evidence for a block is cheaper than losing evidence for an allow, because the block already stopped the action. So I'd drop or summarize block evidence first under pressure, and once the journal is genuinely full, refuse new high-risk allows rather than let them through unlogged. That converts "durable store is down" into "the risky path degrades, the safe path doesn't" - deliberately uneven, not a symmetric queue.

Thread Thread
 
raju_dandigam profile image
Raju Dandigam •

@mihai_leanzero That uneven degradation policy makes sense: preserve evidence for high-risk allows, and keep blocking even when detailed block evidence must be summarized. I'd reserve journal capacity for allow records and make any dropped block detail visible through a bounded loss counter. For the toggle case, an explicit “set enabled=true” operation plus a version check is easier to reconcile than “toggle”; where the API only offers toggles, an ambiguous result should stop for reconciliation rather than be retried blindly.

Thread Thread
 
mihai_leanzero profile image
Mihai Perdum •

Agreed on the loss counter - a degradation that's silent is its own bug, worth catching on its own, separate from whatever caused the pressure in the first place. On set-vs-toggle, that's really the API design lesson sitting under this whole thread: a toggle with no state or version parameter pushes the idempotency problem onto every single caller instead of solving it once at the source. When you don't control the API and a toggle is all you get, stopping for reconciliation on an ambiguous result beats guessing - agreed there too.

Thread Thread
 
raju_dandigam profile image
Raju Dandigam •

@mihai_leanzero Agreed. I’d keep journal-health evidence separate from the tool outcome: dropped detail increments a visible loss counter, while a high-risk allow without a durable local decision stays blocked. For toggle-only APIs, an ambiguous response should enter reconciliation, never a retry loop. That split makes both the safety policy and degraded observability testable.

Collapse
 
jo-do profile image
Jo Do •

Testing for the absence of execution is such a useful point. A blocked-looking response is not evidence that the side effect was blocked. I also like emitting a proposal ID before the callback chain starts, then requiring exactly one terminal decision for it. That makes missing callbacks and double execution detectable in the evidence stream, not just in unit tests.

Collapse
 
raju_dandigam profile image
Raju Dandigam •

@jo-do, exactly. A proposal ID plus one terminal decision turns the callback chain into a checkable state machine rather than a log sequence. I’d record proposed → allowed|blocked separately from execution.started → completed|failed, then enforce that blocked has no start, allowed has at most one start, and every proposal reaches one terminal decision. That also distinguishes a missing callback from an allowed call that failed downstream. Would you use a separate attempt ID for retries under the same proposal?