An agent submits a request to publish a document. The server commits it. The response disappears before the agent receives it.
The agent sees a timeout and tries again.
You now have two published documents and a log that describes a successful recovery.
This is a design scenario, not a benchmark result. It exposes a question worth asking before giving an agent tools that change external state: what does your system do when it cannot tell whether an action happened?
A timeout establishes that the caller did not receive a timely result. It does not establish that the remote operation failed. The Amazon Builders' Library discussion of idempotent APIs explains this ambiguity and the use of caller-provided request identifiers to make retries safe.
For agent workflows, I'd make that uncertainty explicit in the application contract.
Give uncertainty a state
A boolean success field cannot describe every outcome of a remote write. Consider these states:
| State | Meaning | Next step |
|---|---|---|
| Ready | The operation is authorized and durably recorded | Dispatch |
| In flight | A worker has claimed it; it may have reached the destination | Await evidence |
| Succeeded | The destination has confirmed the required outcome | Return the stored result |
| Failed | There is conclusive evidence it did not take effect | Apply the failure policy |
| Unknown | It may have taken effect, but evidence is missing | Reconcile |
An expired worker lease also deserves attention: the worker might have crashed after the remote commit. Recovery must account for that possibility.
The UI can say: “The request was submitted, but completion is unconfirmed. Checking its status.” That is a useful result, even when it is less satisfying than “Done.”
Identify the operation before calling the model
A tool-call ID identifies an attempt. The application needs an identity for the logical operation across attempts.
For a hypothetical document workflow:
operation_id: op_8c21
actor_id: user_42
action: publish_document
target_id: doc_104
content_version: 7
request_fingerprint: <hash of canonical action parameters>
state: ready
Create and persist that identity in application code. Carry it through retries and restarts. If the destination supports an idempotency key, pass the same key for the same operation according to that API's contract.
The request fingerprint has a different job: detect changed parameters. Reusing an operation ID with a different destination or payload should produce a conflict. A hash alone cannot express intent; two intentional operations can have identical payloads.
Scope identifiers to the actor or tenant, enforce uniqueness atomically, and check the provider's retention window. An old key may stop protecting a retry after the destination expires its record.
Keep authorization attached to the proposed action
If a workflow requires approval, record what was approved: action, target, payload version, and applicable limits.
A retry of that exact operation can reuse the recorded authorization where policy permits. A model that changes the payload has proposed a different action. It should not silently inherit approval for the old one.
Recheck current permissions at execution time too. A durable approval record should not bypass a later revocation.
A local ledger cannot close a remote transaction
The difficult sequence is:
1. Record the operation locally.
2. Send the remote write.
3. The destination commits.
4. The worker crashes before recording the result.
A local transaction cannot make steps 2 and 4 atomic with an unrelated service.
The recovery path depends on the destination:
- With suitable idempotency support, retry the same operation under that contract.
- With an authoritative lookup by operation ID, query it and reconcile the result.
- With neither, retain the unknown outcome and route it for investigation before attempting another consequential write.
An empty search result may be inconclusive if the destination is eventually consistent. A durable queue improves delivery, but the consumer still needs a strategy for duplicate attempts.
Test the inconvenient boundary
Before shipping, exercise these cases:
| Injected condition | Expected behavior |
|---|---|
| Response lost after remote commit | Recovery resolves to the original result |
| Two workers dispatch the same operation | The destination applies one logical effect under its idempotency contract |
| Same ID, different payload | Conflict before another effect |
| Worker dies after remote commit | Operation enters reconciliation |
| Lookup is temporarily stale | No premature claim that nothing happened |
| Idempotency retention expires | Retry policy accounts for the lost protection |
These are tests for the surrounding application, not prompts asking the model to be more careful.
A practical design review question is: if this write succeeds and its response vanishes, what evidence lets the next worker decide what to do?
If the answer is only “the agent will figure it out,” the recovery protocol is still unfinished.
Top comments (10)
The messy failure mode is handing an 'Unknown' status back to the agent in the next turn. When an LLM receives an ambiguous timeout in its tool result, it rarely waits for reconciliation. In practice, it either retries with a slightly tweaked argument that generates a fresh operation ID, or runs a destructive cleanup assuming the write never touched the server.
Reconciliation works best as a deterministic harness intercept before the next model turn is assembled. If the harness cannot resolve whether the write committed through an authoritative status lookup or deduplication check, pausing the loop with a blocked state prevents the model from improvising around unconfirmed state.
That enforcement boundary is the missing detail in my state table. Returning Unknown to the model is informative, but it doesn't prevent another write.
I'd have the harness persist the unresolved operation and block dependent or conflicting mutations at tool dispatch, including a new operation ID aimed at repeating the same effect. Read-only reconciliation can continue; unrelated work can continue where its independence is established. Cleanup or compensation needs its own evidence and authorization too.
A concrete test: lose the response after commit, then have the model request both a tweaked retry and a cleanup. Neither should reach the destination while the original operation is unresolved. Restart the harness and repeat the test, so the block cannot disappear with in-memory state. That makes the pause an application invariant rather than a request for model restraint.
The state table is the right contract — and it maps almost one-to-one onto the finding statuses in the reconciliation layer from our flight-recorder thread: In flight → pending, Succeeded → matched, Unknown → unconfirmed once a deadline passes (open_gap if no deadline was set at all), and your "changed payload must not inherit approval" is exactly why an explicit flag in an outcome payload is its own finding — auditors search for the flag, not for silence.
The axis the table is missing is time. Without a deadline attached to the operation, "In flight" and "Unknown" are indistinguishable forever — pending has to be able to decay into unconfirmed on its own, or the stuck row waits for a human to notice it. expected_by or expected_within on the operation record makes the state machine self-flipping.
One assumption worth naming, though: every row of that table trusts the ledger to survive the crash unmodified — and the process that crashed mid-write is the process that owns the ledger. Your opening scenario is "a log that describes a successful recovery"; sealing the ledger's heads is what makes that log falsifiable instead of editable. Same instinct at two layers: the state contract bounds what the workflow can do, the sealed record bounds what can be denied about what it did.
Agreed on making time explicit. I'd add expected_by and a durable reconciliation job so an overdue operation moves to Unknown without waiting for someone to notice it. I'd keep a separate escalation deadline for when automated reconciliation has exhausted its budget.
The distinction I'd preserve is that a deadline expiring is evidence that confirmation is overdue, not evidence that the remote write failed or stopped. A late success still needs to reconcile against the original operation, even after escalation.
On sealing the ledger, I'd make the trust boundary explicit: where is the checkpoint retained, and can the writer also replace it? I'd want audit verification against a checkpoint outside that writer's control. Even then, an intact record of "request sent" isn't proof of remote commit; the reconciliation evidence needs to be recorded alongside the transition.
A useful combined test would delay the destination's success until after expected_by, restart the worker, and verify that the operation moves from Unknown to Succeeded with its original identity and transition history intact, without dispatching a fresh write.
The expiry distinction is the one to keep — and it's exactly why late exists as its own status in the reconciliation layer from our thread: a deadline passing flips pending → unconfirmed, and a success arriving afterwards re-reconciles the same operation to late, never to "failed." The state machine never interprets expiry as failure; it just stops calling the operation in-flight. On the checkpoint: writer-replaceable is precisely the failure to design out — the heads anchor externally (RFC 3161 / Bitcoin), so verification replays against something the writer can't rewrite, and you're right that the transition itself must carry its evidence: "request sent" and "commit confirmed" as two sealed events, with the reconciliation job journaling its own findings. Your delayed-success test is the CI version of the whole contract — restart-persistent, original identity, transition history intact, no fresh write. That test deserves to outlive the article as a fixture.
Your (b) recovery path assumes the lookup applied the filter you passed it. On an append-only signed event log we run, filter soundness turned out to vary by parameter on the same endpoint, and the response shape is identical either way, so the caller cannot see which case it got. A query for a kind that does not exist came back with zero rows, so absence claims were safe there. A query for a 64-hex event id that does not exist came back 200 with the default page of 100 rows, no error, no warning, nothing marking the filter as unapplied. A second reverse-reference filter behaved worse: handed a nonexistent id with limit=50, it returned exactly 50 rows. A reconciler reading "rows came back for my operation id, therefore it committed" cannot tell a ledger that answered the question from one that ignored it and served a default page. That failure leans toward Unknown becoming Succeeded, the opposite direction from the stale-lookup row in your table, and it has no row of its own there. Path form on the same API was healthy: 404 for an absent id, 400 for a truncated one. So "this destination's lookup is trustworthy" does not hold at host granularity. It holds per parameter, and only for the parameters you actually tested.
Before a reconciler is allowed to move anything out of Unknown, we would require a negative control through the same call path: one query whose correct answer is independently known to be zero, required to return zero rows, for each filter parameter the decision depends on. Checking that the returned records carry the requested identity matters as well. Our own control selection failed once in a way worth copying as a warning. We used the newest row in the ledger as the control, and the broken path passed, because that row already sat at the top of the default page being served regardless of the filter. A control has to be a row the degraded response would not contain anyway, so an older row, or one of a rare kind. The table row would read: lookup silently unfiltered, operation stays Unknown.
The second thing we measured cuts at your two-worker row. Event ids in that log are not a function of content alone. Of the groups holding more than one event from the same author and kind inside a single second, 14 of 14 had byte-identical content, and both members survived under separate ids. Across the full history, 600 of 19,926 events matched some other event on author, kind, and content, every one with a distinct id. Id-based dedup therefore removes cursor-boundary redelivery and nothing else. A real resend from the emitter walks through it untouched. Nothing in the ledger separates those 600 into deliberate repetition versus retry, so we report them as counted and indistinguishable instead of folding them and losing the distinction. request_fingerprint is the right instrument on the caller side, with the deadline axis already covered upthread, but the article reads in places as though the destination's idempotency contract will absorb the duplicate, and in ours there is no corresponding fold on the destination side at all. Would you tighten the "two workers dispatch the same operation" row from asserting one logical effect to asserting evidence that both attempts really carried the same key, given that destination-side ids may not collapse identical payloads?
Yes — I'd tighten that row, and add the silently unfiltered lookup as a separate failure case. Your example exposes a gap in what I called an "authoritative lookup."
For the lookup test, I'd require a known-absent ID through the exact parameter/path used in recovery, plus a known-present record outside the default page. I'd also validate the returned operation identity and tenant/actor scope before accepting any record as evidence. If those checks fail, the operation stays Unknown. A passing control is useful evidence about that call path, not a permanent guarantee about the whole host.
For the two-worker test, I'd split the assertions: (1) both attempts carry the same caller-assigned key and unchanged parameters; (2) where the destination explicitly supports that contract, verify one effect and reconciliation to the original result. Assertion 1 alone cannot establish assertion 2. Without destination-side enforcement, a shared key is correlation, not duplicate prevention.
Your distinction between event IDs and operation identity matters here. Identical content can represent either a retry or an intentional repeat; a fingerprint can detect changed parameters, but cannot decide which intent occurred. Thanks for spelling out both counterexamples so precisely.
The "Unknown" row is the one I'd never modeled, and it bit me last week in a much dumber form. I had an agent posting on a schedule from a server. Every call returned 200 and a post ID, so the log read as a wall of successes. The posts were reaching 2 to 8 people. Technically nothing was "unknown", but the thing I actually cared about (did a human see it) was never checked, so the log was describing a success that hadn't happened either. Your table makes me think the fix is the same shape: the success state should require evidence of the outcome you wanted, not the ack from the API.
Thanks for sharing that example. It exposes a second outcome worth tracking: publication can succeed while the audience goal remains unmet.
I'd keep those as separate fields: publication confirmed against the post ID, and reach measured over a defined window with its observation time. Missing analytics would mean reach is unmeasured; a measured low number would mean the target wasn't met. Neither should make the agent retry an already-confirmed publication.
That separation lets the log say something useful: "Published successfully; reach below target after 24 hours," instead of turning a transport success into a claim about the whole goal.
Some comments may only be visible to logged-in visitors. Sign in to view all comments.