Crash-Resilient Long-Running Agents in Microsoft Foundry: Leases, Checkpoints, and Side-Effect Fencing
Day 1 of 100 — Microsoft Foundry 100 Days / 100 Blogs series
Your agent has been running for six minutes. It already called three tools, burned a few thousand tokens analyzing a document, and it's halfway through generating a report. Then the container gets OOM-killed by the orchestrator during a routine scale-in event.
What happens next?
If you built this on a typical "background task" pattern — a Celery worker, a Kubernetes Job, a bare asyncio.create_task — the answer is usually: nothing happens next. The work is gone. The client polling for a result eventually times out or gets a failed status with no indication of how much progress was actually made. Somebody re-runs the whole pipeline, the tool calls that had side effects (sending an email, charging a card, writing to a downstream system) may fire twice, and you've quietly built a distributed systems bug into your AI product.
This is not a hypothetical edge case. The moment an agent's turn takes longer than a few seconds — multi-step tool use, long document generation, deep research loops, human-in-the-loop approval gates — you are running a long-lived, stateful process on infrastructure that will eventually restart it out from under you: redeploys, autoscaling, spot eviction, node patching, or a plain crash. Handling that correctly by hand means building a job ledger, a lease/heartbeat mechanism, idempotency keys, and a resume protocol. Most teams either skip it (and eat duplicate side effects) or spend weeks reinventing it badly.
Microsoft Foundry's Agent Service ships this as a first-class preview capability: durable work identity, lease-based crash detection, and stream replay, wired directly into the AgentServer SDKs for both the Responses protocol and the lower-level task primitives. This article is a deep dive into how it actually works under the hood, how to use it correctly, and — just as important — what it deliberately does not do for you.
Note: The APIs discussed here (
ResponsesServerOptions(resilient_background=True),@multi_turn_task,FoundryStateStore) are in public preview as of this writing. Preview features ship without an SLA and are subject to change — treat production adoption accordingly. (verify current preview status before shipping to production)
Table of Contents
- Why This Matters Now
- Background Execution vs. Resilient Execution
- The Resilient Work Model
- Work Shapes
- Architecture Walkthrough
- Implementation: Building a Crash-Resilient Agent
- Managing Durable State with FoundryStateStore
- Fencing Non-Idempotent Side Effects
- Stream Replay for Reconnecting Clients
- Real-World Scenario: A Long-Running Research Agent
- Production Considerations
- Security Considerations
- Performance, Scale, and Cost Considerations
- Common Mistakes and Pitfalls
- Alternatives and Trade-Offs
- Practical Recommendations
- Conclusion
- References
Why This Matters Now
Agentic workloads are structurally different from the request/response APIs most backend engineers have spent a career hardening. A single "turn" of an agent can span tool calls, retrieval steps, sub-agent delegation, and multi-minute generation — and increasingly, products want agents that keep working after the user has closed the tab, with results delivered later via webhook, push notification, or on next login.
That shift breaks two assumptions baked into most web architectures:
- The HTTP connection is not the unit of work anymore. A response can outlive the request that triggered it.
- The process is not the unit of durability anymore. A response can — and eventually will — outlive the container that started computing it.
Microsoft Foundry addresses the first problem with background execution (stored, pollable, streamable responses). It addresses the second, newer problem with resilient execution — the subject of this article. Foundry introduced this as part of its Foundry Agent Service preview wave (alongside autopilots, human-in-the-loop approval steps, and steerable agents), reflecting a broader industry realization: if agents are going to run for minutes or hours doing real work, "just retry the whole thing" is not an acceptable failure mode for production systems.
Background Execution vs. Resilient Execution
It's tempting to conflate "runs in the background" with "survives a crash." Foundry's documentation is unusually precise about drawing this line, and it's worth internalizing before writing any code:
| Capability | What it provides | What it doesn't provide |
|---|---|---|
| Background execution | Asynchronous work that clients can poll or reconnect to. | Recovery after the process that owns the work stops. |
| Resilient execution | Durable work identity, persisted input, process-loss detection, and handler reentry. | Automatic preservation of every intermediate application state or side effect. |
| Stream replay | Retained events that reconnecting clients can receive from a cursor. | A checkpoint of the agent's internal workflow state. |
Three separate concerns, three separate mechanisms. You can have background execution without resilience (the response just gets marked failed if the container dies). You can have resilience without stream replay (the handler recovers, but a client watching a live SSE stream loses the connection). Foundry lets you compose all three, but understanding that they're independent axes will save you a lot of confusion when something behaves "correctly but not how you expected."
For the Responses protocol, full crash recovery only applies to stored, background responses (store=true, background=true) where the server has explicitly opted in via resilient_background=True. Foreground responses — where the client is expected to be connected — are always marked failed on crash, because there's no meaningful way to resume a conversation whose client half already disappeared.
For the Invocations protocol (the lower-level task primitives), your application owns the request/response/status contract, and you opt in per-task using @task or @multi_turn_task decorators.
The Resilient Work Model
At the center of this system are two identities:
- Work identity — names the logical job or multi-turn conversation. This persists across the whole lifetime of the work, potentially across many process restarts.
- Input identity — names one input or turn within that work. A multi-turn conversation has one work identity and many input identities, one per turn.
Here's the sequence that makes crash recovery possible:
- Before a handler starts, the runtime persists its input and acquires a lease on the work record.
- While the handler runs, the runtime renews the lease periodically (a heartbeat, effectively).
- If the process crashes, gets OOM-killed, or is forcibly terminated, it stops renewing the lease.
- Once the lease expires, a later process (could be a new container instance, could be the same one after restart) can reclaim the lease and reinvoke the registered handler — with the same work identity and input identity.
This is a classic distributed-systems lease pattern (similar in spirit to how distributed job queues like Sidekiq or Celery implement visibility timeouts, or how Kubernetes leader election works), but Foundry wires it directly into the agent hosting layer so you don't have to stand up your own Redis-backed lock manager.
Crucially: recovery re-enters the handler from the beginning of the function. It is not deterministic replay. Your local variables, your in-memory call stack, any partial state you were holding in Python objects — all of that is gone. The only thing that survives is whatever you explicitly persisted. This is the single most important mental model shift: resilience gives you a guaranteed second chance to run the function; it does not give you a time machine back to where you left off. You have to build the "where you left off" part yourself, using checkpoints.
It's also worth explicitly distinguishing recovery from retry:
| Retry | Recovery | |
|---|---|---|
| Triggered by | A failure reported by the running handler | The process disappearing entirely |
| Consumes retry budget? | Yes | No |
| What continues | A fresh attempt | The same durable attempt |
A retry is what happens when your handler catches an exception and the framework decides to try again. Recovery is what happens when there's no handler left to catch anything — the whole process is dead — and a different process picks up the exact same durable attempt later.
Work Shapes
Foundry supports two shapes of resilient work, and choosing the right one depends on the lifetime and concurrency semantics of what you're building:
| Shape | Use when | Lifecycle |
|---|---|---|
| One-shot resilient work | A single operation must survive a process interruption. | One input → one result. Removed after reaching terminal state. |
| Multi-turn chain | A conversation or agent session accepts multiple inputs over time. | One work identity stays active across turns until deleted or retention expires. |
A subtlety worth calling out: calls that reuse the same one-shot work identity converge on the same logical operation rather than kicking off duplicate work. That's an important dedup guarantee if your client-side retry logic might fire the same request twice against a flaky network.
Multi-turn chains introduce an additional concept: steering. A newer turn can queue behind the currently active turn and signal the in-flight handler to wind down early, letting a conversation redirect ("actually, stop and do X instead") without spinning up a second concurrent handler that could race against the first. Conversations stay strictly sequential — steering doesn't fork execution.
Architecture Walkthrough
The diagram below shows the full crash-and-recovery lifecycle for a three-stage agent turn (analyze → generate → refine), which is the exact pattern used in Foundry's official resilient-streaming sample.
Walking through it:
- Process A acquires a lease on the work record and persists the input.
- It completes Stage 1 (analyze) and writes a checkpoint — a discrete, append-only record that says "this stage is done, here's the output."
- Before it can finish Stage 2, the container is killed.
- The lease is not renewed, so eventually it expires.
-
Process B (a fresh container instance) reclaims the lease and is reinvoked by the framework with
context.is_recovery == True, the same work identity, and the same original input. - Process B reads the checkpoint, sees Stage 1 is already done, skips it, and resumes at Stage 2 (generate), then completes Stage 3 (refine).
- The response reaches a terminal state and the runtime cleans up the resilient record.
Note what's not in this diagram: any orchestrator deciding "let's retry this." There's no retry budget consumed. From the runtime's perspective, this is one continuous durable attempt that happened to span two processes.
Implementation: Building a Crash-Resilient Agent
Let's build this for real, using the Responses protocol, which is the surface most teams building conversational or generative agents will use.
Step 1 — Opt in to resilient background execution
Resilience is off by default. You must explicitly enable it:
# main.py
from azure.ai.agentserver.responses import ResponsesAgentServerHost, ResponsesServerOptions
# resilient_background defaults to False. Without this, a crashed background
# response is simply marked "failed" — the framework does NOT reinvoke the handler.
options = ResponsesServerOptions(resilient_background=True)
app = ResponsesAgentServerHost(options=options)
This only matters for requests where the caller sets store=true and background=true. Foreground responses always fail hard on crash — there's no client connection left to resume toward.
Step 2 — Write a handler that checkpoints its own progress
from azure.ai.agentserver.responses import ResponseEventStream
STAGES = ["analyze", "generate", "refine"]
@app.response_handler
async def handler(request, context, cancellation_signal):
# Branch on whether this is a fresh call or a reinvocation after a crash.
if context.is_recovery:
# Seed the in-memory stream from the last durable snapshot instead
# of starting the turn over from scratch.
stream = ResponseEventStream.from_snapshot(context.persisted_response)
start_stage = len(stream.response.output) # completed, checkpointed stages
else:
stream = ResponseEventStream(response_id=context.response_id, request=request)
start_stage = 0
for i in range(start_stage, len(STAGES)):
stage = STAGES[i]
# Cooperative shutdown check — see "Handle graceful shutdown" below.
if context.is_shutting_down:
await context.exit_for_recovery()
return
result = await run_stage(stage, request.input, stream)
# One completed stage = one checkpointed output item. If the process
# dies before this line, the recovered attempt simply reruns the stage.
# If it dies after, the recovered attempt skips it.
stream.checkpoint_output_item(result)
return stream.finalize()
async def run_stage(stage_name: str, agent_input: str, stream: ResponseEventStream):
# Simplified: in production this would call a model, a tool, or both,
# streaming partial tokens through `stream` as they arrive.
...
The key design decision here: each stage boundary is a checkpoint boundary. This is deliberate. Foundry's guidance is explicit — prefer phase boundaries that checkpoint cleanly, where one output item corresponds to one completed phase. If a phase crashes before its checkpoint, it reruns entirely (so keep phases idempotent or cheap to redo); once checkpointed, the recovered attempt skips it permanently.
Step 3 — Test it locally with a simulated crash
Foundry's sample tooling lets you rehearse this exact failure mode on a laptop, which is a nice departure from having to fake OOM kills in a real cluster to validate your recovery logic:
# Initialize the official resilient-streaming sample
azd auth login
azd ai agent init -m https://github.com/microsoft-foundry/foundry-samples/blob/main/samples/python/hosted-agents/bring-your-own/responses/resilient-streaming/azure.yaml
# Force a crash right after stage 0 (analyze) checkpoints
SIMULATE_CRASH_AFTER_STAGE=0 azd ai agent run --no-client
In a second terminal, kick off a stored background response:
curl -sS -X POST http://localhost:8088/responses \
-H "Content-Type: application/json" \
-d '{"input": "renewable energy supply chains", "store": true, "background": true}'
The agent checkpoints stage 0 and then deliberately exits. Restart it:
azd ai agent run --no-client
The framework reinvokes the handler with context.is_recovery == True, restores context.persisted_response, skips the already-checkpointed analyze stage, and completes generate + refine. This local dev loop (files back the state store locally, so you exercise the same code path as production) is genuinely useful — it turns "did I implement recovery correctly?" from a production incident into a five-minute local test.
Step 4 — The lower-level task primitive alternative
If you're not using the Responses protocol — for example, you're building a custom multi-turn workflow with your own request/response contract — use @task or @multi_turn_task directly:
from azure.ai.agentserver.core.tasks import multi_turn_task, TaskContext
@multi_turn_task(name="research")
async def research(ctx: TaskContext[dict]) -> dict:
if ctx.entry_mode == "recovered":
# A previous lifetime didn't finish; resume from the durable watermark.
last_done = ctx.metadata.get("last_done_step", 0)
elif ctx.entry_mode == "resumed":
# A later turn in an existing multi-turn chain.
last_done = ctx.metadata.get("last_done_step", 0)
else: # "fresh"
last_done = 0
for step in range(last_done, TOTAL_STEPS):
await do_step(step, ctx)
ctx.metadata["last_done_step"] = step + 1
await ctx.metadata.flush()
return {"status": "complete"}
Declaring a @task or @multi_turn_task handler automatically enables the startup recovery scan — the runtime looks for abandoned leases on boot and reclaims them. If you register tasks lazily after the host has already started, you must force-enable this explicitly:
from azure.ai.agentserver.core.tasks import set_resilient_tasks_enabled
# Must be called at import time, before host lifespan startup, if tasks
# are registered dynamically rather than via decorators at module load.
set_resilient_tasks_enabled(True)
ctx.entry_mode gives you exactly three states to branch on: "fresh" (first turn, first attempt), "resumed" (a later turn in an existing chain, nothing crashed), and "recovered" (the framework is reinvoking after a previous lifetime failed to finish). Conflating "resumed" and "recovered" is a subtle bug source — they're semantically different even though the recovery code for both often looks similar.
Managing Durable State with FoundryStateStore
Checkpoints need somewhere durable to live, and that's FoundryStateStore — a server-backed, crash-surviving key-value store scoped by a caller-chosen name.
The critical design rule: keep progress markers and checkpoints as separate items. A progress marker is small and mutable (e.g., "we're on step 4"). A checkpoint is a completed, immutable result for a given step. Separating them lets a recovered attempt figure out both "what step comes next" and "what results already exist for prior steps" without racing itself.
from azure.ai.agentserver.core.storage import FoundryStateStore
async def save_step(task_id: str, step: int, result: dict) -> None:
store = await FoundryStateStore.get_or_create(f"workflows/{task_id}")
async with store:
progress = await store.get_item("progress")
# If a later attempt already advanced past this step, don't redo work
# or overwrite a newer checkpoint with a stale one.
if progress and int(progress.value["workflow_step"]) > step:
return
# Checkpoints are append-only: use a deterministic key so a recovered
# attempt can detect "this write already happened" instead of
# silently overwriting it with a possibly-different result.
checkpoint_key = f"checkpoints/{step}"
checkpoint = await store.get_item(checkpoint_key)
if checkpoint is None:
await store.create_item(checkpoint_key, result)
elif checkpoint.value != result:
# Same step, different result on recompute — this is a bug signal,
# not something to silently paper over.
raise RuntimeError("The checkpoint contains a different result.")
# Only advance the progress marker AFTER the checkpoint is durably written.
next_progress = {"workflow_step": step + 1}
if progress is None:
await store.create_item("progress", next_progress)
else:
# Optimistic concurrency: fails loudly if something else already moved forward.
await store.set_item("progress", next_progress, if_match=progress.etag)
Ordering matters here: write the checkpoint before you advance the progress marker. If the process dies between those two writes, the recovered attempt re-derives the same checkpoint key, finds it already exists, and safely continues — instead of silently redoing (and potentially double-executing) the step.
A few operational limits worth knowing up front:
- Item values cap at 1 MB of serialized JSON — this is an index/checkpoint-pointer store, not a blob store. Large artifacts belong in blob storage with a reference stored in the checkpoint.
- Store names can be 1–128 characters and support
/as a hierarchy separator — use it to encode workflow, session, or thread scope (workflows/{task_id},langGraphCheckpoints/{thread_id}). - Each item can carry up to 16 tags, useful for filtering when listing checkpoint history.
Backing a framework checkpointer
If you're already using LangGraph or Microsoft Agent Framework (MAF) for orchestration, you don't need to write custom recovery code at all — point the framework's own checkpointer at FoundryStateStore and its native recovery becomes crash-durable for free:
# LangGraph: one thread maps to one store, scoped per user
async def _store(thread_id: str) -> FoundryStateStore:
return await FoundryStateStore.get_or_create(
f"langGraphCheckpoints/{thread_id}", user_isolation=True
)
One sharp edge here: for the MAF adapter specifically, always set user_isolation=True. MAF's only native grouping concept is workflow_name — a definition name shared across every user running that workflow — so without isolation, get_latest() / list_checkpoints() will return other callers' checkpoints. This is exactly the kind of subtle multi-tenant data leak that's easy to miss in a demo and painful to discover in production.
Fencing Non-Idempotent Side Effects
Checkpointing solves "don't redo expensive computation." It does not, by itself, solve "don't send the same email twice." For any action a downstream system can't deduplicate on its own — sending a notification, charging a payment method, posting to an external webhook — you need an explicit watermark, written before the action and cleared after it commits:
# Fence before the non-idempotent side effect.
context.conversation_chain_metadata["email_sent"] = True
await context.conversation_chain_metadata.flush()
await email_service.send(...)
# A recovered handler checks this watermark before ever attempting to send again.
If the process crashes between the flush and the actual send, the recovered attempt sees email_sent = True, and — depending on your logic — either skips the send (accepting the small risk it never actually happened) or reconciles against the email provider's own idempotency key (safer, but requires the downstream system to support one). This is the same fundamental trade-off every distributed system faces at the boundary of "durable log" and "external side effect": you can get at-least-once or at-most-once semantics cheaply, but exactly-once requires either idempotent downstream APIs or a two-phase commit-style protocol. Foundry gives you the watermark primitive; the semantics you build on top are still your responsibility.
Stream Replay for Reconnecting Clients
Recovery for the agent is one problem. Recovery for the client watching a live stream is a related but separate one. Foundry's guidance here: use a separate stream identity per request or turn — never reuse a multi-turn work identity as the stream identity, because a completed stream closes even though the conversation itself continues across later turns.
A replayable stream retains its events and assigns cursors, so a reconnecting client can pick up from the last cursor it saw rather than from the beginning. When a stream is recovered, the very next response.in_progress event is a snapshot reset — the client should discard any locally accumulated partial output that isn't reflected in that snapshot and treat output indexes as scoped to the current snapshot, not as monotonically increasing across the whole recovery history. Getting this wrong is a classic source of duplicated or garbled tokens showing up in a chat UI after a reconnect.
Real-World Scenario: A Long-Running Research Agent
Consider a common enterprise pattern: a research agent that takes a topic, searches multiple internal and external sources, synthesizes findings, and produces a structured report — a process that can easily run for 3–10 minutes depending on source count and model latency.
Without resilience: The agent runs as a background task on a container that gets recycled mid-run during a routine deployment. The user's polling client eventually sees a failed status. Support gets a ticket. Someone manually re-triggers the job, which redoes all the search calls (burning API quota and cost) and, if the agent had already started drafting an email digest of findings to stakeholders, potentially sends it twice.
With Foundry's resilient background execution:
- The client submits
{"input": "...", "store": true, "background": true}. - The handler runs each research phase (search → synthesize → draft → notify) as a checkpointed stage.
- The deployment recycles the container mid-synthesis.
- The framework detects the abandoned lease, reclaims the work record, and reinvokes the handler with
context.is_recovery == True. - The handler restores
context.persisted_response, sees the search phase already checkpointed, and resumes at synthesize — no repeated search API calls, no wasted tokens. - The notify phase checks its
email_sentwatermark before dispatching, so even if synthesis had to partially rerun, the digest email fires exactly once. - The client, which had disconnected during the deploy, reconnects using the stream identity and cursor, and picks up the SSE stream exactly where it left off.
This is the difference between an agent that's a demo and one you can put a production SLA behind.
Production Considerations
-
This is preview. Treat
resilient_background,FoundryStateStore, and the task primitives as subject to breaking changes. Pin SDK versions and read release notes before upgrading. - Design for reentrant handlers from day one, even before you need crash recovery — it's a much cheaper habit to build early than to retrofit onto a codebase full of stateful, non-idempotent handler bodies.
- Retention and cleanup are partially your job. The runtime cleans up terminal or expired runtime records, but your application still owns cleanup of external checkpoints, sessions, and any application-level data you wrote to blob storage or a database.
- Cooperative shutdown matters. Distinguish graceful shutdown from failure explicitly:
if context.is_shutting_down:
# Leaves the response in_progress for a later lifetime to reclaim —
# do NOT write a terminal (failed/completed) state here.
await context.exit_for_recovery()
return
Writing a terminal state during a graceful shutdown window forecloses recovery entirely; the work will look "done" (or "failed") and nobody will ever pick it back up.
- Monitor lease reclaim rates. A high rate of recovered attempts is a leading indicator of infrastructure instability (aggressive autoscaling, OOM pressure, node churn) — instrument and alert on it, don't just silently rely on recovery to paper over it.
Security Considerations
-
State store scoping is your responsibility.
FoundryStateStorenames are caller-chosen strings — if you don't encode a stable, unpredictable scope (and setuser_isolation=Truewhere applicable), you risk cross-tenant or cross-user data bleed, as explicitly called out for the MAF adapter above. Audit every store name scheme for multi-tenant safety before shipping. - Persisted inputs are durable by design — that's the whole point, but it also means anything in the request payload (which may include PII or sensitive context) lives in the state store for the retention window. Apply the same data classification and retention policies you'd apply to any other durable store, and don't put secrets in the input payload; use a reference/secret-store pattern instead.
- Authorization for recovered handlers. A recovered handler reenters with the original request and metadata — make sure your authorization checks are re-evaluated on recovery rather than assumed from the first invocation, especially in multi-turn chains where a user's permissions could theoretically change mid-conversation.
- Watermarks that gate side effects are security-relevant, not just correctness-relevant. A watermark bug that causes a payment or an approval action to fire twice is a security and compliance incident, not just a UX glitch — treat fencing code with the same rigor as your payment/idempotency layer.
Performance, Scale, and Cost Considerations
- Lease renewal has a cost. Frequent heartbeats mean more calls against the state layer; check the configurable lease/renewal interval against your handler's realistic phase duration so you're not renewing far more often than necessary. (verify current default lease interval and configurability before capacity planning)
- Checkpoint granularity is a cost/latency trade-off. Very fine-grained checkpoints minimize rework on recovery but add write overhead to the state store on every phase transition. Very coarse checkpoints are cheaper per-write but mean more redone work (and redone token spend against the model) on recovery. Size your phases around natural, expensive boundaries — typically "one model call" or "one tool call" per phase, not smaller.
- Recovery reruns burn tokens too. If a phase crashes before its checkpoint commits, that phase's LLM calls are redone — factor this into cost modeling for high-crash-rate environments (e.g., aggressive spot instance usage).
-
Input size limits shape your architecture. The ~10 MiB per-input limit (raising
InputTooLargebefore any network call) and the 1 MB per state-store-item limit both push you toward externalizing large payloads (documents, images, large tool outputs) to blob storage and passing references — which is good practice anyway, but worth designing for up front rather than hitting the limit in production. - Multi-turn chains have a retention lifecycle. A work identity stays active "until the application deletes it or its retention period expires" — plan for both explicit cleanup on conversation end and a sane default retention so abandoned conversations don't accumulate indefinitely in the state store.
Common Mistakes and Pitfalls
-
Assuming resilience is on by default. It isn't, on either the Responses protocol (
resilient_backgrounddefaults toFalse) or lazily-registered tasks (set_resilient_tasks_enabledmust be called explicitly). Silent no-ops here are a classic "worked in the demo, failed in prod during the first real deploy" bug. - Treating recovery as deterministic replay. Recovery reenters the handler from the top; it does not restore local variables or the call stack. Code that assumes "if I got this far, X must already be true in memory" will break on the recovered path.
- Storing large artifacts directly in checkpoint metadata. Metadata is meant to be a small index (an ID, a phase name, a pointer) — not a checkpoint store itself. Cramming conversation history or tool output into it will hit size limits and defeat the purpose of keeping metadata cheap to read/write.
- Reusing a multi-turn work identity as a stream identity. This causes streams to close unexpectedly mid-conversation because the work identity's lifecycle and a single turn's stream lifecycle are different things.
-
Confusing
"resumed"and"recovered"entry modes. They require different handling in some designs (a"resumed"turn is a deliberate new input; a"recovered"turn is an involuntary do-over of an unfinished one) — collapsing the distinction can cause a fresh, intentional turn to be treated as a stale recovery, or vice versa. - Forgetting to fence non-idempotent side effects. Checkpointing computation is not the same as fencing side effects — an email, webhook call, or payment charge needs its own explicit watermark, not just reliance on "the stage was checkpointed."
-
Writing a terminal state during graceful shutdown. This looks like correct error handling but actually prevents recovery — use
exit_for_recovery()instead of letting the handler fall through to afailedstate.
Alternatives and Trade-Offs
It's worth being clear-eyed about when to reach for Foundry's built-in resilience versus rolling your own or using a general-purpose workflow engine:
- Durable workflow engines (Temporal, Azure Durable Functions, AWS Step Functions). These give you deterministic replay, fan-out/fan-in, and durable timers — capabilities Foundry's resilient tasks explicitly say they don't provide ("Design boundaries" in the docs). If your agent's control flow looks more like a complex DAG with branching, parallel fan-out, or long human-wait timers measured in days, a dedicated workflow engine composed alongside Foundry's model/tool layer is likely a better fit than forcing everything through the task primitive.
-
Framework-native checkpointing (LangGraph, MAF) without FoundryStateStore. You can absolutely keep using your framework's own checkpoint backend (Postgres, Redis, local disk). The value of routing it through
FoundryStateStoreis getting crash durability "for free" without operating a separate stateful service — but if you already run Postgres for other reasons, there's no obligation to add another storage dependency. - Simple "just retry the whole job" for cheap, fast, idempotent work. If your agent turn completes in a couple of seconds and has no non-idempotent side effects, the full resilience machinery is probably overkill — a naive rerun is correct and simpler. Reach for checkpointed resilience specifically when turns are long and side effects are non-idempotent and infrastructure churn (redeploys, autoscaling, spot) is a realistic threat to mid-flight work.
- Roll-your-own lease/heartbeat system. Entirely possible, and some teams already have one. Foundry's advantage is that it's integrated with the model/response lifecycle itself (checkpoints line up with response output items, streams replay against the same cursor system) rather than being a bolted-on general-purpose job queue.
Practical Recommendations
- Default new long-running agent handlers to being reentrant-safe, even if you're not enabling resilience yet — it costs little upfront and a lot to retrofit.
- Pick phase boundaries around expensive, natural units of work (one model call, one tool call, one external API round-trip) — not arbitrary code blocks.
- Treat any action your downstream system can't deduplicate as requiring an explicit watermark, full stop — don't rely on "it probably won't crash right there."
- Use the local crash-simulation workflow (
SIMULATE_CRASH_AFTER_STAGE,azd ai agent run --no-client) as a standard part of your test suite for any agent expected to run more than a few seconds, not just a one-off manual check. - Keep state store names encoded with a clear, stable hierarchy (
workflows/{id},{framework}Checkpoints/{thread_id}) from the start — retrofitting a naming scheme onto live production data is painful. - Instrument recovery events (count, phase-at-crash, time-to-reclaim) from day one; they're a proxy for infrastructure stability that you'll want long before you need to debug a specific incident.
Conclusion
Long-running, tool-using agents break a quiet assumption that most backend systems have relied on for years: that the process and the unit of work are roughly the same lifetime. Once an agent's turn can outlive its container, "just retry" stops being an adequate failure strategy, and teams either build a real lease/checkpoint/fencing system or ship agents that silently duplicate side effects during every redeploy.
Microsoft Foundry's resilient long-running agent primitives — durable work/input identities, lease-based crash detection, FoundryStateStore checkpointing, and stream replay — give you the hard distributed-systems plumbing so you can focus on the part that's actually specific to your product: deciding where the checkpoint boundaries are and which side effects need fencing. It's preview software, and the API surface will likely shift, but the underlying model — reentrant handlers, explicit checkpoints, watermarked side effects — is a durable pattern worth adopting even if you end up building it on different infrastructure later.
If you're shipping any agent whose turn can take longer than a few seconds, ask yourself right now: what happens if this container dies halfway through? If the honest answer is "we lose the work and possibly double-fire a side effect," this is the capability to go build against next.
Have you hit crash-recovery issues with long-running agents in production? Drop your war stories in the comments — I'm collecting real failure patterns for a future post in this series on production incident postmortems.
This is Day 1 of the Microsoft Foundry 100 Days / 100 Blogs series — one deep technical dive into a different corner of Microsoft Foundry every day for 100 days. Follow along for the next 99.
References
- Resilience for long-running Microsoft Foundry hosted agents (preview) — Microsoft Learn
- Manage state for long-running agents (preview) — Microsoft Learn
- Deploy a crash-resilient long-running agent (preview) — Microsoft Learn
- Recover long-running work after a crash (preview) — Microsoft Learn
- Long-running agent API reference (preview) — Microsoft Learn
- Microsoft Foundry sample: resilient-streaming (GitHub)
- Microsoft Foundry docs: What's new for August 2026

Top comments (0)