Durable Execution for AI Agents: How to Make Long-Running Agents Survive Failures
What durable execution is, why long-running AI agents need it, and how to persist state, checkpoint steps, retry safely and pause for human approval.
Quick answer
Durable execution records an agent's progress step by step so it can resume exactly where it stopped after a crash, deployment, timeout or rate limit, instead of starting over. Long-running agents need it because they make many slow, failure-prone calls, change external systems and often wait for people. In practice: run the agent loop inside a workflow engine or a framework with checkpointing, treat each model and tool call as a recorded step, make every external action idempotent, model human approvals as waits, and set timeouts and retry policies per step. Short interactive agents can skip it.
Why agents fail differently from web requests
A web request lasts milliseconds; if it fails, the user retries. An agent task can run for minutes or days: research, reconciliation, onboarding, multi-system investigations. Along the way it hits rate-limited models, flaky APIs, deployments that restart servers, and pauses while a person approves something. Anthropic's engineering team described this directly when writing about its multi-agent research system: minor failures can be catastrophic for agents, so the system had to resume from where the agent was rather than restart from the beginning.
Restarting from scratch is not just slow. It re-spends tokens, re-runs tool calls that already changed external systems and may produce a different plan the second time.
How durable execution works
| Concept | What it means for an agent |
|---|---|
| Recorded steps | Each model call and tool call is persisted with its result |
| Replay and resume | After a failure, the run continues from the last completed step using recorded results |
| Retries per step | Failed calls retry with backoff according to policy, without repeating earlier steps |
| Durable timers and waits | The run can sleep for hours or wait for a human signal without holding a server |
| Deterministic orchestration | The orchestration code is replayable; non-deterministic work (model calls, APIs) happens in recorded steps |
Worth noting
Temporal and OpenAI announced an integration in July 2025 that runs OpenAI Agents SDK agents with durable execution, so rate-limited model calls resume when capacity recovers and crashed runs continue where they stopped. Similar patterns exist in other workflow engines and agent frameworks.
Implementation patterns
- Wrap the agent loop in a workflow: each iteration's model call and each tool call becomes a recorded activity
- Make side effects idempotent: pass an idempotency key derived from the run and step to payment, email and record-changing tools
- Model approvals as signals: the workflow waits for an approve/reject event instead of polling or blocking a thread
- Set per-step timeouts and retry policies: different for model calls, internal APIs and third-party APIs
- Persist agent state explicitly: plan, notes and decisions stored as data the run can reload, not only in the model context
- Version workflows: in-flight runs must keep working when you deploy a new version of the agent
- Combine with run limits: durability keeps runs alive; budgets and breakers stop the ones that should not continue (see runaway AI agents)
When you need it and when you do not
| Agent type | Durable execution? | Why |
|---|---|---|
| Chat assistant answering in seconds | No | Retry from the start is fine |
| Agent that only reads and summarizes | Usually no | No side effects; restart is cheap |
| Multi-step agent that changes records or sends messages | Yes | Avoid duplicate side effects after failures |
| Background agents running minutes to hours | Yes | Crashes, deploys and rate limits are likely mid-run |
| Agents waiting for human approval | Yes | Waits can last hours or days |
Building agents that run for more than a few seconds?
ZSpace Labs builds long-running agents on durable workflow engines with idempotent tools, approval waits and run budgets. See AI automation services.
Observability gets easier
A side benefit: the recorded history of each run is an audit trail. You can see every step, input, output, retry and approval, replay a failed run in a test environment and answer "what did the agent do?" precisely. Connect it to your tracing so model-level detail sits alongside workflow history; see AI agent observability.
Conclusion
Long-running agents fail mid-task; durable execution makes that survivable. Record steps, resume instead of restarting, keep side effects idempotent, model approvals as waits and pair durability with hard limits. For coordinating several agents, see AI agent orchestration; for the broader list of production pitfalls, why AI agents fail in production.
Common questions.
A way of running code so that its progress is recorded step by step and it can resume exactly where it stopped after a crash, deployment, timeout or rate limit, without repeating completed steps. Workflow engines such as Temporal provide it.