Skip to content
AI & Automation4 min read

Durable Execution for AI Agents: How to Make Long-Running Agents Survive Failures

What durable execution is, why long-running AI agents need it, and how to persist state, checkpoint steps, retry safely and pause for human approval.

01

Quick answer

Durable execution records an agent's progress step by step so it can resume exactly where it stopped after a crash, deployment, timeout or rate limit, instead of starting over. Long-running agents need it because they make many slow, failure-prone calls, change external systems and often wait for people. In practice: run the agent loop inside a workflow engine or a framework with checkpointing, treat each model and tool call as a recorded step, make every external action idempotent, model human approvals as waits, and set timeouts and retry policies per step. Short interactive agents can skip it.

02

Why agents fail differently from web requests

A web request lasts milliseconds; if it fails, the user retries. An agent task can run for minutes or days: research, reconciliation, onboarding, multi-system investigations. Along the way it hits rate-limited models, flaky APIs, deployments that restart servers, and pauses while a person approves something. Anthropic's engineering team described this directly when writing about its multi-agent research system: minor failures can be catastrophic for agents, so the system had to resume from where the agent was rather than restart from the beginning.

Restarting from scratch is not just slow. It re-spends tokens, re-runs tool calls that already changed external systems and may produce a different plan the second time.

03

How durable execution works

ConceptWhat it means for an agent
Recorded stepsEach model call and tool call is persisted with its result
Replay and resumeAfter a failure, the run continues from the last completed step using recorded results
Retries per stepFailed calls retry with backoff according to policy, without repeating earlier steps
Durable timers and waitsThe run can sleep for hours or wait for a human signal without holding a server
Deterministic orchestrationThe orchestration code is replayable; non-deterministic work (model calls, APIs) happens in recorded steps

Worth noting

Temporal and OpenAI announced an integration in July 2025 that runs OpenAI Agents SDK agents with durable execution, so rate-limited model calls resume when capacity recovers and crashed runs continue where they stopped. Similar patterns exist in other workflow engines and agent frameworks.

04

Implementation patterns

  • Wrap the agent loop in a workflow: each iteration's model call and each tool call becomes a recorded activity
  • Make side effects idempotent: pass an idempotency key derived from the run and step to payment, email and record-changing tools
  • Model approvals as signals: the workflow waits for an approve/reject event instead of polling or blocking a thread
  • Set per-step timeouts and retry policies: different for model calls, internal APIs and third-party APIs
  • Persist agent state explicitly: plan, notes and decisions stored as data the run can reload, not only in the model context
  • Version workflows: in-flight runs must keep working when you deploy a new version of the agent
  • Combine with run limits: durability keeps runs alive; budgets and breakers stop the ones that should not continue (see runaway AI agents)
05

When you need it and when you do not

Agent typeDurable execution?Why
Chat assistant answering in secondsNoRetry from the start is fine
Agent that only reads and summarizesUsually noNo side effects; restart is cheap
Multi-step agent that changes records or sends messagesYesAvoid duplicate side effects after failures
Background agents running minutes to hoursYesCrashes, deploys and rate limits are likely mid-run
Agents waiting for human approvalYesWaits can last hours or days

Building agents that run for more than a few seconds?

ZSpace Labs builds long-running agents on durable workflow engines with idempotent tools, approval waits and run budgets. See AI automation services.

Start a Project
06

Observability gets easier

A side benefit: the recorded history of each run is an audit trail. You can see every step, input, output, retry and approval, replay a failed run in a test environment and answer "what did the agent do?" precisely. Connect it to your tracing so model-level detail sits alongside workflow history; see AI agent observability.

07

Conclusion

Long-running agents fail mid-task; durable execution makes that survivable. Record steps, resume instead of restarting, keep side effects idempotent, model approvals as waits and pair durability with hard limits. For coordinating several agents, see AI agent orchestration; for the broader list of production pitfalls, why AI agents fail in production.

FAQ

Common questions.

A way of running code so that its progress is recorded step by step and it can resume exactly where it stopped after a crash, deployment, timeout or rate limit, without repeating completed steps. Workflow engines such as Temporal provide it.

Get in touch

Have a project in mind?

Whether you're building a new digital product, improving an existing website, or looking to automate part of your business — let's talk.