How to Stop AI Agents Looping and Running Up Costs: Step Limits, Budgets and Circuit Breakers
Why AI agents loop, retry and overspend, and the engineering controls that stop them: step limits, budgets, timeouts, circuit breakers and state checks.
Quick answer
Agents run away for predictable reasons: unclear finish conditions, tool errors treated as temporary, state they cannot see, and other agents bouncing work back. Stop it with controls outside the model: a maximum number of steps per run, a token and cost budget per run and per day, timeouts on every tool call, a retry budget with backoff, circuit breakers that trip on repeated identical actions or errors, validation of state after each action, and escalation to a person when any limit is hit. Then measure cost per completed task, because the cheapest model is not always the cheapest system.
Why agents cost more than you expect
A chatbot answers once. An agent plans, calls a tool, reads the result, reasons again and repeats, and each step usually re-sends the growing conversation. Anthropic's 2025 engineering write-up on its research system put numbers on it: agents typically used about four times more tokens than chat interactions, and its multi-agent system about fifteen times more. That can be worth it for valuable tasks; it becomes a problem when the extra steps are loops rather than progress.
Five ways agents run away
| Pattern | What happens | Typical cause |
|---|---|---|
| Infinite loop | Agent repeats the same plan or tool call | No finish condition; result not recognized as success |
| Repeated tool calls | Same query with tiny variations | Tool returns ambiguous or empty results |
| Retry storm | Many retries against a failing API, often across many runs at once | Errors treated as transient; no backoff or shared limit |
| Circular handoffs | Agents pass a task back and forth | Overlapping responsibilities in multi-agent setups |
| State errors | Agent redoes completed work or acts on stale data | State not persisted or not re-read after actions |
The controls, in order of importance
Every one of these lives in your orchestration code, not in the prompt. Instructions such as "do not repeat yourself" help, but they are not enforcement.
| Control | What it does | How to set it |
|---|---|---|
| Step limit | Caps reasoning/tool iterations per run | From the step distribution of successful runs, plus margin |
| Run budget | Caps tokens and spend per run | From cost per successful run; alert at 80% |
| Daily/tenant budget | Caps total spend per day, customer or feature | From expected volume; hard stop with alert |
| Timeouts | Bounds every tool and model call | Per tool, from normal latency |
| Retry budget | Limits retries with exponential backoff | Small number per call; shared across runs for the same dependency |
| Circuit breaker | Stops calling a failing tool; halts on repeated identical actions | Trip on error rate or N identical calls |
| State validation | Checks the world after each action | Re-read the record; confirm the expected change |
| Escalation | Hands the case to a person with context | Triggered by any limit or breaker |
Key takeaway
A run that hits a limit should end in a useful handoff (what was tried, what failed, what is left), not a silent failure and not another retry.
Design tools that do not invite loops
Many loops start with tools. A search tool that returns an empty list without explanation invites endless rephrasing; an API error that just says "failed" invites retries. Return explicit outcomes ("no orders found for this email; ask the customer for an order number"), mark errors as retryable or not, and make write tools idempotent so a retry cannot duplicate an action. See AI agent tool design.
Cost control beyond loops
Once runaway behaviour is contained, the normal levers apply. The difference for agents is to judge them by cost per completed task, not cost per call.
| Lever | Agent-specific note |
|---|---|
| Model routing | Use smaller models for classification and extraction steps, larger ones for planning; see LLM routing |
| Context reduction | Compact history and trim tool outputs; long contexts are re-sent every step |
| Prompt caching | Keep instructions and tool definitions stable at the start of the context so providers can cache them |
| Fewer tool calls | Task-shaped tools that return what the next step needs in one call |
| Task decomposition | Deterministic code for fixed steps; the model only where judgement is needed |
| Batching | Use batch APIs for non-urgent background work where providers offer discounts |
Agents costing more than they should?
ZSpace Labs audits agent traces for loops, retries and waste, then adds budgets, breakers and routing that cut cost per completed task. See AI automation services.
Why the cheapest model is not always the cheapest system
An illustrative comparison: a small model costs a fraction per token, but on a multi-step task it takes more steps, retries failed tool calls and completes fewer cases correctly, so people handle the rest. A larger model costs more per token, finishes in fewer steps and completes more cases. Cost per completed task, including human handling of failures, can favour the larger model. Measure both on the same test set before deciding; see small language models for when the smaller option wins.
Monitor for runaway behaviour
- Steps per run and tokens per run, by task type
- Repeated identical tool calls within a run
- Retry counts and error rates per tool
- Runs ending at a limit (and why)
- Cost per completed task and daily spend vs budget
- Escalations created by limits
Conclusion
Loops and overspending are engineering problems with engineering solutions. Put hard limits around every run, make tools give clear outcomes, contain failing dependencies with breakers, validate state after actions and escalate with context when limits are hit. Then optimize cost per completed task. For agents that must survive crashes and restarts mid-task, see durable execution for AI agents; for tracing, see AI agent observability and LLM cost optimization.
Test limits and breakers before go-live in an isolated environment; see AI agent sandbox.
Common questions.
Common causes are tool errors the agent treats as transient, ambiguous goals with no clear finish condition, state the agent cannot see changing (so it repeats a step), and two agents handing a task back and forth. The model is not broken; the system lacks stopping rules.