Skip to content
AI & Automation5 min read

How to Stop AI Agents Looping and Running Up Costs: Step Limits, Budgets and Circuit Breakers

Why AI agents loop, retry and overspend, and the engineering controls that stop them: step limits, budgets, timeouts, circuit breakers and state checks.

01

Quick answer

Agents run away for predictable reasons: unclear finish conditions, tool errors treated as temporary, state they cannot see, and other agents bouncing work back. Stop it with controls outside the model: a maximum number of steps per run, a token and cost budget per run and per day, timeouts on every tool call, a retry budget with backoff, circuit breakers that trip on repeated identical actions or errors, validation of state after each action, and escalation to a person when any limit is hit. Then measure cost per completed task, because the cheapest model is not always the cheapest system.

02

Why agents cost more than you expect

A chatbot answers once. An agent plans, calls a tool, reads the result, reasons again and repeats, and each step usually re-sends the growing conversation. Anthropic's 2025 engineering write-up on its research system put numbers on it: agents typically used about four times more tokens than chat interactions, and its multi-agent system about fifteen times more. That can be worth it for valuable tasks; it becomes a problem when the extra steps are loops rather than progress.

03

Five ways agents run away

PatternWhat happensTypical cause
Infinite loopAgent repeats the same plan or tool callNo finish condition; result not recognized as success
Repeated tool callsSame query with tiny variationsTool returns ambiguous or empty results
Retry stormMany retries against a failing API, often across many runs at onceErrors treated as transient; no backoff or shared limit
Circular handoffsAgents pass a task back and forthOverlapping responsibilities in multi-agent setups
State errorsAgent redoes completed work or acts on stale dataState not persisted or not re-read after actions
04

The controls, in order of importance

Every one of these lives in your orchestration code, not in the prompt. Instructions such as "do not repeat yourself" help, but they are not enforcement.

ControlWhat it doesHow to set it
Step limitCaps reasoning/tool iterations per runFrom the step distribution of successful runs, plus margin
Run budgetCaps tokens and spend per runFrom cost per successful run; alert at 80%
Daily/tenant budgetCaps total spend per day, customer or featureFrom expected volume; hard stop with alert
TimeoutsBounds every tool and model callPer tool, from normal latency
Retry budgetLimits retries with exponential backoffSmall number per call; shared across runs for the same dependency
Circuit breakerStops calling a failing tool; halts on repeated identical actionsTrip on error rate or N identical calls
State validationChecks the world after each actionRe-read the record; confirm the expected change
EscalationHands the case to a person with contextTriggered by any limit or breaker

Key takeaway

A run that hits a limit should end in a useful handoff (what was tried, what failed, what is left), not a silent failure and not another retry.

05

Design tools that do not invite loops

Many loops start with tools. A search tool that returns an empty list without explanation invites endless rephrasing; an API error that just says "failed" invites retries. Return explicit outcomes ("no orders found for this email; ask the customer for an order number"), mark errors as retryable or not, and make write tools idempotent so a retry cannot duplicate an action. See AI agent tool design.

06

Cost control beyond loops

Once runaway behaviour is contained, the normal levers apply. The difference for agents is to judge them by cost per completed task, not cost per call.

LeverAgent-specific note
Model routingUse smaller models for classification and extraction steps, larger ones for planning; see LLM routing
Context reductionCompact history and trim tool outputs; long contexts are re-sent every step
Prompt cachingKeep instructions and tool definitions stable at the start of the context so providers can cache them
Fewer tool callsTask-shaped tools that return what the next step needs in one call
Task decompositionDeterministic code for fixed steps; the model only where judgement is needed
BatchingUse batch APIs for non-urgent background work where providers offer discounts

Agents costing more than they should?

ZSpace Labs audits agent traces for loops, retries and waste, then adds budgets, breakers and routing that cut cost per completed task. See AI automation services.

Start a Project
07

Why the cheapest model is not always the cheapest system

An illustrative comparison: a small model costs a fraction per token, but on a multi-step task it takes more steps, retries failed tool calls and completes fewer cases correctly, so people handle the rest. A larger model costs more per token, finishes in fewer steps and completes more cases. Cost per completed task, including human handling of failures, can favour the larger model. Measure both on the same test set before deciding; see small language models for when the smaller option wins.

08

Monitor for runaway behaviour

  • Steps per run and tokens per run, by task type
  • Repeated identical tool calls within a run
  • Retry counts and error rates per tool
  • Runs ending at a limit (and why)
  • Cost per completed task and daily spend vs budget
  • Escalations created by limits
09

Conclusion

Loops and overspending are engineering problems with engineering solutions. Put hard limits around every run, make tools give clear outcomes, contain failing dependencies with breakers, validate state after actions and escalate with context when limits are hit. Then optimize cost per completed task. For agents that must survive crashes and restarts mid-task, see durable execution for AI agents; for tracing, see AI agent observability and LLM cost optimization.

Test limits and breakers before go-live in an isolated environment; see AI agent sandbox.

FAQ

Common questions.

Common causes are tool errors the agent treats as transient, ambiguous goals with no clear finish condition, state the agent cannot see changing (so it repeats a step), and two agents handing a task back and forth. The model is not broken; the system lacks stopping rules.

Get in touch

Have a project in mind?

Whether you're building a new digital product, improving an existing website, or looking to automate part of your business — let's talk.