Skip to content
AI & Automation5 min read

AI Agent Incident Response: What to Do When an Agent Gets It Wrong

An incident response plan for AI agents: detection, kill switches, containment, reversing actions, root cause analysis and preventing repeats.

01

Quick answer

Prepare before anything goes wrong. Every production agent needs a way to detect problems (outcome monitoring and alerts, not just uptime), a tested kill switch that stops it without a deployment, a log detailed enough to list and reverse its actions, a fallback to the manual process, and a runbook naming who decides what. When an incident happens: contain first, assess which actions and customers were affected, reverse or correct them, find the root cause from traces (context, tools, permissions, model change or manipulation), fix the control that failed, re-run evaluations and update the runbook.

02

What counts as an agent incident

TypeExampleTypical cause
Wrong actionRefunds issued twice; records overwrittenNon-idempotent tool, retry loop, bad data
Wrong statementAgent promises a policy that does not existUngrounded answer, outdated source
Data exposureAgent includes another customer's detailsRetrieval without tenant scoping
ManipulationAgent follows instructions hidden in an emailPrompt injection (goal hijack)
Runaway costToken spend jumps tenfold overnightLoop, oversized context, traffic spike
Silent degradationAccuracy drops after a model updateUnevaluated change
03

Before an incident: the minimum preparation

  • Detection: outcome metrics (error rate, escalation rate, refunds, complaints), cost alerts and anomaly alerts on action volume; see AI agent observability
  • Kill switch: a configuration flag or credential revocation that stops the agent in minutes, tested regularly
  • Degraded modes: read-only mode, approval-required mode, or routing all cases to people
  • Action log: every write with enough detail to identify and reverse it
  • Compensating actions: scripted reversals for common writes (void credit, restore record)
  • Runbook: who is on call, who decides to stop the agent, who talks to customers
  • Version records: model, prompt, tool and policy versions per run

Key takeaway

If stopping an agent requires a code deployment, you do not have a kill switch. Test the switch the same way you test backups.

04

During an incident: contain, then assess

Contain first. Switch the agent to a degraded mode or stop it. Do not wait to understand the cause; agents can repeat a mistake hundreds of times while you investigate. If manipulation is suspected, revoke the agent's credentials and those of any connected tools.

Assess impact. Use the action log to list everything the agent did in the affected period: which records, customers, amounts and messages. Narrow by the run IDs and tools involved. This list drives remediation and any notifications.

Reverse and remediate. Run compensating actions where possible, correct records, and contact affected customers with a clear explanation and fix. Some actions (messages sent, information disclosed) cannot be undone; correct them and record what was done.

05

Finding the root cause

Agent incidents rarely have a single bug. Read the traces of affected runs and work through the layers:

LayerQuestion
ContextDid the agent have wrong, stale, missing or injected information?
ToolsDid a tool return an error, partial data or allow an unsafe action?
PermissionsCould the agent do more than its task required?
ApprovalsDid an approval step exist, and was it meaningful?
ChangesDid a model, prompt, tool or data change precede the incident?
InputWas there adversarial content (emails, documents, web pages)?
06

After an incident: fix the control, not just the case

Patching the prompt for the specific case is rarely enough. Fix the control that let the mistake through: narrow a tool, add a validation, scope retrieval per tenant, add a threshold for approval, make a write idempotent, add a cost limit. Add the incident to your evaluation set so it is tested on every future change, re-run evaluations and record the change. The OWASP Top 10 for Agentic Applications is a useful checklist for which class of control failed; see the OWASP agentic guide.

Need agents you can stop, audit and repair?

ZSpace Labs builds kill switches, degraded modes, action logs and runbooks into production agents. See AI automation services.

Start a Project
07

Communicating with customers and stakeholders

Be direct: what happened, who was affected, what you have done and what changes. Do not blame "the AI"; customers and regulators hold the business responsible, as covered in who is responsible when an AI agent makes a mistake. Where personal data is involved, follow your data breach procedures and legal notification duties.

08

Practise

Run a short exercise each quarter for important agents: simulate a manipulation or a runaway loop, trigger the kill switch, list the affected actions from logs and reverse a sample. The first exercise usually reveals a missing log field or a switch nobody can find. See AI red teaming for adversarial testing.

09

Conclusion

Agents will make mistakes; the difference between a minor incident and a serious one is preparation. Detect outcomes, keep a tested kill switch and degraded modes, log every action so it can be reversed, contain before investigating, and fix the control that failed. Then practise, so the plan works when it matters.

The detail behind two steps of this plan is covered separately: how to safely undo autonomous actions and how to build an audit trail for agent actions.

FAQ

Common questions.

Any event where an agent causes or nearly causes harm: wrong actions on records or money, incorrect statements to customers, data exposure, runaway costs, or behaviour suggesting manipulation such as prompt injection.

Get in touch

Have a project in mind?

Whether you're building a new digital product, improving an existing website, or looking to automate part of your business — let's talk.