AI Agent Observability: How to Monitor and Debug Agentic Systems
How to monitor and debug AI agents: traces and spans for model and tool calls, token and cost tracking, latency, errors, evaluation scores, OpenTelemetry GenAI conventions, privacy and incident investigation.
Quick answer
AI agent observability means recording every run as a trace: a root span for the run and child spans for each model call, tool call, retrieval and approval, carrying inputs, outputs, tool arguments, token counts, cost, latency, errors and version information. Aggregate these into metrics and dashboards, attach evaluation scores and user feedback, alert on errors, cost spikes, loops and quality drops, and protect the data, because traces contain customer information. Use the OpenTelemetry GenAI conventions to keep telemetry portable.
Where This Fits
Observability feeds agent evaluation and cost optimization, and records the decisions made by guardrails. General ecommerce observability practices are in ecommerce observability.
Model-level quality, drift and cost monitoring is covered in AI model monitoring.
Observability for non-agent LLM applications such as assistants and RAG systems is covered in LLM observability and tracing.
Why Agents Need Different Monitoring
A traditional service fails loudly: errors, timeouts, crashes. Agents often fail quietly: a plausible answer from the wrong source, an extra refund, a loop that burns tokens before giving up. To find these failures you need to see the content of each step and judge quality, not just measure availability.
Anatomy of an Agent Trace
| Span type | Key attributes |
|---|---|
| Agent run (root) | Run ID, user or tenant, task type, agent and prompt versions, outcome |
| Model call | Provider, model, input and output tokens, latency, finish reason |
| Tool call | Tool name, arguments, result or error, latency, policy decision |
| Retrieval | Query, sources returned, scores, filters applied |
| Approval | Reviewer, decision, edits, wait time |
OpenTelemetry GenAI Conventions
The OpenTelemetry semantic conventions for generative AI define standard operation names (such as chat, invoke_agent and execute_tool) and attributes for providers, models and token usage. They are still evolving, but adopting them keeps telemetry portable between tools and lets traces from model SDKs, frameworks and MCP servers line up. The latest MCP specification also documents trace context propagation through its metadata fields.
Metrics and Dashboards
- Runs per task type, success and escalation rates
- Tokens and cost per run, per task type and per customer
- Latency per step and end to end (p50 and p95)
- Tool error rates and policy denials
- Runs hitting step or cost limits
- Online evaluation scores and user feedback
- Model and prompt version distribution during rollouts
Agents in production but no idea what they are doing?
ZSpace Labs instruments agents with end-to-end tracing, cost tracking and quality alerts so issues are found before customers report them.
Alerts That Matter
Alert on changes that affect customers or cost: rising escalations or errors, sudden token or cost spikes, latency beyond budget, loops hitting limits, surges in policy denials (possible abuse or a broken tool) and drops in evaluation scores after a release or a provider model update. Tie alerts to owners and runbooks.
Debugging and Incident Investigation
When something goes wrong, the trace should answer: what did the agent receive, what did it retrieve, which tools did it call with which arguments, what came back, what policies decided and what it output. Replay the run against a fixed version to reproduce it. After the fix, add the case to the evaluation set. For incidents involving customers, record who was affected and what was changed so it can be corrected.
Privacy and Security of Telemetry
Traces hold prompts, documents and customer data. Redact or hash sensitive fields where possible, restrict access by role, encrypt at rest, set retention limits and exclude secrets entirely. If you use a third-party observability service, check where data is stored and how it is processed, and cover it in your privacy documentation.
Tooling Options
You can send OpenTelemetry traces to your existing observability platform, use LLM-specific observability tools that add prompt views, datasets and evaluations, or use your model provider's tracing features. Many teams combine a general platform for infrastructure with an LLM tool for content and quality. Choose based on data residency, cost and how well traces connect to evaluation.
How to Instrument an Agent Step by Step
- 1. Create a run ID and propagate it through every service and tool
- 2. Instrument model calls with model, tokens, latency and finish reason
- 3. Instrument tool calls with arguments, results, errors and policy decisions
- 4. Record versions of prompts, tools, models and retrieval indexes
- 5. Add redaction and retention rules
- 6. Build dashboards for success, cost, latency and errors
- 7. Add alerts with owners and runbooks
- 8. Connect traces to evaluation and feedback
What to Log and What Not To
| Data | Log it? | Notes |
|---|---|---|
| Model, version, tokens, latency, cost | Always | Core operational data |
| Tool name, arguments, result status | Always | Redact sensitive argument values |
| Prompts and model outputs | Usually | Redact personal data; restrict access; set retention |
| Retrieved document IDs and scores | Always | IDs rather than full text where possible |
| Full retrieved text | Sometimes | Useful for debugging; high privacy cost |
| Secrets, tokens, credentials | Never | Strip before logging |
| User and tenant identifiers | Always | Needed for access control and audits |
Example Trace
A simplified trace shows how one run breaks down. Even this level of detail answers most debugging questions: where time went, which tool failed and how much the run cost.
invoke_agent support_agent v12 run=run_91c total=6.8s cost=$0.031
├─ chat model=small-model in=1,820 out=96 0.9s -> tool_call get_order
├─ execute_tool get_order(ORD-104233) 0.3s ok
├─ chat model=small-model in=2,410 out=88 0.8s -> tool_call list_payments
├─ execute_tool list_payments(ORD-104233) 0.4s ok (2 payments)
├─ chat model=large-model in=2,950 out=210 2.6s -> tool_call refund_payment
├─ policy refund_payment amount=59.90 ALLOW (under auto limit)
├─ execute_tool refund_payment(pi_b, 59.90) 1.1s ok
└─ chat model=small-model in=3,300 out=140 0.7s -> final replyWorked Example
An illustrative scenario, not a client case: costs for a research agent double overnight with no deploy. Traces show runs now average 18 steps instead of 7, because a search tool started returning errors that the agent retries repeatedly. The team adds a loop detector, a retry cap on that tool and an alert on step-count spikes, and asks the tool's owner to fix the error.
Common Mistakes
- Logging only final answers
- No cost attribution per run or customer
- Storing full prompts with personal data indefinitely
- No version information in traces
- Alerts on uptime only
Want to see exactly why an agent did what it did?
Talk to ZSpace Labs about AI observability and agent operations and telemetry and backend integration.
Conclusion
You cannot run agents responsibly without seeing inside them. Trace every step, track cost and quality, alert on what matters and protect the data. Related: evaluation, LLM cost optimization and guardrails.
Common questions
The ability to see what an agent did and why: traces of every model call, tool call, retrieval and approval in a run, with tokens, cost, latency, errors and quality scores, so problems can be detected and debugged.