Skip to content
AI & Automation

LLM Observability: How to Monitor AI Application Quality and Performance

How to observe LLM applications in production: traces and spans across retrieval, model and tool calls, correlation IDs, token usage, cost, latency, errors, retrieval quality, output quality, user feedback and debugging multi-step workflows.

Quick answer

LLM observability means capturing enough data to explain any production response: a trace for every request with spans for retrieval, prompt assembly, model calls, tool calls and validation, each carrying inputs, outputs, versions, tokens, latency and errors, joined by correlation IDs. On top of traces, track metrics for latency, error rates, cost and cache use, plus quality signals from sampled evaluation and user feedback. Protect the data, because traces often contain sensitive content.

Where This Fits

This article covers observability and tracing for LLM applications such as assistants, RAG systems and AI features. Agent-specific concerns are in AI agent observability, model quality and drift in AI model monitoring, and the wider operating practice in LLMOps.

Why Traditional Monitoring Falls Short

A conventional dashboard might show a support assistant with a 99.9% success rate and healthy latency while it confidently gives customers an outdated refund policy. Nothing failed technically: retrieval returned an old document, the model used it faithfully, and the response was well formed. Only by seeing the retrieved context, prompt version and output together can you find the cause.

LLM applications also have costs that vary per request, behaviour that changes when a provider updates a model, and multi-step flows where one weak step spoils the result. Observability for these systems therefore combines classic telemetry with content, versions and quality.

Anatomy of an LLM Trace

A trace represents one user request. Inside it, spans represent steps, nested where one step calls another. Every span carries timing, status and attributes; LLM spans add model, prompt version, token counts and, where permitted, input and output content.

Example: simplified trace for one request (illustrative)
trace_id: 7f3c...  user: u_482 (hashed)  feature: support_answer  release: 2026.10.2
├─ span api.request              1,842 ms  status=ok
│  ├─ span retrieval.search         146 ms  index=kb_v14 top_k=8 hits=8
│  ├─ span prompt.build               3 ms  prompt=support_answer@v12
│  ├─ span llm.chat               1,512 ms  model=<provider/model>  in=3,904 tok  out=412 tok
│  │                                       cost=$0.0071  finish=stop
│  ├─ span output.validate           11 ms  schema=ok  citations=3/3 found
│  └─ span response.stream          165 ms
feedback: thumbs_down  reason="policy outdated"
Each span answers one debugging question: what was retrieved, what was sent, what came back, what was checked.

Standardizing With OpenTelemetry

The OpenTelemetry semantic conventions for generative AI define standard attribute names for model calls, such as the provider, requested and response model, token usage and operation type. Instrumenting against these conventions lets you send the same data to different backends and combine AI spans with the rest of your distributed tracing.

The conventions are still evolving, so pin library versions and expect some attribute changes over time. Many frameworks and SDKs offer automatic instrumentation, but check what content they capture by default before enabling them in production.

What to Measure

CategoryMetricsTypical alert
LatencyTime to first token, total time, p50 and p95 per featurep95 above budget
ErrorsProvider errors, timeouts, rate limits, validation failuresError rate above baseline
Usage and costInput and output tokens, cost per request and per feature, cache hit rateDaily spend or cost per request spike
RetrievalHit counts, empty results, score distributions, source freshnessRising empty-result rate
QualitySampled judge scores, feedback ratio, edits, retries, escalationsScore drop after a release
SafetyPolicy flags, injection detections, blocked tool callsAny spike or new pattern

Can't tell why your AI feature gives bad answers?

ZSpace Labs instruments LLM applications end to end so every answer can be explained. See our AI engineering services.

Start a Project

Debugging Multi-Step Workflows With Traces

Good traces turn vague complaints into specific causes. A practical routine: find the trace from the user's report or feedback, check retrieval first (were the right documents found, and were they current?), then the assembled prompt (was the context included and the right prompt version used?), then the model output (did it ignore or misread the context?), then validation and post-processing.

For workflows that span services, queues or external APIs, propagate the trace context through every hop, including message headers on queues, so asynchronous steps join the same trace. Where a hop cannot carry trace context, log a correlation ID that links the records. Patterns that recur, such as a document source that is often stale, become fixes and new evaluation cases.

Quality Signals Without Labels

Most production outputs never receive a ground-truth label, so combine several signals. Explicit feedback is valuable but sparse and skewed toward strong reactions. Implicit signals include users editing drafts heavily, retrying, abandoning or escalating to a person. Sampled evaluation scores a small fraction of traces with the same judges used before release, giving a consistent trend. Watch all three by release version so you can tell whether a change helped. Feedback design is covered in AI feedback UX.

Protecting Trace Data

Traces can become the largest store of sensitive data in an AI system: user questions, retrieved documents and model outputs. Redact or hash identifiers, avoid capturing secrets and payment data, restrict trace access by role, set short retention for full content and keep aggregate metrics longer. Check what your instrumentation captures by default. Privacy guidance is in AI data privacy and leakage risks in AI data leakage.

Choosing Tooling

Teams typically choose between dedicated LLM observability platforms, which provide trace views of prompts and outputs, evaluation and feedback features, and general observability backends that receive OpenTelemetry data alongside the rest of the system. Examples of LLM-focused tools include Langfuse, LangSmith and MLflow tracing. Consider data residency and self-hosting options, cost at your trace volume and how well the tool fits your evaluation workflow.

Advantages and Limitations

Observability shortens incident investigations from days to minutes, makes cost visible per feature and shows whether releases improve quality. Its limits: storing content is expensive and sensitive, sampled quality scores are estimates and instrumentation adds some overhead and maintenance as conventions change.

How to Set Up Observability Step by Step

  • 1. Instrument every model, retrieval and tool call as spans with OpenTelemetry
  • 2. Attach versions (release, prompt, model, index) to every trace
  • 3. Propagate trace context across services and queues
  • 4. Record tokens and cost and build per-feature dashboards
  • 5. Add feedback capture linked to trace IDs
  • 6. Score a sample of traces with your evaluation judges
  • 7. Set alerts for errors, latency, cost and quality drops
  • 8. Apply redaction, access control and retention to trace data

Sampling and Retention Strategy

Capturing every request in full detail is expensive and multiplies privacy risk. A common approach: keep metrics and span metadata (timings, token counts, versions, status) for all requests; keep full content for a sample, for all requests with negative feedback or errors, and for flagged policy events; and keep full content only for a short period unless needed for an investigation or evaluation set.

Tail-based sampling, where the decision to keep a trace is made after it completes, lets you keep all slow, failed or flagged traces while sampling normal ones. Document what is captured and for how long, and align it with your privacy notices.

Observability for Cost Governance

Because every model span records tokens and cost, traces become the most accurate source for AI spend by feature, team, customer or tenant. Tag requests with feature and tenant identifiers, build dashboards for cost per request and per successful task, and alert on sudden changes, such as a prompt edit that doubled context size. These views turn cost discussions from guesses into data. Cost levers are covered in LLM cost optimization, and gateway-level attribution in AI platform engineering.

Example Instrumentation Attributes

Whichever tools you use, agree a small, consistent set of attributes on every AI span so dashboards and queries work across features. Align names with the OpenTelemetry generative AI conventions where they exist, and add your own for product context.

AttributeExamplePurpose
Provider and modelprovider, requested and response modelCompare versions, detect silent changes
Token usageinput and output tokens, cached tokensCost and efficiency
Prompt versionsupport_answer@v12Tie behaviour to releases
Feature and tenantfeature=support_answer, tenant=t_93Attribution and isolation checks
Retrieval detailsindex version, chunk IDs, scoresDebug grounding
Outcomevalidation status, finish reason, feedbackQuality signals

Worked Example

An illustrative scenario, not a client case: an HR assistant's thumbs-down rate doubles in a week with no errors logged. Filtering traces with negative feedback shows most involve leave questions, and the retrieval spans show an archived 2024 leave policy ranking above the current one after a re-index. The team excludes archived documents from the index, adds the failing questions to the evaluation set and adds an alert on retrieval of documents past their review date.

Common Mistakes

  • Logging only errors and latency, not content and versions
  • Traces that stop at service boundaries or queues
  • Capturing full prompts with personal data and keeping them indefinitely
  • No link between user feedback and the trace it refers to
  • Dashboards nobody owns or reviews

Want observability that explains every AI answer?

Talk to ZSpace Labs about LLM tracing and monitoring built on open standards and your existing stack.

Start a Project

Conclusion

LLM observability gives you the evidence to answer why a response happened, what it cost and whether quality is improving. Trace every step with versions, measure cost and quality alongside latency, protect the data and turn what you find into tests.

FAQ

Common questions

The ability to understand what an LLM application did and why, from data it emits in production: traces of each request's steps, metrics for latency, errors, tokens and cost, logs, and quality signals such as evaluation scores and user feedback.

Related services
Relevant industries
Get in touch

Have a project in mind?

Whether you're building a new digital product, improving an existing website, or looking to automate part of your business — let's talk.

Keep exploring
AI & Automation
7 min read

AI Agent Observability: How to Monitor and Debug Agentic Systems

How to monitor and debug AI agents: traces and spans for model and tool calls, token and cost tracking, latency, errors, evaluation scores, OpenTelemetry GenAI conventions, privacy and incident investigation.

Read article
AI & Automation
6 min read

AI Model Monitoring: How to Monitor Models in Production

How to monitor AI models in production: input and output monitoring, data and concept drift, sampled quality scoring, feedback and outcomes, latency, errors and cost, alerting, and when to adjust prompts or retrain.

Read article
AI & Automation
10 min read

LLMOps: A Complete Guide to Operating AI Applications in Production

What LLMOps is and how to run it: prompt and configuration management, evaluation, deployment, observability, cost control, security, governance and continuous improvement for applications built on large language models.

Read article