LLM Application Reliability: How to Handle Failures in Production
How to make LLM applications reliable: handling provider outages, timeouts, rate limits and malformed outputs with retries, backoff, fallback models, circuit breakers, queues, graceful degradation and human escalation.
Quick answer
Reliable LLM applications assume dependencies will fail. Set timeouts on every model call, retry only transient errors with backoff and jitter, use circuit breakers to stop hammering failing providers, keep evaluated fallback models ready, validate structured outputs and repair or retry once when they are malformed, move long work to queues, degrade gracefully when AI is unavailable and escalate to people with context when automated paths fail. Test all of it with injected faults.
Where This Fits
Reliability builds on deployment architecture and LLM gateways, which centralize many of these controls. Detection depends on LLM observability. Escalation design is in human-in-the-loop AI.
Failure Modes and Responses
| Failure | How you detect it | Response |
|---|---|---|
| Provider outage or 5xx errors | Error rate, provider status | Circuit breaker, fallback provider, degraded mode |
| Timeout | Call exceeds deadline | Retry once if budget allows; stream; move to queue |
| Rate limit (429) | Error code, retry-after header | Backoff and retry, queue, shed low-priority work |
| Malformed or truncated output | Schema validation, finish reason | Repair, retry with error, safe fallback response |
| Tool or API failure | Tool error responses | Retry if idempotent, alternative path, explain to user |
| Wrong or ungrounded answer | Validation, citation checks, feedback | Refuse or ask a clarifying question, escalate |
Timeouts, Retries and Backoff
Every model call needs a deadline that fits the user experience: an interactive answer might allow tens of seconds with streaming, a background job much longer. Set an overall request budget too, so retries cannot multiply waiting time.
Retry only errors that are likely to be transient: timeouts, rate limits and server errors. Use exponential backoff with jitter so many clients do not retry in lockstep, honour retry-after headers and cap attempts at a small number. Never retry a request that triggered side effects, such as sending an email, unless the operation is idempotent with a key that prevents duplicates.
Circuit Breakers and Fallbacks
When a provider is failing, continuing to send it traffic adds latency and load for no benefit. A circuit breaker counts failures; past a threshold it opens and sends traffic to a fallback or degraded path; after a cooling period it lets a few requests through to test recovery.
Fallback models must be chosen and evaluated before you need them. A different provider's comparable model, a smaller model for simpler tasks or a cached response for common questions can all work. Prompts may need adapting per model. Gateways often implement routing, retries and fallback centrally; see LLM gateway and LLM routing.
AI features failing when providers wobble?
ZSpace Labs designs resilient AI architectures with fallbacks, queues and graceful degradation. See our AI engineering services.
Handling Malformed Outputs
Structured output features, such as JSON schema enforcement offered by several providers, greatly reduce malformed responses but do not remove the need for validation: values can still be wrong, out of range or truncated when the output hits a token limit. Check the finish reason, validate against your schema and business rules, and on failure try a single retry that includes the validation error, or a narrower prompt. If that fails, return a safe response rather than passing bad data downstream.
Graceful Degradation
Decide in advance what users get when AI is unavailable. A search page can show normal results without a generated summary. A drafting tool can let users write manually and offer to generate later. A support assistant can show contact options and collect the question for an agent. Communicate plainly: a short message that the AI feature is temporarily unavailable is better than a frozen interface. Interface patterns are covered in AI error handling UX.
Queues for Resilience
Asynchronous work absorbs spikes and outages. Document processing, batch enrichment and agent runs can sit in a queue, retry with backoff and resume when providers recover, with users notified when results are ready. Use idempotency keys, dead-letter queues for repeatedly failing jobs and visibility into queue depth and age.
Human Escalation
Some failures should end with a person. Escalate when validation keeps failing, confidence is low, the request is sensitive or the user asks. Hand over the conversation, retrieved context and what the system tried, so the person can continue rather than restart.
Advantages and Limitations
Reliability patterns turn provider incidents into minor slowdowns rather than outages, protect budgets and keep user trust. They add complexity, and fallbacks bring their own risks: a weaker model can produce worse answers, and multiple providers mean more contracts, data reviews and prompts to maintain.
How to Improve Reliability Step by Step
- 1. Add timeouts and an overall request budget to every AI call
- 2. Retry transient errors with backoff, jitter and a cap
- 3. Validate outputs and define a single repair attempt
- 4. Choose and evaluate a fallback model or provider
- 5. Add circuit breakers, ideally in a gateway
- 6. Design degraded modes and user messaging
- 7. Inject faults in staging and run game days
Reliability Targets for AI Features
Set explicit service level objectives for AI features, as you would for other services: availability of the feature (including degraded modes), latency percentiles for time to first token and total response, and error or validation failure rates. Some teams add quality objectives, such as sampled evaluation scores staying above a threshold. Error budgets help balance shipping speed against reliability: when a feature burns its budget, prioritize reliability work over new prompts or models.
Remember that your reliability is bounded by your providers'. Read their status history and service terms, and design fallbacks for the gaps.
Testing Failure Handling
Reliability features that are never exercised tend to fail when needed. In staging, inject faults at the gateway or client level: return errors, add latency, simulate rate limits, truncate outputs and return malformed JSON. Confirm that retries stop at their limits, circuit breakers open and close, fallbacks produce acceptable answers, queues absorb bursts and the interface shows the right messages. Run periodic game days with the on-call team. Design of the user-facing side is covered in AI error handling UX.
Idempotency and Side Effects
Retries are safe only when repeating an operation has no extra effect. For model calls that only generate text, retrying is harmless apart from cost. For tool calls that create tickets, send emails, issue refunds or update records, a retry after a timeout can duplicate the action if the first attempt actually succeeded. Use idempotency keys that the receiving system checks, record completed actions in the agent's state before continuing, and make the agent check whether an action already happened before repeating it.
Design workflows so side effects happen once, near the end, after validation and confirmation, rather than scattered through many steps. Queued workflows with explicit states (pending, confirmed, executed) make recovery after failures straightforward. Orchestration patterns are covered in AI agent orchestration.
Worked Example
An illustrative scenario, not a client case: an ecommerce site's AI product Q&A goes down during a provider incident on a sale day, leaving spinning widgets on product pages. Afterwards the team adds a 10-second timeout, a circuit breaker in their gateway, an evaluated fallback model for product questions and a degraded mode that shows the product FAQ. A fault-injection test in staging confirms the page stays usable when the primary provider returns errors.
Common Mistakes
- No timeouts, so slow providers block threads and users
- Retrying everything, including non-idempotent actions
- Fallback models that were never evaluated
- Passing malformed outputs downstream
- No user-facing message when AI is unavailable
Want to stress-test your AI features?
Talk to ZSpace Labs about a reliability review with fault injection, fallback evaluation and degraded-mode design.
Conclusion
LLM applications depend on external models that will sometimes be slow, limited or wrong. Bound every call, retry carefully, fall back to evaluated alternatives, validate outputs, degrade gracefully and hand difficult cases to people.
Common questions
Provider outages and elevated errors, timeouts on long generations, rate limit errors, malformed or truncated structured outputs, tool call errors, and answers that are technically valid but wrong or ungrounded.