LLM Application Deployment: How to Move an AI App Into Production
How to deploy an LLM application to production: environment separation, secrets, provider access through a gateway, containers and serverless options, streaming, scaling, access control, rollback and a production readiness checklist.
Quick answer
To deploy an LLM application, run it as a normal backend service with separate development, staging and production environments, keep provider keys in a secrets manager and call models only from the server or a gateway. Put prompts and model settings in versioned configuration behind flags, stream interactive responses, move long tasks to queues, enforce authentication, rate limits and budgets, add tracing and alerts, and pass an evaluation and security review before release with a tested rollback plan.
Where This Fits
This guide covers getting a working prototype into production. The wider practice is in LLMOps, rollout strategies in AI application release management, failure handling in LLM application reliability and backend integration patterns in AI API integration.
A Reference Architecture
Most production LLM applications share the same layers. Clients talk only to your backend. The backend handles authentication, business logic and orchestration. An AI layer, often behind a gateway, handles model access, prompts, retrieval and tools. Operations services handle secrets, tracing, feature flags and deployment.
Environments and Configuration
Use separate environments with separate provider keys, budgets, vector indexes and data. Staging should mirror production configuration closely, including model versions, so evaluation results transfer. Development can use cheaper models for iteration, but final evaluation should use the production model.
Store prompts, model identifiers, generation settings and retrieval parameters as versioned configuration rather than constants in code. Promote configuration between environments the same way you promote builds, and record which version is live in each environment. See prompt versioning for a workflow.
Secrets and Provider Access
Model provider keys are high-value credentials: a leaked key can run up large bills or expose data. Keep them in a secrets manager, inject them at runtime, rotate them and scope them per environment and, where providers allow, per project. Never ship keys in mobile apps or front-end code.
As usage grows, route all model calls through an LLM gateway that holds keys centrally and enforces rate limits, budgets, logging and fallback. This also makes switching providers or models a configuration change.
Hosting Options
| Option | Good for | Watch for |
|---|---|---|
| Serverless functions | Spiky traffic, short requests | Timeouts on long generations, cold starts, streaming support |
| Managed containers | Most applications, streaming, steady traffic | Autoscaling settings, concurrency per instance |
| Kubernetes | Many services, self-hosted models, platform teams | Operational overhead |
| Background workers + queue | Document processing, agent runs, batch jobs | Status reporting, retries, idempotency |
| Edge runtimes | Low-latency routing, light pre-processing | Runtime limits, data residency |
Have a prototype that needs to become a product?
ZSpace Labs takes AI prototypes to production with secure architecture, evaluation and monitoring. See AI application development.
Streaming, Timeouts and Long Tasks
Interactive features should stream tokens so users see progress within a second or two. Streaming affects your stack: load balancers, proxies and serverless platforms must support long-lived connections, and validation of the full output has to happen after streaming completes or on structured chunks.
Set explicit timeouts on model calls and total request time. Tasks that take minutes, such as processing large documents or running agents, belong in background queues with progress updates, retries and idempotency keys so a retry does not repeat side effects.
Scaling and Limits
Provider rate limits, measured in requests and tokens per minute, are often the first scaling constraint rather than your own servers. Track usage against limits, request increases ahead of launches, spread load across deployments or regions where providers support it, and queue or shed non-urgent work under pressure.
Protect yourself from abuse and runaway costs with per-user and per-tenant rate limits, maximum input sizes, maximum output tokens and daily budgets with alerts. Cost levers are covered in LLM cost optimization.
Access Control and Data Handling
Authenticate every request and pass the user's identity through to retrieval and tools so permissions are enforced at the data layer, not by the model. Decide what data may be sent to which providers, configure retention and regional settings, and redact sensitive fields where possible. See AI data privacy and AI data leakage.
Production Readiness Checklist
- Evaluation results meet agreed thresholds on the production configuration
- Security review done: injection, output handling, tool permissions, secrets
- Rate limits, input limits, output limits and budgets configured
- Tracing with versions, tokens, latency and cost on every request
- Alerts for errors, latency, cost spikes and quality signals
- Fallbacks for provider failure and a user-facing degraded mode
- Rollback tested for code, prompts and model settings
- Data retention and deletion paths documented
- An owner and on-call process for the feature
Advantages and Limitations
Deploying through a disciplined architecture keeps keys safe, costs bounded and behaviour reversible. It adds components, such as a gateway, configuration store and tracing, that small teams may find heavy at first. Start with the essentials (server-side calls, secrets, limits, tracing and versioned prompts) and add the rest as usage grows.
How to Deploy Step by Step
- 1. Move model calls server-side and remove keys from clients
- 2. Externalize prompts and settings into versioned configuration
- 3. Set up environments with separate keys, data and budgets
- 4. Containerize or choose serverless based on request length and streaming
- 5. Add limits, tracing and alerts
- 6. Run evaluation and security review on staging
- 7. Release behind a flag to a small share of users, then expand
Infrastructure as Code and CI/CD
Define AI infrastructure the same way as the rest of your stack: gateway configuration, secrets references, queues, vector databases, autoscaling rules and alerts in infrastructure-as-code templates reviewed through pull requests. This makes environments reproducible and changes auditable.
In CI, run unit tests, the fast evaluation subset and security checks on every change; build container images, scan them and deploy to staging automatically; run the full evaluation suite there; then promote to production behind a flag. Keep prompt and model configuration deployable independently of code so behaviour fixes do not wait for a full release. The release side is covered in AI application release management.
Multi-Provider and Regional Deployment
Production applications often need more than one model provider: for fallback during outages, for regional data residency or for routing different tasks to different models. Abstract provider calls behind a gateway or internal client, keep prompts adaptable per model and evaluate each provider-model pair you might use. Regional requirements may mean deploying the application, vector store and model endpoints in the same region and verifying that logs and traces stay there too. Cost and routing considerations are in LLM routing and LLM gateway.
Example Deployment Configuration
Keeping AI-specific settings in one configuration file per environment makes reviews and promotion simple. The example below is illustrative; adapt names to your stack.
environment: production
gateway:
base_url: https://ai-gateway.internal
api_key_secret: secrets/ai-gateway/prod
features:
support_answer:
prompt: support_answer@v12
model: <provider/model-version>
fallback_model: <provider-b/model-version>
max_input_tokens: 8000
max_output_tokens: 600
timeout_seconds: 20
retrieval: { index: kb_v14, top_k: 8, rerank: true }
rollout: { flag: support_answer_v12, percent: 25 }
limits:
per_user_requests_per_minute: 20
daily_budget_usd: 400
observability:
tracing: otel
capture_content: sampled_redacted
retention_days: 14Worked Example
An illustrative scenario, not a client case: a startup's AI report generator calls a model directly from the browser with an embedded key, and a scraped key leads to unexpected charges. The team moves calls behind an authenticated API, stores keys in a secrets manager, moves report generation to a background queue with progress updates, adds per-account daily limits and puts prompts behind a flag. The next release goes to 5% of accounts first.
Common Mistakes
- Calling providers from client code with embedded keys
- Using one provider key and budget for every environment
- Serverless timeouts cutting off long generations
- No limits on input size, output tokens or spend
- No way to roll back a prompt without a full redeploy
Want a production review before launch?
Talk to ZSpace Labs about a production readiness review covering architecture, security, evaluation and operations.
Conclusion
Deploying an LLM application is ordinary service deployment plus control over prompts, models, providers and cost. Keep model access server-side, version configuration, plan for long requests, limit usage and make every change reversible.
Common questions
The application code deploys like any service, but behaviour also depends on prompts, model versions, retrieval indexes and provider availability. Deployment therefore includes configuration management, provider access, streaming, cost limits and the ability to roll back prompts and models independently of code.