AI Inference Optimization: How to Reduce Latency and Serving Costs
How to optimize AI inference: measuring latency and throughput, choosing smaller or specialized models, batching, caching, quantization, speculative decoding, hardware utilization, prompt and output length and workload-specific trade-offs for hosted and self-hosted models.
Quick answer
Optimize inference by measuring time to first token, tokens per second, total latency, throughput and cost per request by feature, then working through levers in order of effort: shorten prompts and outputs, cache repeated prefixes and answers, route simple tasks to smaller models, stream and parallelize. When self-hosting, use an efficient serving engine with continuous batching and paged KV caches, evaluate quantization and speculative decoding, tune parallelism and keep GPUs well utilized. Re-evaluate quality after every change.
Where This Fits
This is the hub for our AI infrastructure cluster. Related guides cover model serving, batching and caching, quantization, GPU optimization, self-hosting, edge deployment and platform engineering. Spending controls are in LLM cost optimization and model selection in LLM routing.
Understand Where Time and Money Go
LLM inference has two phases. Prefill processes the input prompt in parallel and builds the key-value (KV) cache; it dominates time to first token and grows with prompt length. Decode generates output tokens one at a time, each depending on the previous; it dominates total time for long outputs and is usually limited by memory bandwidth rather than raw compute. Cost follows tokens processed and, when self-hosting, GPU time.
Measure per step and per feature before optimizing. A request that seems slow because of the model may actually be slow because of retrieval, sequential tool calls or a 6,000-token prompt full of unused context.
Levers by Layer
| Layer | Lever | Typical effect | Trade-off |
|---|---|---|---|
| Workload | Shorter prompts, trimmed context | Faster prefill, lower cost | Risk of removing useful context |
| Workload | Shorter outputs, structured formats | Faster decode | Less detail |
| Workload | Caching prefixes and responses | Lower latency and cost on repeats | Invalidation complexity |
| Model | Smaller or specialized model | Large speed and cost gains | Quality on harder tasks |
| Model | Quantization | Less memory, often faster | Possible quality loss |
| Serving | Continuous batching, paged KV cache | Higher throughput | Tuning, some latency trade-off |
| Serving | Speculative decoding | Faster generation | Extra complexity and memory |
| Hardware | Right GPU and parallelism, high utilization | Lower cost per token | Capacity planning |
Workload Optimizations Come First
The cheapest optimizations change what you ask the model to do. Remove unused instructions and examples, retrieve fewer but better chunks, summarize long histories, and ask for concise or structured outputs. Run independent model calls in parallel rather than in sequence. Stream outputs so users see progress early. Move non-urgent work to batch processing, which several providers offer at lower prices.
AI features too slow or too expensive?
ZSpace Labs profiles AI workloads and applies the right optimizations, from prompts to serving infrastructure. See AI engineering services.
Model Choice and Routing
Smaller models are faster and cheaper, and many tasks such as classification, extraction and routing do not need the largest model. Route tasks by difficulty, use cascades where a small model handles most requests and escalates uncertain ones, and consider distilled or fine-tuned small models for high-volume narrow tasks. Evaluate on your own data; see LLM routing.
Serving Optimizations for Self-Hosted Models
Modern inference engines implement many optimizations for you. vLLM, for example, documents PagedAttention for efficient KV cache memory, continuous batching with chunked prefill, prefix caching, speculative decoding and support for many quantization formats. Alternatives include SGLang and NVIDIA TensorRT-LLM. Benchmark engines on your models, prompt lengths and concurrency, since results depend heavily on workload.
Batching and caching are covered in LLM batching and caching, quantization in LLM quantization and GPU tuning in GPU optimization for AI.
Latency vs Throughput
Interactive features need low time to first token and steady token rates for each user; offline jobs need maximum throughput at minimum cost. These goals conflict: larger batches raise throughput but can slow individual requests. Separate workloads where possible, with interactive traffic on capacity tuned for latency and batch jobs on capacity tuned for throughput, or use priority scheduling.
Measuring the Right Things
- Time to first token at p50 and p95 per feature
- Output tokens per second per request
- Total request latency including retrieval and tools
- Throughput in requests and tokens per second
- Cost per request and per successful task
- GPU utilization and memory use when self-hosting
- Quality scores after each optimization
Advantages and Limitations
Inference optimization improves user experience and often cuts costs substantially. Its risks are quality loss that goes unnoticed and complexity that outweighs savings. Benchmarks rarely transfer between workloads, so measure on your own traffic and keep evaluation in the loop.
How to Optimize Inference Step by Step
- 1. Instrument latency, tokens and cost per step and feature
- 2. Trim prompts and outputs
- 3. Add caching for repeated prefixes and answers
- 4. Route tasks to the smallest adequate model
- 5. Optimize serving if self-hosting: engine, batching, quantization
- 6. Tune hardware and autoscaling
- 7. Re-run evaluation after each change
Optimizing RAG and Agent Latency
In retrieval and agent applications, the model call is only part of the latency. Retrieval, reranking, tool calls and sequential reasoning steps add up. Run independent retrievals and tool calls in parallel, cache embeddings for frequent queries, rerank only a short list, limit agent steps and use smaller models for routing and planning steps where evaluation allows. Stream a partial answer or show progress while slower steps finish. Trace each step so you know which one to optimize; see LLM observability.
Speculative Decoding
Speculative decoding speeds up generation by proposing several tokens cheaply, with a small draft model or other methods, and having the main model verify them in one pass. When proposals are accepted, several tokens are produced for the cost of one main-model step. Gains depend on how predictable the output is and on the engine's implementation; structured or repetitive outputs often benefit more. Engines such as vLLM support several speculative decoding methods, but measure on your workload, since it adds memory use and complexity.
Optimization by Workload Type
| Workload | Priority | Most useful levers |
|---|---|---|
| Interactive chat | Time to first token | Streaming, prefix caching, smaller models, warm capacity |
| Long document Q&A | Prefill cost | Better retrieval, prompt caching, context trimming |
| Structured extraction | Cost and accuracy | Small or fine-tuned models, batching, constrained output |
| Agents | Total task time and cost | Parallel tools, step limits, routing per step |
| Offline batch jobs | Throughput and cost | Batch APIs, large batches, off-peak capacity |
Worked Example
An illustrative scenario, not a client case: a document Q&A feature has slow first responses. Traces show prompts averaging many thousands of tokens because the full retrieved documents are included. Retrieving fewer, better-ranked chunks, enabling prompt caching for the fixed system prompt and routing simple lookup questions to a smaller model reduce time to first token and cost per answer, with evaluation scores unchanged.
Common Mistakes
- Optimizing serving before trimming prompts
- Measuring averages instead of percentiles
- Applying aggressive quantization without evaluation
- Mixing interactive and batch traffic on the same capacity
- Trusting published benchmarks for your workload
Need a performance review of your AI stack?
Talk to ZSpace Labs about inference profiling and optimization for hosted and self-hosted models.
Conclusion
Inference optimization works best from the outside in: change the workload, then the model, then the serving stack and hardware, measuring latency, cost and quality at every step.
Common questions
Reducing the latency and cost of running trained models to produce outputs, by changing the workload, the model, the serving software and the hardware, while keeping quality within acceptable limits.