LLM Quantization: How to Make Language Models Smaller and Faster
How LLM quantization works: precision formats from FP16 to 4-bit, weight and activation quantization, methods such as GPTQ and AWQ, formats such as GGUF, KV cache quantization, quality trade-offs, hardware compatibility, evaluation and deployment.
Quick answer
Quantization stores model weights, and optionally activations and the KV cache, in fewer bits so models need less memory and bandwidth and often run faster. BF16 or FP16 is the usual baseline; FP8 and INT8 typically cost little quality on supported hardware; 4-bit formats such as GPTQ, AWQ, GGUF variants and newer 4-bit floating-point formats save more but need careful evaluation. Choose formats your hardware and serving engine accelerate, calibrate with representative data and compare quality, latency and cost against the unquantized model on your own tasks.
Where This Fits
Quantization is one lever in AI inference optimization. Hardware considerations are in GPU optimization for AI, local and device deployment in AI edge deployment and running open-weight models in LLM self-hosting.
Why Quantization Helps
A model with 8 billion parameters needs roughly 16 GB just for weights in 16-bit precision, about 8 GB in 8-bit and about 4 to 5 GB in 4-bit formats including overhead. Smaller weights fit on smaller or fewer GPUs, leave more memory for the KV cache (and therefore more concurrent users or longer contexts) and reduce the bytes read per generated token, which speeds up memory-bound decoding.
Precision Formats
| Format | Bits per weight | Typical use | Notes |
|---|---|---|---|
| BF16 / FP16 | 16 | Baseline serving and training | Reference quality |
| FP8 | 8 | Weights and activations on newer GPUs | Often small quality impact; needs hardware support for speed |
| INT8 | 8 | Weights, sometimes activations | Widely supported |
| INT4 (GPTQ, AWQ) | 4 | Weight-only for GPU serving | Larger memory savings; evaluate quality |
| 4-bit floating point (NVFP4, MXFP4) | 4 | Newest GPU architectures | Hardware-specific |
| GGUF variants | 2 to 8 | llama.cpp, CPUs, laptops, edge | Many levels; lower levels lose more quality |
Methods
Post-training quantization converts a trained model without retraining, using a small calibration dataset. GPTQ quantizes weights layer by layer to minimize output error; AWQ scales weights to protect those most important to activations. Quantization-aware training simulates low precision during training or fine-tuning for better quality at low bit widths, at higher cost. QLoRA fine-tunes adapters on top of a 4-bit base model, reducing fine-tuning memory.
Serving engines support many formats; vLLM, for example, lists FP8, MXFP8/MXFP4, NVFP4, INT8, INT4, GPTQ/AWQ, GGUF and others. Check that your engine accelerates, rather than merely loads, the format you choose.
Want to run larger models on smaller hardware?
ZSpace Labs evaluates quantization options against your quality and latency targets. See AI infrastructure services.
Quality Trade-Offs
Quality loss is uneven. Quantized models may perform nearly identically on common tasks but degrade on reasoning, long contexts, code, maths, less common languages or strict output formats. Smaller models tend to be more sensitive than larger ones. Calibration data matters: calibrate on text resembling your workload. Never rely on a single benchmark score.
KV Cache Quantization
For long contexts and high concurrency, the KV cache can use more memory than the weights. Some engines support storing it in FP8 or other reduced formats, increasing the number of tokens that fit. Evaluate long-context tasks specifically, since errors can accumulate over long sequences.
Hardware Compatibility
Speed gains require hardware and kernel support. FP8 compute is available on recent data-centre GPU generations; 4-bit floating-point formats need the newest architectures; integer formats have broad support through optimized kernels. On CPUs and Apple silicon, llama.cpp and GGUF are common. On phones and embedded devices, mobile runtimes have their own supported formats; see AI edge deployment.
Evaluating a Quantized Model
- Run your application's evaluation set on both baseline and quantized models
- Check segments: long inputs, languages, structured outputs, reasoning tasks
- Measure memory, time to first token, tokens per second and throughput
- Test at realistic concurrency, not single requests
- Compare cost per successful task, not just per token
- Keep the baseline available for rollback
Advantages and Limitations
Quantization often makes self-hosting and edge deployment practical and can reduce serving costs considerably. It can quietly reduce quality on specific tasks, gains depend on hardware support and quantized community models vary in reliability. Treat it like a model change: evaluate, roll out gradually and monitor.
How to Quantize Step by Step
- 1. Define quality and performance targets
- 2. Pick formats your hardware and engine accelerate
- 3. Prefer official quantized releases or quantize with representative calibration data
- 4. Evaluate on your tasks against the baseline
- 5. Benchmark memory, latency and throughput at realistic load
- 6. Roll out gradually with monitoring
- 7. Re-evaluate when models, engines or hardware change
Quantization for Fine-Tuning
Quantization also changes fine-tuning economics. QLoRA, described in Dettmers et al., fine-tunes low-rank adapters on top of a 4-bit quantized base model, making it possible to adapt larger models on a single GPU. Quality of the resulting model should be evaluated against your tasks like any fine-tune. When deploying, you can serve the adapter on a quantized or full-precision base, and results can differ slightly, so evaluate the exact serving configuration.
Quantization on Edge Devices
On phones, laptops and embedded hardware, quantization is often required rather than optional, because memory and power are tight. Mobile and edge runtimes support specific integer and floating-point formats, and neural processing units may only accelerate certain operations and bit widths. Test the exact format on target devices for accuracy, latency, memory and battery or thermal behaviour under sustained use. See AI edge deployment.
Choosing a Quantization Level
| Situation | Reasonable starting point |
|---|---|
| Production GPU serving, quality-sensitive | FP8 or INT8 where hardware supports it, evaluate |
| Model does not fit at 8-bit | 4-bit weight-only (GPTQ or AWQ), evaluate carefully |
| Laptops and CPUs | GGUF at a mid-level quantization, test lower levels |
| Phones and embedded devices | Formats supported by the device runtime and accelerator |
| Long contexts at high concurrency | Consider KV cache quantization too |
Worked Example
An illustrative scenario, not a client case: a team wants to self-host a mid-sized open-weight model on GPUs that cannot fit it in 16-bit at the required concurrency. An FP8 version meets quality targets on their evaluation set and fits with room for KV cache. A 4-bit version uses even less memory but drops on structured extraction tasks, so they deploy FP8 for production and use the 4-bit version only for internal experimentation.
Common Mistakes
- Choosing the lowest bit width without evaluation
- Formats the hardware cannot accelerate
- Calibration data unlike the real workload
- Testing single requests instead of realistic load
- Downloading quantized models from unverified sources
Considering quantized models for production?
Talk to ZSpace Labs about a quantization evaluation on your tasks and hardware.
Conclusion
Quantization trades bits for efficiency. Choose formats your hardware accelerates, calibrate on realistic data and let evaluation on your own tasks decide how far to go.
Common questions
Representing a model's weights, and sometimes activations or KV cache, with fewer bits than the original format, for example 8 or 4 bits instead of 16, to reduce memory and often increase speed.