vLLM is an inference engine for large language models, not a model architecture or a “VLLM model” that must use one universal quantization scheme. The right format depends on the model checkpoint, GPU generation, sequence lengths, concurrency, latency target, and acceptable quality loss.
For most new production deployments in 2026, start with FP8 on supported NVIDIA GPUs or AWQ/GPTQ 4-bit weights when GPU memory is the primary constraint. Use FP16 or BF16 when quality and compatibility matter more than memory savings, and treat INT8 or bitsandbytes as workload-specific choices rather than automatic defaults.
If you need the fundamentals first, the guide to model quantization techniques and deployment trade-offs explains the terminology used here.
Quick recommendation
- Best default on Hopper and newer NVIDIA GPUs: FP8, provided the model has a compatible checkpoint and calibration is sound.
- Best for fitting a large model into limited VRAM: AWQ 4-bit or GPTQ 4-bit, after validating output quality.
- Best for maximum compatibility and predictable quality: BF16 or FP16.
- Best for rapid experiments on consumer GPUs: AWQ, GPTQ, or bitsandbytes, depending on the model and vLLM version.
- Best for CPU-heavy or unusual deployments: verify support before choosing; many popular vLLM paths are optimised for CUDA GPUs.
Quantization reduces the precision used for model weights, activations, or both. It can lower VRAM consumption and improve throughput, but it does not automatically make every request faster. Dequantization overhead, kernel availability, KV-cache memory, prompt length, and batch size all affect the result.
The formats that matter in vLLM
FP16 and BF16
FP16 and BF16 are reduced-precision floating-point formats, not aggressive weight quantization. They generally preserve model quality well and have broad hardware support. BF16 is often preferable on newer datacenter GPUs because it provides a wider exponent range and is less prone to overflow during computation.
Choose BF16 or FP16 when:
- You have enough VRAM for the model and KV cache.
- The application is sensitive to small quality changes.
- You want the simplest debugging and serving path.
- You are establishing a baseline before testing lower-bit formats.
They remain useful for Indian deployments where reliability, multilingual quality, and straightforward operations matter more than squeezing every last rupee from GPU capacity.
FP8
FP8 is usually the strongest performance-oriented option on supported recent GPUs. It can reduce memory use and improve throughput while retaining better quality than many 4-bit approaches, particularly when activations are handled correctly. Results depend heavily on the exact FP8 variant, checkpoint, calibration method, and GPU architecture.
FP8 is a good candidate for high-concurrency serving, especially when you operate modern NVIDIA hardware in a cloud or colocated inference cluster. Test long prompts and generation-heavy workloads separately: a benchmark based only on short prompts can hide KV-cache and latency problems.
AWQ 4-bit
AWQ, or Activation-aware Weight Quantization, protects weights that are especially important to activation behaviour. AWQ checkpoints are popular for serving because they can provide substantial memory savings while retaining useful quality. They are often a practical choice for running 7B, 13B, or larger models on constrained GPUs.
AWQ is attractive when:
- VRAM is the binding constraint.
- You need several concurrent requests on a single GPU.
- A trusted AWQ checkpoint already exists for your model.
- You can accept some quality loss compared with BF16.
Do not assume that every AWQ checkpoint is interchangeable. Confirm the model’s architecture, tokenizer, quantization metadata, and vLLM compatibility.
GPTQ 4-bit
GPTQ is another widely used post-training weight-quantization method. It can make large models viable on smaller GPUs, but quality and speed vary across checkpoints and kernels. GPTQ is often a sensible fallback when a well-tested checkpoint is available and AWQ is not.
Compare GPTQ and AWQ on your own prompts rather than relying on generic claims. For a customer-support assistant, measure factuality, refusal behaviour, and Indian language performance—not just tokens per second.
For a deeper comparison with transformer-serving choices, see which quantization format is best for Transformers. Teams evaluating other runtimes may also benefit from the practical guide to EXL2 quantization for LLMs, although EXL2 is not a universal replacement for vLLM formats.
INT8 and bitsandbytes
INT8 can offer a useful quality-to-memory compromise, but its value depends on kernel support and the model path used. bitsandbytes is convenient for experimentation and loading quantized models, yet convenience does not guarantee the best production throughput in vLLM.
Use INT8 or bitsandbytes when:
- The model has a validated checkpoint and serving path.
- You need an intermediate option between BF16 and 4-bit weights.
- Your GPU and vLLM release support the required kernels efficiently.
Avoid selecting INT8 merely because “8-bit” sounds like the safest option. Benchmark it against BF16 and FP8 using the same prompts, batch sizes, and concurrency.
Quantization format comparison
| Format | Memory saving | Typical quality | Best fit | Main caution |
|---|---:|---|---|---|
| BF16/FP16 | Low to moderate | Highest baseline | Compatibility and quality | Requires more VRAM |
| FP8 | Moderate to high | High when calibrated | Modern NVIDIA GPUs, high throughput | Hardware and checkpoint support |
| INT8 | Moderate | Usually high | Validated intermediate deployments | Kernel and model-path variation |
| AWQ 4-bit | High | Good to variable | VRAM-constrained serving | Checkpoint and kernel compatibility |
| GPTQ 4-bit | High | Good to variable | Existing GPTQ ecosystems | Speed and quality vary by model |
| bitsandbytes 4/8-bit | High | Variable | Prototyping and flexible loading | May not be the fastest production path |
How to choose for your deployment
Start by recording four constraints: model size, available GPU memory, target concurrency, and quality threshold. Include KV-cache requirements; a model that fits in VRAM at batch size one may fail under real traffic with long Indian-language prompts or retrieved documents.
Then run an apples-to-apples test:
1. Load the same model and tokenizer in BF16 or FP16 as the quality baseline.
2. Test FP8, AWQ, GPTQ, or INT8 checkpoints that explicitly match the architecture.
3. Measure time to first token, inter-token latency, tokens per second, peak VRAM, and concurrent requests.
4. Evaluate representative prompts in English and the languages your users actually speak, including code-switching and transliteration.
5. Check structured JSON, tool calls, refusals, citations, and long-context behaviour.
6. Repeat tests with production-like prompt lengths and traffic patterns.
For a startup building a cost-conscious serving layer, custom vLLM inference for startups covers architecture decisions around batching, autoscaling, and operational cost. Quantization should be one part of that design, not a substitute for capacity planning.
Common mistakes
- Choosing a quantization method before checking GPU and vLLM support.
- Comparing different model checkpoints instead of different formats.
- Measuring only throughput while ignoring first-token latency and quality.
- Forgetting KV-cache memory and maximum context length.
- Mixing tokenizer files or applying an incompatible quantization configuration.
- Assuming a 4-bit model is always faster than BF16; dequantization and kernel behaviour may reverse the result.
- Deploying without a rollback path to the higher-precision baseline.
The post-training quantization guide explains why calibration data and layer sensitivity affect these outcomes.
Bottom line
There is no single best format for vLLM. Use BF16 or FP16 as your baseline, FP8 for supported modern GPUs and throughput-focused serving, and AWQ or GPTQ 4-bit when VRAM savings are the priority. Select the winner with production-like benchmarks and multilingual quality tests, then pin the model, vLLM version, CUDA stack, and quantization configuration for reproducible deployments.
FAQ
Is AWQ better than GPTQ for vLLM?
Neither is universally better. AWQ can be an excellent default when a high-quality compatible checkpoint and kernel are available; GPTQ may be preferable when your chosen model has stronger GPTQ support. Benchmark both on the actual workload.
Does quantization reduce answer quality?
It can. The effect varies by model, bit width, calibration data, task, and language. Test factuality, instruction following, tool use, and regional-language prompts—not only general benchmark scores.
Should I use FP8 or 4-bit quantization?
Choose FP8 when you have supported modern hardware and want strong throughput with relatively small quality degradation. Choose 4-bit when memory capacity or deployment cost is the dominant constraint.
Is GGUF suitable for vLLM?
GGUF is primarily associated with llama.cpp-style ecosystems. It is not automatically the best or most compatible format for vLLM. Use a checkpoint and quantization method explicitly supported by your vLLM release and GPU stack; see what GGUF quantization is for context.
How often should I retest a quantized deployment?
Retest after changing the model, quantization checkpoint, vLLM version, CUDA or driver stack, GPU type, context length, or batching policy. These changes can affect both quality and performance.