0tokens

Apply for AI Grants India

Financial support for innovators building the future of AI in India.

Apply now

Chat · which quantization format is best for transformers

Which Quantization Format Is Best for Transformers?

  1. aigi

    Transformers rarely fail in production because they are too accurate; they fail because they are too large, slow, expensive, or difficult to serve. Quantization reduces the precision used to store and compute model values, allowing a transformer to run with less memory and, on compatible hardware, lower latency and power consumption.

    The short answer is: INT8 is the safest general-purpose choice, INT4 is usually the best choice for memory-constrained LLM serving, FP8 is attractive on newer data-centre GPUs, and GPTQ, AWQ, or EXL2 are implementation-specific formats rather than universal winners. The right decision depends on the target accelerator, model architecture, batch size, context length, and acceptable quality loss.

    Quantization in practical terms

    A model trained in FP32 or BF16 stores weights with much higher numerical precision than inference usually requires. Quantization maps those values to a smaller representation, commonly INT8, INT4, FP8, or a packed low-bit format, along with scaling factors that help reconstruct useful values during computation.

    Quantization can reduce:

    • Memory use: INT8 weights require roughly half the storage of FP16 or BF16; INT4 requires roughly one-quarter.
    • Bandwidth pressure: Smaller weights move faster between memory and compute units.
    • Serving cost: More requests can fit on the same GPU or CPU.
    • Energy consumption: Lower data movement can matter as much as lower arithmetic cost.

    These gains are not automatic. A format only improves latency when the serving stack and hardware have an optimized kernel for it. An INT4 model running through an unoptimized dequantization path may be slower than an INT8 model with native support.

    The main formats and methods

    INT8: the dependable baseline

    INT8 is often the best first experiment for encoder models, classifiers, embedding models, and many production APIs. It offers a strong balance between quality, memory reduction, and hardware support across modern CPUs, GPUs, and edge accelerators.

    Use INT8 when you need:

    • Stable accuracy with limited calibration effort.
    • Broad deployment support.
    • Fast inference for BERT-style encoders, rerankers, and smaller language models.
    • A straightforward path through ONNX Runtime, TensorRT, OpenVINO, or PyTorch tooling.

    For transformer LLMs, weight-only INT8 can be useful, but it may not deliver the memory savings required for large models. Activation outliers and long-context workloads also require careful calibration.

    INT4: the practical choice for large LLMs

    INT4 is usually the strongest option when GPU memory is the primary constraint. Weight-only INT4 can make a model fit on a single local GPU or reduce the number of GPUs required for serving. It is widely used for instruction-tuned models and local inference.

    The trade-off is quality and kernel dependence. Four-bit quantization is more sensitive to layer selection, group size, calibration data, and outlier handling. For many LLMs, group-wise weight-only INT4 is a better starting point than quantizing weights and activations aggressively.

    Consider INT4 when:

    • The model does not fit comfortably in FP16 or BF16.
    • You are serving with vLLM, TensorRT-LLM, llama.cpp, or another optimized runtime.
    • Slight quality degradation is acceptable in exchange for lower cost.
    • You are building an on-premise or edge deployment with limited VRAM.

    FP8: efficient on newer accelerators

    FP8 retains a floating-point structure while using fewer bits than FP16 or BF16. It can offer an attractive quality and performance balance on hardware designed for FP8 matrix operations, particularly recent data-centre GPUs and compatible inference stacks.

    FP8 is not automatically the best format for a laptop GPU, CPU, or older accelerator. Verify native kernel support before choosing it. If the runtime silently converts FP8 operations back to a higher precision, the expected speed and cost benefits may disappear.

    GPTQ, AWQ, and EXL2: specialised LLM formats

    GPTQ and AWQ are commonly associated with post-training, weight-only quantization for causal language models. They can produce strong results at 4-bit precision, but their performance depends on the runtime and kernel implementation. AWQ often prioritises preserving salient weights, while GPTQ uses an optimisation procedure to reduce reconstruction error.

    EXL2 supports flexible bit rates and is designed for efficient inference in compatible ExLlama-based stacks. It can be useful when you need to trade model quality against VRAM at a finer granularity than a fixed INT4 checkpoint. Learn more in this practical guide to EXL2 quantization.

    Do not treat these names as interchangeable file extensions. A GPTQ checkpoint may be excellent in one runtime and inconvenient in another. Check supported kernels, tensor parallelism, context-length behaviour, and batching before downloading or converting a model.

    PTQ versus QAT

    Post-training quantization (PTQ) is the default for most teams because it does not require retraining. You calibrate the model with representative text or inputs, quantize it, and evaluate the result. The quality of the calibration set matters: use production-like languages, prompts, document lengths, and domain terminology.

    Our guide to post-training quantization explains the workflow and common failure modes in more detail.

    Quantization-aware training (QAT) simulates low-precision behaviour during fine-tuning. It is more expensive, but can recover accuracy when PTQ causes unacceptable degradation. QAT is worth considering for safety-critical classifiers, narrow domain models, multilingual systems, and deployments where a one-point quality loss has a measurable business cost.

    Static quantization uses calibration statistics ahead of inference, while dynamic quantization determines some activation scales at runtime. The differences are covered in this guide to static quantization. In practice, dynamic methods are convenient for CPU inference, whereas static or weight-only methods are often preferred for high-throughput GPU serving.

    A decision guide for builders

    Choose the format using this sequence:

    1. Identify the hardware. Record GPU or CPU model, available memory, supported data types, and inference engine. Hardware support matters more than a format's headline bit count.
    2. Set the quality floor. Evaluate task accuracy, exact-match scores, perplexity, refusal behaviour, and structured-output validity—not only a single benchmark.
    3. Start with the least aggressive option. Try BF16 or FP16, then INT8, then INT4 if memory or cost remains a problem. For supported modern GPUs, benchmark FP8 as well.
    4. Use representative calibration data. Include Indian English, regional languages, code, legal or government terminology, and long documents if those appear in production.
    5. Measure end-to-end performance. Test time to first token, tokens per second, p95 latency, throughput, peak VRAM, prompt length, and concurrency.
    6. Validate the actual runtime. Compare the exact checkpoint, kernel, batch size, and serving configuration you intend to deploy.

    For edge and vision workloads, quantization is only one part of deployment. Pair it with architecture and graph optimisations, as discussed in this guide to optimising vision transformers for edge deployment.

    Recommendations by use case

    • BERT, RoBERTa, and encoder APIs: Start with INT8 PTQ; use QAT if classification or extraction quality drops.
    • Small LLM on CPU: Test INT8 and INT4 in llama.cpp or an equivalent CPU-optimised runtime.
    • Large LLM on a constrained GPU: Start with AWQ or GPTQ INT4, then compare against EXL2 where supported.
    • Modern data-centre GPU: Benchmark FP8, INT8, and weight-only INT4 using the intended serving engine.
    • Multilingual or domain-specific model: Prefer INT8 first and use a calibration set that reflects actual languages and terminology.
    • Mobile or edge accelerator: Select the vendor-supported format, often INT8, rather than the smallest theoretical representation.

    Final recommendation

    There is no single best quantization format for every transformer. Choose INT8 for reliability and broad compatibility, INT4 for fitting large LLMs into limited memory, FP8 for compatible modern accelerators, and GPTQ/AWQ/EXL2 when their runtime ecosystem matches your deployment. Validate quality and latency on your own workload before committing to a checkpoint. A slightly larger model with stable output and native kernels is often better than a smaller model that requires expensive dequantization or loses important behaviour.

    Last updated 23 September 2026

AIGI may be inaccurate. Replies seeded from the guide above.