0tokens

Apply for AI Grants India

Financial support for innovators building the future of AI in India.

Apply now

Chat · what is 8 bit quantization

What Is 8-Bit Quantization? A Practical Guide

  1. aigi

    What is 8-bit quantization?

    8-bit quantization converts higher-precision numbers—typically FP32 or FP16—into 8-bit integer values. An 8-bit unsigned integer can represent 256 levels from 0 to 255, while a signed INT8 value usually represents -128 to 127. A model stores weights and often activations in these compact values, then uses a scale and sometimes a zero point to approximate the original numbers during inference.

    The result is a smaller, potentially faster model that is easier to run on CPUs, mobile devices, edge hardware, and inference accelerators. It is not simply “compressing a file”: quantization changes how arithmetic is performed and can affect accuracy. For a broader explanation of the design choices, see this guide to model quantization.

    How INT8 quantization works

    A common affine quantization scheme maps a real value to an integer using:

    q = round(x / scale) + zero_point

    Here, x is the original floating-point value, q is the stored integer, scale controls the step size, and zero_point represents the real value that maps to integer zero. To recover an approximation of the original value:

    x ≈ scale × (q - zero_point)

    The quantizer first observes a tensor’s range, such as its minimum and maximum values. It then divides that range into a limited number of representable levels. Values between levels are rounded to the nearest one. This rounding introduces quantization error, so the scale and range must be chosen carefully.

    Some systems use symmetric quantization, where zero is centred and the range is approximately -127 to 127. Others use asymmetric quantization, which can represent ranges that are not balanced around zero more efficiently. Per-tensor quantization uses one scale for an entire tensor; per-channel quantization uses separate scales, commonly for each output channel of a weight tensor, and often preserves accuracy better.

    Why use 8-bit quantization?

    For a model whose weights are stored in FP32, switching to INT8 can reduce weight storage by roughly four times before accounting for scales, zero points, metadata, and runtime buffers. Compared with FP16, the weight reduction is approximately two times. The practical benefits include:

    • Lower memory use: More models can fit in RAM, cache, or accelerator memory.
    • Lower bandwidth pressure: Smaller tensors move faster between memory and compute units.
    • Faster inference: CPUs and accelerators with INT8 instructions can process more operations per cycle.
    • Lower energy use: Reduced data movement and integer arithmetic can improve efficiency, especially at the edge.
    • Lower serving cost: Smaller instances and better throughput can reduce the cost of production inference.

    The gains depend on the hardware and software stack. An INT8 model will not automatically run faster if the runtime dequantizes tensors back to floating point or lacks optimised kernels. Benchmark the complete deployment path rather than relying on file size alone. For a budgeting view, compare these results with how quantization reduces inference cost.

    Weight-only, dynamic, and static quantization

    There is no single 8-bit workflow. The right method depends on the model, hardware, and inference pattern.

    • Weight-only quantization: Weights use INT8 while activations remain FP16 or FP32. This is relatively simple and useful when memory bandwidth is the main bottleneck, although it may deliver less speed than fully integer execution.
    • Dynamic quantization: Weights are quantized ahead of time, while activation ranges are calculated during inference. It requires less calibration and is often practical for transformer layers on CPUs. Learn more in this guide to dynamic quantization.
    • Static or post-training quantization: Both weights and activations are quantized using ranges gathered from a representative calibration dataset. It can provide efficient INT8 execution, but poor calibration data can cause avoidable accuracy loss. See static quantization and post-training quantization.
    • Quantization-aware training: The training process simulates low-precision effects, allowing the model to adapt before deployment. This is useful when post-training conversion causes unacceptable degradation; the quantization-aware training guide covers the trade-offs.

    What happens to accuracy?

    Accuracy loss is usually caused by outliers, narrow activation distributions, sensitive layers, or an unsuitable calibration set. A model may retain its headline benchmark score while becoming worse on the Indian languages, document types, camera conditions, or user queries that matter in production. Evaluate on a representative validation set, not only on a generic test dataset.

    Useful checks include:

    • Compare FP32 or FP16 and INT8 accuracy on the same examples.
    • Measure task-specific metrics such as F1, word error rate, retrieval recall, BLEU, or perplexity.
    • Inspect layer-wise error to identify sensitive operations.
    • Test long inputs, rare tokens, small objects, low-light images, and regional language content where relevant.
    • Measure p50 and p95 latency, throughput, peak memory, startup time, and energy on the target device.

    If degradation is concentrated in a few layers, keep those layers in FP16 or FP32 and quantize the rest. Mixed-precision deployment often provides a better production compromise than forcing every operation into INT8.

    A practical INT8 deployment workflow

    1. Define the constraint: Set targets for latency, memory, throughput, accuracy, and power.
    2. Choose the runtime: Confirm that your target CPU, GPU, NPU, or edge accelerator supports the required INT8 operators.
    3. Prepare calibration data: Use a small but representative, legally sourced sample of real inputs. Include relevant Indian languages, accents, image conditions, or document formats when applicable.
    4. Select the scheme: Start with per-channel weight quantization and choose dynamic or static activation quantization based on runtime support.
    5. Convert and validate: Export the model, check operator coverage, and compare outputs against the higher-precision baseline.
    6. Benchmark end to end: Test the actual serving stack, including tokenisation, preprocessing, batching, data transfer, and post-processing.
    7. Monitor after release: Track quality drift, latency, memory, and failure cases. Quantization should be treated as a deployment change, not a one-time file conversion.

    For large language models, INT8 is one option among several formats. The best choice may depend on the serving engine and GPU memory target; compare it with EXL2 quantization, GPTQ quantization, or AWQ quantization before committing.

    Limitations and common mistakes

    8-bit quantization is not universally faster, and it does not guarantee identical outputs. Common mistakes include using unrepresentative calibration data, quantizing outlier-heavy layers without inspection, ignoring unsupported operators, and comparing benchmarks across different batch sizes or hardware. Do not confuse 8-bit weights with a fully INT8 execution graph: activations, accumulators, and some sensitive operations may remain at higher precision.

    Bottom line

    8-bit quantization is a practical way to reduce AI model memory use and inference cost while retaining much of the quality of FP16 or FP32 models. Its success depends on the quantization scheme, calibration data, runtime kernels, target hardware, and validation metrics. Start with a representative benchmark, keep sensitive layers at higher precision when necessary, and choose the format that meets your real deployment constraints—not merely the smallest model file.

    Last updated 23 September 2026

AIGI may be inaccurate. Replies seeded from the guide above.