0tokens

Apply for AI Grants India

Financial support for innovators building the future of AI in India.

Apply now

Chat · how do quantized models work

How Do Quantized Models Work? A Practical Guide

  1. aigi

    Quantization makes neural networks smaller and cheaper to run by representing weights, activations, and sometimes intermediate computations with fewer bits. Instead of storing every value as a 32-bit floating-point number, a model may use FP16, INT8, INT4, or another compact format. The result can be lower memory use, faster inference, and reduced power consumption—provided the target hardware and model support the chosen format.

    For Indian teams building voice assistants, document tools, computer-vision systems, or on-device products, quantization is often the difference between a prototype that works in a cloud notebook and a product that can run affordably at scale.

    What quantization changes

    A full-precision model stores a real value such as 0.4173 directly in floating-point format. Quantization maps a range of floating-point values to a smaller set of representable numbers. A common affine mapping is:

    q = round(x / scale) + zero_point

    Here, x is the original value, q is the integer representation, and scale and zero_point describe how to translate between the two spaces. During inference, kernels use the quantized values directly or convert them temporarily when an operation requires higher precision.

    The important distinction is between:

    • Weights: learned parameters stored in the model.
    • Activations: intermediate outputs produced for each input.
    • Inputs and outputs: data entering or leaving the model.
    • Accumulators: wider intermediate values used to prevent overflow during operations such as matrix multiplication.

    Quantization is therefore not simply “making a model smaller”. It changes the numerical representation and the kernels used to execute the model.

    Common precision formats

    • FP32: High numerical precision and broad compatibility, but large memory and compute requirements.
    • FP16 or BF16: Half-precision floating point. Often a strong choice on modern GPUs and accelerators, with limited accuracy impact.
    • INT8: Eight-bit integer weights and activations. A mature option for CPUs, mobile chips, and many edge accelerators.
    • INT4: Four-bit weights, frequently used for large language models. It can substantially reduce memory use, but accuracy and kernel support require closer testing.
    • Mixed precision: Different layers use different formats. Sensitive layers may remain in FP16 or FP32 while others use INT8 or INT4.

    A smaller bit-width does not automatically mean faster inference. Speed depends on whether the deployment device has optimized kernels, enough memory bandwidth, and support for the required format.

    How a quantized model runs

    1. Select the quantization scope

    You first decide whether to quantize weights only, weights and activations, or selected layers. Weight-only quantization is common for language models because it reduces memory pressure while leaving some computation in a higher-precision format. Full INT8 quantization can deliver better CPU efficiency when the hardware supports integer matrix multiplication.

    2. Estimate ranges and calibrate

    The system needs representative ranges for weights and activations. Calibration runs sample inputs through the model and records distributions, often using minimum and maximum values or percentile-based clipping. The calibration dataset should resemble real traffic: Indian accents for speech, local scripts for OCR, and production-like documents for retrieval or classification.

    Poor calibration can damage accuracy even when the model looks correct on a small test set. Include noisy, long, multilingual, and edge-case examples rather than calibrating only on clean benchmark data.

    3. Convert and execute operations

    Weights are converted into the target format. At inference time, optimized kernels perform operations such as matrix multiplication using integer or low-precision arithmetic. Accumulators are usually wider than the inputs—for example, INT8 values may accumulate into INT32—to reduce overflow risk. Biases and selected normalization layers may remain at higher precision.

    Some runtimes dequantize values between operations. That can preserve compatibility but reduce the speed and energy gains. A practical deployment test must measure the complete graph, not just the model file size.

    4. Validate the output

    Compare the quantized model with the original model using task-specific metrics. For a vision model, measure accuracy, recall, and latency across lighting and camera conditions. For a language model, assess perplexity, exact-match tasks, instruction following, hallucination rates, and regional-language quality. For a voice system, test word error rate across accents, background noise, and code-switching.

    Teams working on computer vision models on GitHub should also compare performance on the actual device, because desktop results rarely predict mobile or edge behavior.

    Post-training quantization versus quantization-aware training

    Post-training quantization (PTQ) converts an already trained model. It is fast, inexpensive, and usually the first method to try. Dynamic quantization estimates some activation information at runtime, while static quantization uses calibration data in advance. PTQ often works well for robust architectures and INT8 deployment.

    Quantization-aware training (QAT) simulates quantization during training. The model learns to tolerate rounding and clipping, usually through fake-quantization operations that imitate lower-precision arithmetic while retaining trainable parameters. QAT takes more compute and engineering effort but can recover accuracy when PTQ causes a meaningful regression.

    A sensible workflow is: establish a full-precision baseline, try FP16 or dynamic PTQ, calibrate static INT8, inspect layer-level errors, and move to QAT or mixed precision only where necessary. Model architecture choices also matter; understanding customizable neural network architectures helps when selecting layers that are likely to be sensitive.

    Benefits and trade-offs

    Quantized models can provide:

    • Smaller artifacts: INT8 weights are roughly one-quarter the size of FP32 weights before runtime overheads; INT4 can reduce them further.
    • Lower memory bandwidth: Smaller tensors move more efficiently between memory and compute units.
    • Lower latency: Optimized integer or low-precision kernels can accelerate inference.
    • Lower serving cost: More requests can fit on the same infrastructure, including Indian cloud and edge deployments.
    • Offline capability: Compact models are easier to package into mobile, kiosk, agricultural, industrial, and field-service applications.

    The trade-offs include accuracy loss, outlier sensitivity, unsupported operators, conversion complexity, and hardware-specific behaviour. Quantization can also expose weaknesses in models that were already poorly calibrated or overfit.

    A deployment checklist for 2026

    Before shipping, record a baseline and test the quantized candidate against it:

    • Measure model size, peak RAM, cold-start time, median and p95 latency, throughput, and power draw.
    • Test on the exact CPU, GPU, NPU, or mobile chipset used by customers.
    • Evaluate representative Indian languages, accents, names, units, scripts, and connectivity conditions where relevant.
    • Check every supported export path, such as ONNX Runtime, TensorFlow Lite, PyTorch, or the vendor SDK.
    • Inspect sensitive layers and keep them at FP16 or FP32 if mixed precision improves results.
    • Add regression tests for safety, privacy, refusal behaviour, and output formatting.
    • Monitor drift after release; quantization does not remove the need for model evaluation.

    For teams building agent products, efficient inference is only one part of production readiness. Pair it with the controls described in how to secure autonomous AI workflows. Developers choosing an implementation stack can compare options in this guide to AI frameworks for Indian student entrepreneurs.

    When should you quantize?

    Quantize when memory, latency, battery life, bandwidth, or serving cost is limiting the product. Start early enough to discover hardware and accuracy constraints, but do not optimize before establishing a reliable full-precision baseline. If a model is already fast enough and accuracy is mission-critical, the operational risk of aggressive quantization may outweigh the savings.

    For large models, weight-only INT4 or mixed-precision approaches are often practical. For compact classifiers and vision models running on CPUs or mobile devices, static INT8 may offer the best balance. The correct choice is empirical: benchmark the complete application with realistic inputs, not just the compressed checkpoint.

    FAQ

    Does quantization always reduce accuracy?
    No. FP16 may have negligible impact, and well-calibrated INT8 models can remain close to FP32 performance. Lower-bit formats are more likely to require QAT, mixed precision, or task-specific tuning.

    Is INT4 always faster than INT8?
    No. INT4 usually saves memory, but speed depends on hardware kernels, packing, dequantization overhead, and batch size.

    Can every neural network be quantized?
    Most models can be converted to some lower-precision format, but unsupported operations and sensitive layers may require a mixed-precision design.

    What is the best first step?
    Create a full-precision baseline, choose a representative calibration set, test FP16 or INT8 PTQ, and compare quality and end-to-end latency on the target device.

    Last updated 23 September 2026

AIGI may be inaccurate. Replies seeded from the guide above.