0tokens

Apply for AI Grants India

Financial support for innovators building the future of AI in India.

Apply now

Chat · how to quantize a model for edge ai

How to Quantize a Model for Edge AI

  1. aigi

    Edge deployment is rarely blocked by model quality alone. A model that performs well on a workstation may exceed an Indian smartphone’s memory budget, drain a battery, or miss the latency target for a camera, factory sensor, or offline language application. Quantization addresses this gap by representing model weights and activations with fewer bits, commonly moving from FP32 to FP16, INT8, or—in selected workloads—INT4.

    For builders, the objective is not simply the smallest file. It is the best trade-off between accuracy, latency, memory, power consumption, and hardware support. This guide explains how to quantize a model for edge AI, select a method, calibrate it correctly, and verify that the resulting model works on the device that matters.

    What quantization changes

    Quantization maps floating-point values to a smaller numerical range. In an INT8 model, weights and activations are represented with 8-bit integers plus scale and zero-point parameters. The runtime reconstructs an approximation of the original values during computation.

    Common options include:

    • FP16: Usually reduces model size and can accelerate inference on supported GPUs, NPUs, and mobile processors, with limited accuracy impact.
    • INT8: The most practical starting point for production edge inference. It often delivers substantial memory and latency improvements on CPUs and NPUs.
    • INT4: Useful for large language models and memory-constrained deployments, but more sensitive to calibration, layer selection, and runtime support.
    • Dynamic quantization: Weights are quantized ahead of time while some activation scaling happens during inference. It is convenient for certain transformer and CPU workloads.
    • Static post-training quantization: A representative calibration dataset determines activation ranges before deployment.
    • Quantization-aware training (QAT): Simulated quantization is introduced during training so the model can adapt to low-precision arithmetic.

    Quantization is different from pruning and distillation. Those methods change the number of parameters or the learned representation; quantization changes numerical precision. They can be combined, but each adds its own validation requirements.

    Decide whether quantization is worth it

    First record a baseline on the target device—not only on a laptop. Measure:

    • Model size on disk and peak RAM usage
    • Cold-start and warm inference latency
    • Throughput at the intended batch size, often one for edge applications
    • Energy or battery impact over a realistic workload
    • Accuracy, recall, precision, F1, word error rate, or task-specific quality

    A camera application may prioritise p95 latency and thermal stability, while an offline Hindi speech or document-processing tool may accept slightly higher latency to preserve recognition quality. If you are building a vision system, establish a strong baseline with the workflow described in how to build computer vision models on GitHub. For mobile-specific constraints, compare quantization with the broader practices in this AI model optimization for mobile devices guide.

    A practical workflow for INT8 quantization

    1. Freeze the export path

    Choose the deployment format and runtime before quantizing. TensorFlow Lite, ONNX Runtime, ExecuTorch, Core ML, TensorRT, Qualcomm AI Engine, and vendor Android runtimes do not support exactly the same operators or data types. Exporting to an intermediate format and quantizing later can also change results.

    Check whether the target accelerator supports fully integer execution. A model described as INT8 may still fall back to FP32 for unsupported operators, reducing the expected speed and increasing memory use.

    2. Clean and representative the calibration data

    Static quantization needs a representative dataset to estimate activation ranges. Use roughly 100–1,000 samples as a starting point, then adjust based on model complexity and validation results. The data should reflect deployment conditions:

    • Different Indian lighting, accents, scripts, camera angles, and network-independent inputs where relevant
    • Normal and difficult examples, including long-tail cases
    • The same preprocessing, resizing, tokenisation, normalisation, and channel order used in production

    Do not use random synthetic inputs unless they genuinely match the operating distribution. Poor calibration can damage accuracy even when the original model is strong.

    3. Start with post-training quantization

    Post-training quantization is usually the fastest experiment. Try FP16 first if the hardware has strong half-precision support. Then test dynamic quantization or INT8 static quantization, depending on the model and runtime.

    For an ONNX workflow, the process typically involves exporting a validated model, running a calibration reader, applying static quantization, and loading the result with ONNX Runtime on the target platform. For TensorFlow Lite, provide a representative dataset to the converter and inspect whether the output is fully integer. PyTorch users should select the backend and deployment path early; eager-mode, FX graph-mode, and newer export-based flows have different operator coverage.

    Keep every artifact versioned: training checkpoint, exported model, calibration data hash, quantization configuration, runtime version, and benchmark results. Reproducibility matters when a vendor SDK changes kernel selection or operator support.

    4. Use QAT when accuracy drops

    If INT8 post-training quantization causes a material quality loss, identify where it occurs before retraining the entire model. Compare per-class metrics, confusion matrices, activation ranges, and layer-level sensitivity. Attention projections, embedding layers, first and last layers, and detection heads can be especially sensitive.

    QAT inserts fake-quantization operations during training. Fine-tune from the original checkpoint with a lower learning rate, monitor both task loss and quantized validation metrics, and export only after the fake-quantized model stabilises. Mixed precision is often better than forcing every layer to INT8: retain sensitive operations in FP16 or FP32 while quantizing the rest.

    Validate on real edge hardware

    A desktop benchmark cannot predict deployment performance. Test the exact phone, single-board computer, industrial gateway, or embedded accelerator used by customers. Measure p50 and p95 latency, sustained throughput, RAM, temperature, and power after a warm-up period.

    Also test failure cases:

    • Unsupported operators and silent CPU fallbacks
    • Different input resolutions and sequence lengths
    • Battery-saver and thermal-throttling modes
    • Offline operation and app cold starts
    • Multilingual or code-mixed inputs if the product serves Indian users

    For local language applications, quantized small language models can be a practical alternative to cloud inference. Compare the result with guidance on deploying large language models locally, and evaluate the specific behaviour of open-source small language models for Hindi rather than relying only on parameter count.

    Common mistakes to avoid

    • Quantizing before fixing the baseline: Compression cannot repair data leakage, poor preprocessing, or an underperforming model.
    • Using an unrepresentative calibration set: This is one of the most common causes of activation saturation.
    • Measuring only file size: A smaller model may not be faster if the runtime lacks an optimised kernel.
    • Ignoring fallback operators: Mixed FP32 execution can erase much of the expected benefit.
    • Evaluating only average accuracy: Track minority classes, regional data, rare conditions, and confidence calibration.
    • Assuming INT4 is automatically better: It can reduce memory substantially, but accuracy and hardware support vary widely.

    A production checklist

    Before release, confirm that you have:

    • A baseline and quantized model evaluated on the same held-out data
    • A representative calibration set that is versioned and privacy-safe
    • Per-class and worst-case quality checks, not just an aggregate score
    • Device benchmarks for latency, RAM, thermal behaviour, and energy
    • A fallback or rollback plan for unsupported hardware
    • Monitoring for drift and a process to recalibrate or retrain
    • Clear documentation of runtime, operators, precision, and model licence

    Quantization should be treated as part of the deployment pipeline, not a one-time file conversion. For multimodal or language products, also test generation quality, repetition, and instruction-following; numerical compression can expose weaknesses that task accuracy alone misses. Teams working on Indian-language NLP can pair deployment tests with benchmarking NLP models for Telugu and Sanskrit to ensure that optimisation does not disproportionately affect lower-resource languages.

    FAQ

    Does quantization always reduce accuracy?

    No. FP16 often has negligible impact, and well-calibrated INT8 can preserve production quality. Sensitive models may need QAT, mixed precision, or selective exclusion of specific layers.

    Should I choose dynamic or static quantization?

    Dynamic quantization is simpler and can work well for CPU-based transformer workloads. Static quantization generally offers more predictable performance for vision and speech models when you have representative calibration data and hardware with INT8 kernels.

    How many calibration samples do I need?

    Start with a few hundred representative samples, then test stability by changing the sample subset. More data is not automatically better; coverage of real input conditions matters more than raw volume.

    Can I quantize an LLM to INT4?

    Yes, but quality depends on architecture, calibration, group size, activation handling, and runtime kernels. Evaluate perplexity and task-level outputs, including Indian-language prompts, on the target device before committing.

    Conclusion

    The reliable way to quantize a model for edge AI is to begin with a measured baseline, choose a precision supported by the target hardware, calibrate with realistic data, and validate quality and performance together. Start with FP16 or INT8 post-training quantization; move to QAT, mixed precision, or INT4 only when the deployment constraints justify the additional complexity.

    For Indian builders, this approach makes offline-first products more practical across varied phones, gateways, cameras, and embedded systems—without treating model compression as a substitute for disciplined engineering.

    Last updated 23 September 2026

AIGI may be inaccurate. Replies seeded from the guide above.