What is model quantization?
Model quantization reduces the numerical precision used to store and compute a machine learning model’s weights and activations. A model trained in FP32 (32-bit floating point) may be deployed with FP16, INT8, or—in carefully tested cases—4-bit representations.
The objective is not simply to make a file smaller. Quantization can reduce memory bandwidth, improve latency, lower power consumption, and cut inference cost. The real result depends on the model architecture, hardware kernel support, runtime, and workload. An INT8 model on hardware with optimised integer arithmetic may be substantially faster; on unsupported hardware, it may deliver little benefit or even run more slowly.
For teams building products in India, this matters when models must run on affordable Android phones, clinic and factory devices, intermittent-connectivity deployments, or cost-sensitive cloud infrastructure. Quantization is one part of a broader AI model optimisation for mobile devices strategy, alongside pruning, distillation, compilation, and efficient serving.
What changes during quantization?
Quantization maps a continuous floating-point range to a finite set of lower-precision values. A common affine mapping is described using a scale and zero-point:
- Scale controls the numerical range represented by the quantized values.
- Zero-point maps real value zero to an integer value, which is useful for asymmetric data distributions.
- Weights are learned parameters and are often easier to quantize than activations.
- Activations vary with each input, so poor calibration can create larger errors.
- Accumulator precision is the precision used for intermediate results; INT8 inputs may still require INT32 accumulation.
Quantization is therefore not equivalent to rounding every number independently. It is a deployment transformation that must be supported consistently by the model format, operators, runtime, and target processor.
Main approaches
Post-training quantization
Post-training quantization (PTQ) is applied after a model has been trained. It is usually the fastest approach for an initial deployment candidate and does not require a full retraining cycle.
- Dynamic-range quantization quantizes weights in advance while calculating some activation ranges at runtime. It is simple and often useful for CPU-based language or classification models.
- Static PTQ uses a representative calibration dataset to estimate activation ranges before deployment. It generally provides better latency and memory behaviour than dynamic quantization.
- Weight-only quantization stores weights at lower precision while keeping activations at higher precision. This is widely used for large language models, particularly when memory bandwidth is the bottleneck.
- Float16 or bfloat16 conversion reduces storage and can accelerate inference on compatible GPUs, NPUs, and modern CPUs without the accuracy risks of aggressive integer quantization.
A calibration set should reflect real production inputs—not merely clean training examples. For Indian deployments, include expected language mixes, accents, lighting conditions, device cameras, network-quality artefacts, and regional data patterns where relevant.
Quantization-aware training
Quantization-aware training (QAT) inserts simulated quantization operations during training or fine-tuning. The model learns to tolerate the rounding and clipping errors it will encounter at inference time. QAT is more expensive to implement, but it can preserve accuracy when PTQ causes unacceptable degradation.
QAT is especially valuable for sensitive computer vision, speech, medical, and multilingual workloads. Teams can begin with a pretrained FP32 model, fine-tune with fake-quantisation nodes, and export only after validating the intended backend. A small, representative fine-tuning run is often enough; full retraining is not always necessary.
Mixed-precision quantization
Different layers have different sensitivity to precision loss. Mixed-precision strategies retain FP16 or FP32 in fragile layers and use INT8 or lower precision elsewhere. This often produces a better accuracy–latency compromise than applying one format to the whole network.
For language models, attention projections, embeddings, and output heads may need separate treatment. For vision models, the first and last layers can be sensitive. Measure rather than assume: layer sensitivity analysis can identify where higher precision earns its memory cost.
Benefits and trade-offs
Quantization can deliver:
- Lower memory use: smaller weights make on-device deployment possible and reduce model-loading time.
- Lower latency: integer or half-precision kernels can process more values per instruction.
- Lower serving cost: smaller models reduce bandwidth, memory pressure, and accelerator demand.
- Better energy efficiency: reduced data movement is often as important as reduced arithmetic.
- Offline capability: compact models can run without sending sensitive data to a server.
The trade-offs are equally practical. Accuracy may fall unevenly across classes, languages, demographic groups, or difficult inputs. Unsupported operators may trigger slower fallback kernels. A quantized model can also produce different confidence calibration, which matters in ranking, fraud detection, and medical decision-support systems.
A practical workflow for 2026 deployments
1. Define the target first. Record device type, RAM, processor, runtime, latency target, throughput, battery budget, and acceptable accuracy loss.
2. Establish an FP32 baseline. Measure task metrics, peak memory, cold-start time, steady-state latency, throughput, and energy where possible.
3. Choose a deployment format. Common options include TensorFlow Lite, ONNX Runtime, PyTorch export paths, and vendor-specific formats. Confirm supported operators before quantizing.
4. Create a representative calibration set. Keep it separate from evaluation data and cover production distributions, edge cases, and Indian language or geography requirements.
5. Try the least invasive method first. Compare FP16, dynamic PTQ, static INT8, and weight-only approaches before investing in QAT.
6. Validate layer and subgroup behaviour. Check per-class recall, language-wise performance, long-tail examples, calibration, and safety thresholds—not only overall accuracy.
7. Benchmark on real hardware. Desktop benchmarks frequently misrepresent Android, ARM edge boards, NPUs, or cloud instances. Test cold and warm runs.
8. Set rollback criteria. Version the model, calibration data, runtime, and delegate. Keep the FP32 or higher-precision model available for controlled fallback.
For computer vision projects, quantization should be evaluated alongside data and pipeline choices; teams may find the workflow in how to build computer vision models on GitHub useful when structuring reproducible experiments. For large language models, compare quantization against the constraints of small language models, including Hindi-focused options discussed in this guide to open-source small language models for Hindi.
Tools and deployment checks
TensorFlow Lite supports representative-dataset calibration and common integer and float conversion paths. PyTorch users can choose among eager-mode, FX graph, and export-oriented workflows depending on the model and release. ONNX Runtime provides graph optimisation and execution providers for different CPUs, GPUs, and accelerators.
Before shipping, verify:
- all operators use the intended quantized kernels;
- preprocessing and postprocessing remain numerically consistent;
- the runtime does not silently dequantize large sections of the graph;
- model size, peak RAM, latency, throughput, and energy improve on target hardware;
- monitoring can detect drift after deployment.
Quantization is not a substitute for model design. Distillation, pruning, batching, caching, compiler optimisation, and retrieval design may offer larger gains for a particular workload. Treat it as an engineering experiment with measurable acceptance criteria.
FAQ
Does quantization always reduce accuracy?
No. FP16 conversion may have negligible impact, and well-calibrated INT8 or QAT models can remain close to the FP32 baseline. The result depends on the task and model.
Should I use PTQ or QAT?
Start with PTQ for speed and simplicity. Use QAT when the measured accuracy loss is material or the application has strict reliability requirements.
Is INT8 always faster than FP16?
No. Hardware and runtime support determine performance. Benchmark both formats on the device or cloud accelerator that will serve the model.
Can quantization be used for language models?
Yes. Weight-only and mixed-precision methods are common, but evaluate generation quality, multilingual performance, context length, and tokens per second—not just file size.
Where should Indian AI teams begin?
Choose one production-like model and target, establish a baseline, calibrate on representative Indian data, and publish a comparison of accuracy, latency, memory, and cost. That evidence is more useful than claiming a generic percentage improvement.
Apply for AI Grants India
If your team is building an efficient AI product for Indian users, AI Grants India can help you identify relevant grant opportunities and prepare a stronger deployment case.