Post-training quantization (PTQ) converts a trained AI model from high-precision formats such as FP32 into lower-precision representations, commonly INT8, without full retraining. The result is usually a smaller model that requires less memory and can run faster and more efficiently on supported CPUs, GPUs, NPUs, and edge accelerators.
For Indian teams building voice, vision, language, and industrial AI products, PTQ is often the fastest optimisation to test before investing in quantization-aware training or custom hardware. It can reduce cloud inference costs, improve latency for users on inconsistent networks, and make offline or on-device deployment practical.
Why post-training quantization matters
A model’s parameter values and intermediate activations are typically stored as floating-point numbers. Floating-point arithmetic is flexible, but it increases memory movement and compute requirements. Quantization represents these values with fewer bits and uses a scale and zero-point to map between floating-point and integer ranges.
The practical benefits include:
- Lower memory use: INT8 weights generally require about one-quarter of the storage of FP32 weights, before accounting for metadata and runtime overhead.
- Lower latency: Integer kernels can accelerate inference when the target processor supports them efficiently.
- Reduced energy consumption: Less data movement and computation can extend battery life and lower the power budget of edge devices.
- Lower serving cost: Smaller models reduce bandwidth, storage, and sometimes the number of cloud instances needed.
- Easier offline inference: Useful for field operations, factories, vehicles, and applications where sensitive data should remain on-device.
Quantization does not automatically make every model faster. A model may remain bottlenecked by memory access, unsupported operators, or conversions between FP32 and INT8. Always benchmark on the hardware and runtime you intend to ship.
How PTQ works
A typical PTQ workflow has five stages:
1. Export the trained model. Freeze the production checkpoint and export it to a supported format such as ONNX, TensorFlow Lite, OpenVINO, TensorRT, or a vendor-specific runtime.
2. Choose a target representation. Common choices include FP16, INT8, INT4, or mixed precision. INT8 is usually the safest starting point; INT4 can offer greater compression but is more sensitive to model architecture and task requirements.
3. Calibrate the model. Run representative examples through the model so the tool can estimate activation ranges. Calibration data should reflect real inputs, languages, lighting, accents, document types, or operating conditions—not merely convenient training samples.
4. Convert weights and activations. The quantizer applies scales and zero-points to weights, activations, and sometimes biases. Some layers may remain in higher precision if integer execution would create unacceptable error.
5. Validate and benchmark. Compare quality, latency, memory, throughput, power, and failure cases against the original model on the target device.
For language models, calibration and evaluation should cover token distributions, long prompts, multilingual inputs, and output quality. For computer vision, include varied illumination, camera quality, object scale, and backgrounds. Indian deployments may also need explicit testing for code-mixed text, regional languages, low-bandwidth conditions, and inexpensive Android hardware.
Main PTQ approaches
Weight-only quantization
Only model weights are quantized, while activations remain in FP16 or FP32. This is relatively low-risk and particularly useful for large language models, where weight memory is a major constraint. It may reduce storage substantially without delivering the full latency benefit of end-to-end integer execution.
Static or calibrated activation quantization
Weights and activations are quantized using ranges measured during calibration. This can produce strong speedups on INT8 hardware, but poor calibration data may cause saturation or excessive rounding error.
Dynamic quantization
Weights are quantized ahead of time, while activation scales are calculated during inference. Dynamic quantization is convenient for some CPU workloads, especially transformer layers, but runtime overhead can reduce its advantage.
Mixed-precision quantization
Different layers use different precisions. Sensitive attention, embedding, normalisation, or output layers may stay at FP16 or FP32 while robust layers use INT8 or INT4. This approach often provides a better quality-size trade-off than forcing every operation into one format.
Per-tensor and per-channel quantization
Per-tensor quantization uses one scale for an entire tensor. Per-channel quantization uses separate scales, often one per output channel, and usually preserves accuracy better for convolutional networks and models with uneven weight distributions. The right choice depends on runtime support and operator implementation.
PTQ versus quantization-aware training
PTQ is applied after training and is faster to test because it does not require updating model weights. It is a strong first option when the model is robust, calibration data is available, and deployment time is limited.
Quantization-aware training (QAT) simulates quantization during training so the model can adapt to rounding and clipping. QAT generally requires more engineering, compute, and labelled data, but it can recover quality when PTQ causes a material drop—particularly in sensitive vision, speech, and small-model use cases.
A sensible sequence is: establish FP32 quality, test FP16, run INT8 PTQ, investigate outliers, and move to QAT only if the business or product threshold is not met.
How to evaluate a quantized model
Do not rely on model size alone. Create a deployment scorecard covering:
- Task metrics such as accuracy, F1, word error rate, mean average precision, or perplexity.
- Tail latency, not only average latency, under realistic batch sizes.
- Peak RAM, persistent storage, and download size.
- Energy use and thermal throttling on the actual device.
- Unsupported operators and fallback execution in the selected runtime.
- Robustness on difficult slices, including noisy audio, low-light images, rare classes, and Indian language or code-mixed inputs.
- Privacy and data-residency requirements when deciding between on-device and cloud inference.
Use a fixed evaluation set and retain the FP32 model as the reference. If performance drops, inspect layer-wise error, activation histograms, outlier channels, and operations that silently fall back to floating point.
Teams working with public or sensitive datasets should also document provenance and validation; guidance on auditing AI training data integrity is useful when calibration data becomes part of a production release process. For language products, representative low-resource language datasets for AI training in India can help expose quality regressions that English-heavy calibration sets conceal.
Common mistakes to avoid
- Using random calibration data: Calibration must represent production traffic and edge cases.
- Quantizing before establishing a baseline: Without a trusted FP32 benchmark, quality loss is difficult to diagnose.
- Assuming INT8 guarantees speed: Confirm kernel and operator support on the target chip.
- Ignoring preprocessing and post-processing: Image resizing, tokenisation, decoding, and data transfer can dominate end-to-end latency.
- Over-quantizing sensitive layers: Keep selected layers at higher precision when mixed precision is supported.
- Skipping rollout safeguards: Use shadow traffic, canary devices, quality monitoring, and a rollback path.
A practical deployment checklist
1. Define latency, memory, power, and quality targets.
2. Select the exact runtime and hardware before choosing a quantization scheme.
3. Build a small but representative calibration set with documented consent and provenance.
4. Export and quantize using a reproducible configuration.
5. Compare FP32, FP16, INT8, and—only where justified—INT4 variants.
6. Benchmark complete application flows, not just model execution.
7. Test difficult regional, device, and connectivity conditions.
8. Ship with monitoring, versioned artefacts, and a rollback plan.
PTQ is also part of a broader efficiency strategy. Teams designing edge deployments may benefit from understanding energy-efficient AI training chips, while reproducible model delivery is easier when open-source AI model training scripts on GitHub are versioned alongside export and benchmarking steps.
FAQ
Does PTQ require retraining?
No. Basic PTQ uses an existing trained model and calibration data. If quality loss is too high, quantization-aware training may be required.
How much accuracy is lost?
There is no universal number. Some models show negligible change with INT8; others are sensitive to activation outliers, small datasets, or aggressive INT4 quantization. Measure task-specific quality.
Should every model be quantized to INT8?
No. FP16, weight-only INT8, mixed precision, or even FP32 may be better depending on the hardware, model, latency target, and quality requirements.
Can PTQ be used for large language models?
Yes. Weight-only and mixed-precision methods are widely used, but evaluate long-context behaviour, multilingual prompts, generation quality, and memory bandwidth on the serving stack.
Is PTQ useful for startups in India?
Often. It can make a proof of concept viable on lower-cost cloud instances or consumer devices, but savings should be verified against local hardware prices, traffic patterns, and support costs.
PTQ is best treated as an engineering experiment with measurable acceptance criteria—not a one-click compression step. When calibration, runtime compatibility, and evaluation are handled carefully, it can turn a research model into a practical product without immediately rebuilding the training pipeline.