Quantization can make an AI model smaller, faster, and cheaper to run, but lower numerical precision can also change its behaviour. A model that looks strong on a validation set may lose recall on rare classes, produce less reliable probabilities, or fail to meet latency targets on the hardware you actually deploy.
This guide explains how to evaluate quantized models systematically. It is designed for teams moving models to phones, edge devices, GPUs, CPUs, or constrained cloud services in India, where memory limits, variable connectivity, and hardware diversity often matter as much as benchmark accuracy.
What changes when a model is quantized?
Quantization represents weights, activations, or both with fewer bits. Common configurations include FP16, INT8, INT4, and mixed-precision formats. Quantization may be applied after training (post-training quantization) or incorporated during training (quantization-aware training).
The impact depends on the model and task:
- Weights-only quantization usually reduces storage and can be effective for large language models.
- Weight-and-activation quantization can deliver larger inference gains but is more sensitive to calibration quality.
- Static quantization uses representative calibration data to determine activation ranges.
- Dynamic quantization estimates some ranges at runtime and can be simpler to adopt.
- Mixed precision preserves higher precision in layers that are especially sensitive.
Quantization is not the same as pruning, distillation, or general model compression, although these techniques can be combined. If the model will run locally, pair evaluation with a deployment plan such as how to deploy large language models locally, rather than relying only on desktop benchmarks.
Establish a trustworthy baseline
Evaluate the original floating-point model and the quantized candidate with exactly the same test data, preprocessing, decoding settings, and post-processing. Record the model version, tokenizer, library, compiler, batch size, thread count, and hardware.
Your baseline report should include:
- Task metrics on a fixed, held-out test set
- Latency distribution, not just average latency
- Peak RAM or VRAM and model file size
- Throughput at realistic batch sizes
- Energy use or battery impact where relevant
- Output examples and known failure categories
Report absolute change and relative change. For example, “F1 fell from 0.842 to 0.831” is more useful than “performance fell by 1.3%” without context. Keep a small regression set of difficult examples so future quantization experiments remain comparable.
Measure task quality with the right metrics
Accuracy alone is rarely sufficient. Select metrics based on how the model is used and how errors affect people.
- Classification: accuracy, macro-F1, per-class precision and recall, balanced accuracy, and confusion matrices
- Detection: precision, recall, mAP, performance by object size, and missed-object rates
- Segmentation: IoU, Dice score, boundary quality, and performance across image conditions
- Speech and OCR: word or character error rate, script-wise performance, and robustness to accents or noise
- Translation: BLEU or chrF alongside human review, terminology accuracy, and adequacy
- Language models: exact match, task-specific scores, factuality, refusal behaviour, toxicity, and human preference
- Retrieval systems: recall@k, precision@k, MRR, nDCG, and answer-grounding quality
For Indian deployments, break results down by language, script, geography, device class, and connectivity conditions where applicable. A quantized model may preserve aggregate performance while degrading on Marathi, Telugu, Sanskrit, or low-resource Hindi data. Teams working on multilingual systems can compare evaluation practices with benchmarking NLP models for Telugu and Sanskrit.
Test calibration and confidence, not just predictions
Quantization can alter logits and confidence scores even when the predicted label remains unchanged. Measure calibration with reliability diagrams, expected calibration error, Brier score, and negative log-likelihood where probabilities drive decisions.
Check whether confidence still separates correct from incorrect predictions. This matters in medical triage, document processing, fraud detection, and any workflow where uncertain cases should be routed to a human. If calibration worsens, consider temperature scaling or another calibration method on a separate validation split; do not tune on the final test set.
For generative models, inspect changes in token probabilities, output length, repetition, hallucination rate, and refusal patterns. Evaluate fixed prompts as well as paraphrased prompts, and preserve decoding parameters across baseline and quantized runs.
Evaluate calibration data carefully
Representative data is one of the most important inputs to post-training quantization. It should reflect production distributions, including long inputs, rare classes, different image lighting, background noise, code-switching, and regional language variation.
Use a documented calibration sample and test sensitivity to its size and composition. Avoid leaking test examples into calibration. Compare activation ranges and identify layers with outliers. If a few layers account for most quality loss, selective higher precision or mixed precision may be more effective than abandoning quantization entirely.
For vision workloads, test conditions that resemble field use rather than clean benchmark images. This is especially important when quantized models support applications built with computer vision models on GitHub or mobile cameras.
Benchmark on target hardware
A quantized model is only useful if it improves the deployment that matters. Benchmark the actual runtime and accelerator: CPU, GPU, NPU, mobile chipset, server instance, or edge board. Framework-level estimates can differ substantially from production results.
Record:
- Cold-start and warm-start latency
- p50, p95, and p99 latency
- Throughput under realistic concurrency
- Peak memory and sustained memory use
- Power draw, thermal throttling, and battery impact
- Load time, binary size, and network transfer cost
- Failures, timeouts, and out-of-memory events
Test realistic input lengths and batch sizes. For an Indian product serving intermittent or low-bandwidth regions, measure offline behaviour and model update size as well as server latency. If the model is destined for managed infrastructure, compare the complete endpoint path using guidance such as deploying ML models on AWS Lambda in India, not just raw inference time.
Check robustness and safety regressions
Run both ordinary stress tests and targeted slices. Include distribution shifts, corrupted inputs, missing fields, noisy audio, low-resolution images, long prompts, and unusual token sequences. Compare not only the number of errors but also whether the error type changes.
For high-impact use cases, add:
- Fairness and subgroup performance checks
- Abstention or human-review thresholds
- Privacy and data-leakage tests
- Adversarial or jailbreak evaluations where relevant
- Monitoring for drift after deployment
For medical imaging, for example, preserve sensitivity on rare but consequential findings rather than optimising only overall accuracy. Domain-specific evaluation should accompany model selection, including work involving reasoning models for medical image analysis.
A practical evaluation workflow
1. Freeze the baseline model, data, preprocessing, and runtime.
2. Create a representative calibration set without contaminating the test set.
3. Quantize one configuration at a time and record all settings.
4. Run task, calibration, robustness, and subgroup evaluations.
5. Benchmark on target hardware at production-like loads.
6. Investigate layer-level or slice-level regressions.
7. Try mixed precision, better calibration data, or quantization-aware training if quality loss is unacceptable.
8. Set release gates for quality, p95 latency, memory, failure rate, and safety.
9. Shadow-test the candidate against the baseline before switching traffic.
10. Monitor post-release drift and maintain a rollback path.
A useful release policy might require no more than a defined drop in macro-F1, no regression beyond a p95 latency budget, and zero critical failures on a protected test suite. The thresholds should reflect business and safety consequences, not arbitrary percentage targets.
Common mistakes to avoid
- Comparing quantized and floating-point models with different preprocessing
- Reporting only average latency
- Benchmarking on a developer laptop instead of target hardware
- Using an unrepresentative calibration set
- Ignoring rare classes and regional language slices
- Treating model size reduction as proof of production benefit
- Tuning thresholds on the final test set
- Replacing a quantized model without a rollback and monitoring plan
FAQ
Does a smaller model always run faster?
No. Speed depends on kernel support, memory bandwidth, accelerator compatibility, batch size, and runtime overhead. Benchmark the complete serving path.
What is a reasonable accuracy loss after INT8 quantization?
There is no universal threshold. Some models show negligible change; others require calibration improvements or quantization-aware training. Set limits per task and risk level.
Should I evaluate INT4 models differently from INT8 models?
The core workflow is the same, but INT4 generally requires closer checks of outlier handling, perplexity or task quality, long-context behaviour, and hardware kernel support.
Which tools can I use?
PyTorch, TensorFlow Lite, ONNX Runtime, OpenVINO, TensorRT, and vendor-specific mobile or edge runtimes can support quantization and benchmarking. Always pin versions and document compiler and runtime settings.
How should I evaluate quantized language models?
Use task benchmarks, perplexity where meaningful, human review, safety tests, long-context prompts, multilingual slices, and end-to-end latency and memory measurements. For Hindi-focused deployments, compare against relevant small language models for Hindi rather than relying on English-only scores.
Apply for AI Grants India
If you are building an efficient AI product, evaluation infrastructure, or an India-focused deployment, learn about AI Grants India and review available support for research and implementation.