0tokens

Apply for AI Grants India

Financial support for innovators building the future of AI in India.

Apply now

Chat · how to benchmark quantized models

How to Benchmark Quantized Models in Production

  1. aigi

    Quantization can make an AI model smaller, cheaper, and faster, but those benefits are not automatic. A model that looks efficient on a laptop may deliver little improvement on an Indian-language workload, a shared cloud GPU, or an edge device with limited memory. The right benchmark compares the same model, data, runtime, and hardware before and after quantization—and records the trade-offs clearly.

    This guide explains how to benchmark quantized models for research prototypes and production systems, including language, vision, and multimodal workloads.

    Start with a fair baseline

    Benchmark the original model before quantization. Keep the model architecture, tokenizer or preprocessing pipeline, input shapes, and evaluation data unchanged. Record:

    • Model format and framework version
    • Precision, such as FP32, FP16, INT8, or 4-bit
    • Parameter count and serialized file size
    • Hardware, operating system, driver, and accelerator
    • Runtime and key optimisation settings
    • Batch size, sequence length, image resolution, and concurrency
    • Warm-up count and number of measured iterations

    A baseline prevents misleading comparisons. For example, an INT8 model running through an unoptimised compatibility path can be slower than FP16 on a device with strong floating-point support. Measure the deployment configuration you actually intend to use, not just the quantized file size.

    If the model handles Hindi, Telugu, Sanskrit, or other Indian languages, include representative multilingual data rather than relying on English-only benchmarks. For task-specific evaluation, compare your results with methods used in benchmarking NLP models for Telugu and Sanskrit. Vision teams can also use domain-relevant datasets and workflows described in how to build computer vision models on GitHub.

    Define the workload before choosing metrics

    A benchmark is useful only when it reflects the product. Specify the target scenario in advance:

    • Interactive inference: single requests, low batch size, and strict time-to-first-token or response-latency targets
    • Batch inference: sustained throughput with fixed batch sizes
    • Edge inference: limited RAM, thermal throttling, battery constraints, and offline operation
    • Server inference: concurrent users, queueing, autoscaling, and accelerator utilisation
    • Generative models: prompt length, output length, token generation rate, and KV-cache memory
    • Vision models: image resolution, preprocessing time, detection thresholds, and post-processing

    For Indian deployments, test realistic network conditions separately from model execution. A Mumbai user’s end-to-end response time may be dominated by network routing or queueing even when local inference is fast. Report both model latency and service latency.

    Measure quality loss, not just accuracy

    Quantization can affect different classes, languages, and input lengths unevenly. Evaluate the quantized model against the FP32 or FP16 baseline using a fixed test set and confidence intervals where possible.

    Useful quality measures include:

    • Accuracy, F1, precision, recall, and AUROC for classification
    • mAP, IoU, and recall for detection or segmentation
    • Word error rate or character error rate for speech systems
    • Exact match, pass rate, and task-specific scores for language models
    • Perplexity and token-level loss for language-model regression tests
    • Human review for safety, factuality, and multilingual fluency

    Do not hide regressions behind an aggregate score. Break results down by language, script, class, image condition, prompt length, and difficult examples. A small average loss may conceal a serious decline in Devanagari OCR, low-light images, or medical edge cases. Teams working with regional-language models can use open-source small language models for Hindi as a useful comparison point when designing multilingual tests.

    Benchmark latency correctly

    Latency measurements are easily distorted by cold starts, compilation, caching, and background processes. Use a repeatable protocol:

    1. Pin the model to a known device and runtime.
    2. Run warm-up requests until compilation and memory allocation stabilise.
    3. Measure enough iterations to report median, p90, p95, and p99 latency.
    4. Separate preprocessing, model execution, post-processing, and network time.
    5. Test batch sizes and concurrency levels that match production.
    6. Repeat runs at different times to identify thermal or shared-resource effects.

    For autoregressive models, report time to first token and tokens per second, while specifying prompt and output lengths. For computer vision, report images per second as well as per-image latency. A high throughput result at batch size 32 is not evidence of a good interactive experience at batch size 1.

    Track memory, power, and cost

    Quantization is often chosen to fit a model into a constrained environment. Measure peak resident memory during actual inference, not only the model file size. Include weights, activations, runtime buffers, tokenizer memory, and KV cache where relevant.

    Also record:

    • Cold-start time and disk-to-memory load time
    • CPU, GPU, NPU, and accelerator utilisation
    • Energy per inference or per 1,000 requests
    • Thermal throttling during sustained runs
    • Cloud instance cost per million inferences or million tokens
    • Failure rate under memory pressure and concurrency

    For cloud deployments, compare the complete serving stack. A cheaper quantized model may require more replicas if its runtime lacks an efficient kernel. Conversely, a slightly larger model may reduce total cost through better accelerator support. When planning deployment, pair benchmark results with the operational considerations in how to deploy large language models locally or how to deploy ML models on AWS Lambda in India.

    Use the right runtime and tooling

    Choose a runtime that supports the quantization scheme natively. Common options include TensorFlow Lite, ONNX Runtime, TensorRT, OpenVINO, llama.cpp, ExecuTorch, and vendor-specific mobile or NPU runtimes. Confirm that the runtime is using integer kernels or the intended accelerator; unsupported operators may silently fall back to FP32 on the CPU.

    Useful tooling includes:

    • Runtime profilers for operator-level timing
    • perf, Nsight Systems, or platform profilers for hardware utilisation
    • Power and thermal monitors on edge devices
    • Load generators for concurrency and throughput tests
    • MLPerf workloads for standardised cross-system comparisons
    • Experiment-tracking tools for configurations, results, and artifacts

    Inspect the operator profile. If 90% of execution remains in an unquantized layer, further reducing weight precision may not improve end-to-end performance. For multimodal applications, benchmark the complete pipeline, including the vision encoder and decoder; resources on evaluating vision models for video understanding illustrate why task-level testing matters.

    Build a reproducible benchmark report

    Store every benchmark result with the model hash, quantization configuration, calibration dataset version, runtime commit, hardware details, and test command. Calibration data should represent production inputs and exclude evaluation labels where appropriate. Run at least three independent repetitions and report distributions rather than a single best number.

    A practical results table should include:

    • Model and precision
    • Quality score and worst affected slice
    • p50, p95, and p99 latency
    • Throughput at defined batch sizes
    • Peak memory and power
    • Cost per request or token
    • Failure rate and unsupported operators

    Set release gates before testing—for example, no more than 1% F1 loss, p95 latency below 150 ms, and peak memory under 2 GB. If a model fails a gate, test mixed precision, better calibration data, quantization-aware training, or selective dequantization instead of automatically choosing a smaller bit width.

    Final checklist

    Before shipping a quantized model, verify that it:

    • Matches or exceeds the baseline on critical quality slices
    • Meets p95 and p99 latency targets on production hardware
    • Fits within memory and thermal limits during sustained use
    • Uses the intended accelerator without silent fallbacks
    • Remains reliable across languages, users, and input conditions
    • Has a reproducible benchmark artifact and rollback path

    The goal is not to produce the smallest model. It is to select the precision and runtime combination that delivers acceptable quality at the lowest practical latency, energy use, and cost for your users.

    Last updated 23 September 2026

AIGI may be inaccurate. Replies seeded from the guide above.