0tokens

Apply for AI Grants India

Financial support for innovators building the future of AI in India.

Apply now

Chat · how to run quantized models offline

How to Run Quantized Models Offline: A Practical Guide

  1. aigi

    Offline inference is often the right architecture when connectivity is expensive, unreliable, or unacceptable for privacy. For Indian teams building on-device assistants, document tools, healthcare applications, industrial monitoring, or vernacular AI, quantization can make a model fit within local memory and deliver predictable latency.

    Quantization reduces the numerical precision used by model weights and, in some cases, activations. A float32 model may be converted to float16, int8, or lower-bit formats. The result is usually a smaller model with lower memory bandwidth requirements, although speed gains depend on the processor and runtime. Quantization is not a substitute for profiling: an int8 model can be slower than a float16 model if the device lacks efficient integer kernels.

    This guide explains how to run quantized models offline in a way that is reproducible, measurable, and suitable for production deployment in 2026.

    Choose the right quantization format

    Start with the target device, not the model file. The main options are:

    • Float16: Usually offers a good compromise between size, accuracy, and compatibility on modern GPUs, mobile NPUs, and some CPUs.
    • Int8: A common choice for CPU and edge inference. It can reduce memory substantially, but requires representative calibration data when activations are quantized.
    • Dynamic-range quantization: Quantizes weights while calculating some activation scales at runtime. It is easy to apply and useful for many CPU workloads.
    • Weight-only 4-bit or 8-bit quantization: Particularly useful for large language models. It reduces memory pressure, but actual speed depends on the backend and supported kernels.

    For a small vision classifier, TensorFlow Lite or ONNX Runtime may be sufficient. For local language-model inference, use a runtime and model format designed for the architecture rather than treating a generic checkpoint as deployable. Teams exploring local language applications can also compare this workflow with how to deploy large language models locally.

    1. Define the offline deployment target

    Record these constraints before converting the model:

    • Operating system and processor architecture, such as ARM64 Android, x86 Linux, or Raspberry Pi.
    • Available RAM, storage, and—if relevant—GPU or NPU memory.
    • Maximum acceptable cold-start time and per-request latency.
    • Expected input size, batch size, and concurrent users.
    • Whether the application must work fully offline or only tolerate intermittent connectivity.
    • Data-handling requirements, including whether inputs, logs, or model updates may leave the device.

    A model that works on a developer laptop may fail on a 4 GB edge computer because the runtime, application, and operating system also consume memory. Reserve headroom for peak usage rather than sizing the device to the compressed model alone.

    2. Export and quantize the model

    Use the framework’s supported export path and keep the original model as a reference. Common choices include TensorFlow Lite, ONNX Runtime, OpenVINO, and llama.cpp-compatible formats for supported language models. Check operator support before conversion; unsupported layers may be left in floating point or cause conversion to fail.

    For post-training quantization, a representative dataset should reflect actual production inputs. For an Indian-language application, include the scripts, spelling variation, code-switching, accents, image quality, and domain vocabulary expected in the field. A calibration set made only from clean English samples can produce misleadingly good benchmarks.

    A TensorFlow Lite conversion may look like this:

    import tensorflow as tf
    
    model = tf.keras.models.load_model("model.keras")
    converter = tf.lite.TFLiteConverter.from_keras_model(model)
    converter.optimizations = [tf.lite.Optimize.DEFAULT]
    
    # Use real, preprocessed production-like samples.
    def representative_dataset():
        for sample in calibration_samples:
            yield [sample.astype("float32")]
    
    converter.representative_dataset = representative_dataset
    converter.target_spec.supported_ops = [
        tf.lite.OpsSet.TFLITE_BUILTINS_INT8
    ]
    converter.inference_input_type = tf.int8
    converter.inference_output_type = tf.int8
    
    with open("model-int8.tflite", "wb") as file:
        file.write(converter.convert())

    The exact APIs vary by framework and release. Treat conversion warnings as engineering tasks, not harmless noise. Verify input and output scales, zero points, tensor shapes, and preprocessing rules after export.

    3. Select an offline runtime

    The runtime must support both the model format and the device’s acceleration path.

    • TensorFlow Lite: Practical for Android, embedded Linux, and microcontroller-oriented deployments. Test delegates such as CPU, GPU, or vendor acceleration where available.
    • ONNX Runtime: Useful when one model must run across Windows, Linux, Android, or edge hardware. Select execution providers deliberately and confirm that quantized operators are actually being used.
    • OpenVINO: A strong option for Intel CPU, GPU, and VPU deployments, especially where operator-level optimization matters.
    • PyTorch ExecuTorch: Consider for PyTorch-based mobile and edge applications, subject to operator and backend support.
    • llama.cpp and related runtimes: Common for local, quantized language models in GGUF format. Match the quantization variant and context length to available memory.

    For production, bundle the runtime and model with the application or device image. Avoid downloading dependencies during first launch if the product must operate without network access. This is especially important for field deployments with intermittent connectivity, such as rural data collection or remote industrial sites.

    4. Validate accuracy and performance separately

    Quantization quality is not established by file size. Build a test set that mirrors real usage and compare the original and quantized models on:

    • Task accuracy, F1 score, word error rate, BLEU or translation quality, depending on the application.
    • Safety and refusal behaviour for sensitive use cases.
    • Latency at cold start and steady state.
    • Peak RAM, model load time, power draw, and thermal throttling.
    • Output consistency across supported operating systems and hardware.

    For language models, measure answer quality at the intended context length and with the actual prompt templates. For vision systems, test blur, low light, compression, camera variation, and regional conditions. A computer-vision pipeline may also benefit from guidance on building computer vision models on GitHub, particularly when reproducible preprocessing and evaluation are part of the repository.

    If accuracy falls beyond the product threshold, try per-channel weight quantization, a better calibration set, selective quantization, or quantization-aware training. Do not assume that the lowest-bit model is the best model. A slightly larger int8 or float16 variant may offer materially better results and lower engineering risk.

    5. Package a reliable offline application

    Store the model with a version identifier, checksum, licence information, tokenizer or label map, preprocessing configuration, and hardware requirements. Keep these assets together; a correct model paired with an old tokenizer can silently produce bad results.

    Add an offline startup check that confirms:

    • The model file exists and passes its checksum.
    • Required operators and delegates are available.
    • Input dimensions and data types match the application.
    • Enough memory is available for inference.
    • The application can provide a useful error when acceleration is unavailable.

    For updates, use signed packages and an atomic replacement process. The device should retain the last known-good model if an update is interrupted. If connectivity is occasional, design updates as optional content rather than a prerequisite for core inference.

    Common mistakes to avoid

    • Benchmarking only on a laptop: Measure on the exact target board or phone.
    • Ignoring preprocessing: Normalisation, tokenisation, image resizing, and channel order must match training.
    • Assuming compression equals speed: Runtime kernels and memory bandwidth determine practical latency.
    • Using an unrepresentative calibration set: Quantization scales should reflect production data.
    • Shipping without observability: Record local, privacy-preserving metrics such as latency, failures, and model version.
    • Skipping licence review: Confirm that model weights, runtime libraries, and redistribution terms permit commercial or public deployment.

    A practical decision path

    For a small CPU-based classifier or detector, start with int8 post-training quantization and ONNX Runtime or TensorFlow Lite. For a mobile application, compare float16 and int8 on the actual handset. For a local language model, estimate memory for weights, runtime overhead, KV cache, and the chosen context length before selecting a 4-bit or 8-bit format. Teams working with Hindi or other Indian languages should evaluate language quality directly; model size alone says little about script handling or code-switched input. Related options include open-source small language models for Hindi.

    Offline quantization works best as a complete deployment discipline: choose hardware first, convert with representative data, validate quality and latency, package every dependency, and plan safe updates. That approach turns a compressed checkpoint into a dependable product component rather than a file that merely runs once on a developer machine.

    FAQ

    Can every model be quantized to int8?
    No. Unsupported operators, sensitive layers, or numerical instability may require mixed precision or a different export path.

    Does quantization always improve speed?
    No. Gains depend on hardware kernels, runtime support, memory access, and batch size. Benchmark on the deployment device.

    Can quantized models run without internet access?
    Yes. Bundle the model, runtime, tokenizer, and all required assets locally. Ensure the application does not depend on remote authentication, downloads, or telemetry.

    How much calibration data is needed?
    There is no universal number. Use a diverse, production-like sample and increase it until validation metrics stabilise. Preserve privacy when collecting samples.

    Should I use post-training quantization or quantization-aware training?
    Try post-training quantization first. Use quantization-aware training when the accuracy loss is too large and retraining is feasible.

    Apply for AI Grants India

    If you are building an offline AI product for Indian users, AI Grants India can help you discover funding and support opportunities for product development, pilots, and deployment.

    Last updated 23 September 2026

AIGI may be inaccurate. Replies seeded from the guide above.