Quantization is often the difference between an AI feature that works in a lab and one that works reliably on an affordable Android phone. For mobile-first products in India, the deployment target is rarely a single flagship device: it is a wide range of chipsets, Android versions, memory limits, languages, network conditions, and battery expectations.
This guide explains how to deploy quantized models for mobile-first India products—from selecting a quantization strategy to benchmarking on real devices and managing updates after launch. The objective is not simply to make a model smaller. It is to meet a defined accuracy, latency, memory, battery, and offline-performance target for the users you intend to serve.
Start with a deployment budget
Before converting a model, define the product constraints. A voice assistant, document scanner, agriculture app, and on-device moderation feature will need different trade-offs.
Set targets for:
- Cold-start latency: How quickly the first prediction is available after the app opens.
- Warm latency: Inference time after the runtime and model are already loaded.
- Peak memory: Include the model, runtime, input buffers, and other app processes.
- Download size: Consider users installing over limited or expensive connections.
- Battery impact: Measure repeated inference, not just one successful request.
- Accuracy by use case: Test important Indian languages, accents, scripts, lighting conditions, and local document formats.
- Offline behaviour: Decide which features must work without a network connection and which can fall back to a server.
A useful baseline is to segment devices by capability rather than optimise for an average phone. Test at least one entry-level, one mid-range, and one higher-end Android device commonly used by your target audience. For broader mobile optimisation principles, see this AI model optimisation guide for mobile devices.
Choose the right quantization method
Quantization reduces numerical precision, commonly from FP32 to FP16, INT8, or lower formats. Lower precision can reduce storage and memory use, but it may also affect accuracy and operator compatibility.
FP16
FP16 is usually the least disruptive starting point. It can reduce model size substantially and work well when the device has suitable GPU or accelerator support. CPU gains vary, so benchmark rather than assume.
INT8 post-training quantization
INT8 post-training quantization is attractive when you need a fast, relatively simple conversion. It requires representative calibration data for activation ranges. That data should reflect real inputs: Hindi-English code-switching, regional speech, low-light images, compressed camera frames, or the document layouts your customers actually submit.
Quantization-aware training
Use quantization-aware training when post-training INT8 conversion causes unacceptable quality loss. The training process simulates quantization effects, allowing the model to adapt before export. This usually requires more engineering and retraining time, but it is valuable for sensitive tasks such as OCR, wake-word detection, speech recognition, and classification at decision boundaries.
Do not assume that every layer should use the same precision. Mixed-precision designs can preserve accuracy in sensitive layers while quantizing the rest. Validate the complete exported model, because unsupported operators may silently force slower execution or partial fallback to the CPU.
Select an inference stack that matches the product
Your deployment runtime should support the target operating systems, hardware delegates, model operators, and update process. Common options include TensorFlow Lite, LiteRT-compatible tooling, ONNX Runtime Mobile, and ExecuTorch-based workflows. Choose based on measured device performance and team familiarity—not popularity alone.
A practical export checklist includes:
- Confirm that every model operation is supported by the chosen runtime.
- Enable hardware acceleration only after testing correctness and stability.
- Package separate model variants when device capability differs significantly.
- Keep preprocessing identical between validation and mobile inference.
- Version the model, tokenizer, label map, normalisation constants, and runtime together.
- Fail safely when an accelerator is unavailable or a model update cannot be downloaded.
For generative features, quantization is only one part of the design. If the model is too large for reliable on-device use, compare local inference with a hybrid architecture or review approaches to deploying Mistral-7B on consumer hardware. Agentic features also need careful limits around tool calls, storage, and network retries; this production guide for open-source AI agents covers related operational concerns.
Build a representative calibration and test set
A calibration set should be small enough to run repeatedly but diverse enough to expose deployment failures. Include data from the product’s actual geographies and languages, with consent and appropriate privacy controls.
For India-focused applications, test for:
- Multiple English accents and code-switching patterns.
- Hindi and other supported Indian languages, including spelling variation and transliteration.
- Low-bandwidth uploads, noisy audio, and aggressive image compression.
- Entry-level devices with limited RAM and thermal throttling.
- Camera variation, glare, shadows, skewed documents, and regional formats.
- Long sessions, backgrounding, interruptions, and app restarts.
Track quality separately for each important cohort. An overall accuracy score can hide a severe regression for one language, device tier, or user workflow. If your product handles multilingual visual or text inputs, compare the model against relevant open-source vision-language models for Indian languages, while keeping the benchmark task-specific.
Benchmark on real phones, not only emulators
Emulators are useful for functional tests, but they do not reproduce thermal limits, memory pressure, battery behaviour, or vendor-specific drivers. Create a repeatable device test matrix and record:
- P50, P90, and P99 latency.
- Peak and sustained memory use.
- CPU, GPU, and accelerator utilisation.
- Battery drain over a realistic session.
- Accuracy before and after quantization.
- Crash rate, timeouts, and accelerator fallback frequency.
- Performance after the device heats up.
Measure cold and warm inference separately. Run enough consecutive predictions to reveal throttling. Test with other common apps open, because users will not dedicate the whole device to your model. Establish an acceptance gate before launch—for example, a maximum P95 latency on the lowest supported tier and a minimum quality score for every supported language.
Design for offline-first and hybrid operation
Quantized models are especially useful when connectivity is intermittent, but offline support needs more than a local model. Cache required assets, queue events safely, expose clear sync states, and avoid blocking the main user journey on a server response.
A hybrid pattern can work well:
- Run fast classification, wake-word detection, OCR pre-processing, or redaction locally.
- Send only necessary, consented data to the server for heavier inference.
- Provide a useful degraded mode when the network is unavailable.
- Compress telemetry and upload it in batches rather than during active inference.
- Encrypt sensitive local data and define retention and deletion rules.
For privacy-sensitive conversational products, pair local inference with the principles in this guide to building privacy-first chat apps.
Ship, monitor, and update safely
Model deployment does not end at app release. Instrument the pipeline so you can distinguish a model problem from a device, runtime, network, or product problem. Monitor latency by device tier, crash-free sessions, fallback rates, battery impact, confidence distributions, and task completion—not just API availability.
Use staged rollouts and keep the previous model available for rollback. Validate updates for package size, signature, compatibility, and migration behaviour. Avoid collecting raw user inputs by default; use aggregated metrics, redacted samples, and explicit consent where examples are necessary for debugging.
Common mistakes to avoid
- Quantizing before establishing an FP32 quality and latency baseline.
- Calibrating only on clean, high-resource training data.
- Testing on a flagship phone and claiming broad Android support.
- Ignoring preprocessing, postprocessing, and tokenisation costs.
- Shipping a model that falls back to a slow CPU path for key operators.
- Measuring one inference instead of sustained performance under heat.
- Treating download size as the only constraint while ignoring RAM and battery.
- Updating weights without versioning the runtime and preprocessing code.
The strongest mobile AI products in India treat quantization as a product-engineering discipline: define the device and user target, calibrate on representative data, benchmark sustained performance, and operate the model after release. That approach produces smaller models, but more importantly, dependable experiences across the devices and connectivity conditions that shape the market.