Mobile deployment changes the definition of a good AI model. A model that is accurate in a cloud notebook can still fail in a phone app because it is too large, slow at batch size one, expensive in battery use, or incompatible with the device’s accelerator. For Indian products serving intermittent connectivity, regional languages, field workers, or privacy-sensitive data, on-device inference can also reduce network dependence and keep sensitive inputs on the handset.
The goal is not simply to minimise file size. You need an acceptable trade-off between accuracy, latency, memory, energy, thermal behaviour, and device coverage. This guide explains how to make that trade-off systematically in 2026.
Start with a deployment budget
Define constraints before changing the model. Record the target phones, operating-system versions, input shape, expected inference frequency, and whether inference runs continuously or only after a user action. A rural diagnostics app, for example, may prioritise offline reliability and battery life over peak throughput; a camera filter may need predictable frame latency.
Set measurable budgets such as:
- Model package size and peak runtime memory
- P50 and P95 latency at batch size one
- Battery consumption during a representative session
- Cold-start time and time to first inference
- Minimum accuracy, recall, or word-error-rate targets
- Thermal performance after sustained use
Benchmark on representative low-, mid-, and high-end Android devices rather than relying on a flagship phone. Chipsets, drivers, and accelerator support vary substantially across the Indian market. Treat the slowest supported device as a product constraint, not an afterthought.
Choose an efficient baseline
The cheapest optimisation is selecting a model that is appropriate for the task. A compact architecture often beats an oversized model that requires aggressive compression. For vision, consider MobileNet, EfficientNet-Lite, small detection models, or a task-specific encoder. For speech and language, use a small multilingual or domain-specific model instead of sending every request through a general-purpose model.
If your product involves image or video understanding, first establish a clean evaluation set covering Indian lighting, scripts, accents, skin tones, camera quality, and connectivity conditions. Guidance on building computer vision models on GitHub can help structure the training and reproducibility workflow. For regional-language applications, compare compact models against the language coverage discussed in open-source small language models for Hindi, rather than assuming an English-first model will transfer well.
Quantise weights and activations
Quantisation reduces the numerical precision used by the model. FP32 weights can often be represented as FP16 or INT8, reducing storage and improving throughput on supported CPUs, GPUs, and NPUs.
- Dynamic-range quantisation is a fast first experiment and commonly quantises weights while choosing activation ranges at runtime.
- Full integer post-training quantisation uses a representative calibration dataset to quantise weights and activations. It is often a strong choice for edge vision and classification models.
- Quantisation-aware training (QAT) simulates reduced precision during training and is useful when post-training conversion causes unacceptable accuracy loss.
- Mixed precision keeps sensitive operations at higher precision while using INT8 or lower precision elsewhere.
Calibration data must resemble production inputs. A few clean laboratory samples will not represent low-light phone images, code-switched speech, or noisy field recordings. After conversion, compare class-level metrics, not only overall accuracy. Quantisation can disproportionately harm minority languages, rare medical findings, or low-frequency intents.
For small language models and local assistants, 4-bit quantisation can make deployment practical, but memory is not determined by weight size alone. KV cache, tokenizer tables, temporary buffers, and context length can dominate runtime memory. Local LLM deployment techniques are covered in how to deploy large language models locally.
Prune, distil, and simplify
Pruning removes parameters that contribute little to the output. Unstructured pruning creates sparse weights but may not produce real speed gains unless the runtime and hardware exploit sparsity. Structured pruning removes channels, attention heads, filters, or layers and usually delivers more predictable improvements because tensor dimensions become smaller.
Use an iterative process: prune a modest amount, fine-tune, evaluate on the full validation suite, and repeat. Do not report only parameter reduction. Measure actual latency and energy on the target device; a sparse model can be smaller without being faster.
Knowledge distillation is useful when a compact student model must retain the behaviour of a larger teacher. Train the student on ground-truth labels plus teacher logits or intermediate representations. Distillation is especially effective for intent classification, speech commands, image classification, and specialised document analysis. It can also preserve useful behaviour after quantisation, provided the teacher’s errors and biases are reviewed rather than copied blindly.
Export through a mobile runtime
Keep the model graph simple and verify every operator during conversion. Unsupported operations can force CPU fallback, create costly copies between accelerators, or prevent compilation altogether. Common deployment paths include:
- Core ML for Apple devices, with attention to Neural Engine, GPU, and CPU partitioning.
- LiteRT/TensorFlow Lite for Android and edge deployment, using delegates where supported.
- ONNX Runtime Mobile when a cross-platform ONNX pipeline fits the model and required operators.
- Platform-native runtimes or vendor SDKs when a specific NPU provides material gains.
The best runtime is the one that delivers stable performance on your device matrix, not the one with the most impressive benchmark. Inspect delegate logs, verify numerical outputs before and after conversion, and pin compatible runtime versions. For voice products, model execution is only one part of the system; audio capture, voice activity detection, decoding, and streaming often determine perceived latency. See the architecture considerations in how to build a voice agent.
Optimise the surrounding pipeline
Pre-processing and post-processing can cost as much as inference. Resize and normalise images using efficient native or hardware-backed routines. Avoid unnecessary format conversions between camera buffers and model tensors. For audio, use streaming windows and reuse buffers instead of repeatedly allocating memory.
Run inference off the UI thread, but control concurrency. Multiple simultaneous requests can increase memory pressure and trigger thermal throttling. Reuse interpreters and tensors where the framework allows it. For video, process only the required frame rate and use temporal sampling rather than analysing every frame.
For interactive applications, measure time to first result, not only average inference time. A smaller model that loads instantly may provide a better experience than a marginally faster model with a long initialisation step.
Profile on real devices
Create a repeatable benchmark harness that records latency percentiles, peak memory, CPU/GPU/NPU utilisation, temperature, battery drain, and accuracy. Test cold and warm starts, airplane mode, background load, low battery, and sustained inference. Include older supported phones and different Android vendors.
Use representative end-to-end inputs. A vision model should be tested through the camera pipeline; a speech model should be tested with microphones, interruptions, accents, and noisy environments. For medical or public-service use cases, retain audit logs for model version, runtime, device class, and confidence threshold without collecting unnecessary personal data.
Ship safely and monitor quality
Use staged rollout and keep a server-side or human-review fallback where the consequences of an incorrect prediction are high. Store model files securely, sign updates, and separate model delivery from app releases when your platform permits it. Document supported devices, precision format, known failure modes, and rollback steps.
Optimisation is complete only when the deployed model meets its product budget without unacceptable quality loss. Revisit the benchmark after OS, runtime, or chipset updates; accelerator behaviour can change across driver versions.
Practical checklist
- Define accuracy, latency, memory, battery, and device-coverage budgets.
- Establish a representative evaluation set before compression.
- Try an efficient baseline before pruning a large model.
- Compare FP16, INT8, and QAT using production-like calibration data.
- Prefer structured pruning when you need predictable hardware speedups.
- Check operator support and accelerator fallback after conversion.
- Benchmark end-to-end performance on real Android and iOS devices.
- Test sustained workloads for heat and battery impact.
- Monitor subgroup and language-specific quality after optimisation.
- Roll out gradually with a documented fallback path.
For teams building larger multimodal systems, mobile inference may be one component of a broader edge-cloud design. AI model optimisation for mobile devices offers a related technical reference, while this guide focuses on the practical path from trained model to reliable phone deployment.