On-device inference changes the engineering problem. A model that performs well on a workstation may be too slow, too large, or too power-hungry for a phone. AI model optimization for mobile devices is the discipline of adapting the model, runtime, and application so inference remains responsive under real memory, thermal, battery, and connectivity constraints.
For Indian products, the target is rarely a single flagship handset. Your model may need to work across entry-level Android phones, mid-range Snapdragon and MediaTek devices, and newer iPhones. That makes representative-device testing and graceful degradation as important as benchmark accuracy.
Start with a deployment budget
Before changing the model, define measurable limits for each product flow:
- Latency: Set separate targets for first inference, steady-state inference, and end-to-end user-perceived response. Camera effects may need a stable frame rate; document scanning can tolerate a longer one-off delay.
- Memory: Track peak working memory, not just the model file size. Activations, delegates, camera buffers, tokenizer state, and duplicate tensors can trigger Android process termination.
- Power and thermals: Measure sustained workloads. A model that is fast for ten seconds may throttle after several minutes of voice, video, or navigation use.
- Package and download size: Consider compressed download size, unpacked size, and whether models can be delivered on demand.
- Accuracy by cohort: Evaluate Indian languages, accents, lighting conditions, network-independent flows, and lower-end camera hardware—not only aggregate validation accuracy.
Create a baseline on physical devices before optimization. Record p50 and p95 latency, peak RAM, energy per task, startup time, and output quality. Emulator results are useful for functional testing but should not decide hardware strategy.
Select an efficient model architecture
Optimization is most effective when efficiency is considered during model selection. For vision, compact backbones such as MobileNet, EfficientNet-Lite, and modern mobile-oriented CNN or transformer variants can reduce compute without requiring aggressive compression. For speech and language, use a task-specific small model where possible instead of embedding a general-purpose model for a narrow workflow.
Reduce unnecessary work in the application pipeline as well. Crop or resize inputs before inference, skip frames when the scene has not changed, cache embeddings, and avoid running multiple models when one shared encoder can serve several tasks. For Indian-language applications, a compact local classifier or speech model can handle common requests while a larger service handles uncertain cases. Open-source small language models for Hindi can help teams compare local model options before committing to a custom stack.
Use quantization deliberately
Quantization reduces the precision used for weights and, in some cases, activations. FP16 often provides a straightforward size and bandwidth improvement on compatible GPUs and NPUs. INT8 can deliver stronger memory and performance benefits, especially for CPUs and mobile accelerators, but requires careful calibration.
Choose the method according to risk and workflow:
- Dynamic-range or weight-only quantization: Fast to apply and useful for early experiments, though activation computation may remain expensive.
- Post-training static quantization: Calibrates activations using representative samples. It is a strong production option for fixed pipelines.
- Quantization-aware training: Simulates quantization during training and is preferable when INT8 causes material accuracy loss.
- 4-bit quantization: Valuable for on-device language models, but memory savings do not automatically translate into speed. Kernel and accelerator support matter.
Your calibration set should represent production inputs: Hindi and regional-language speech, varied accents, noisy streets, low-light images, and the camera resolutions users actually submit. Compare accuracy by language and device tier after conversion. A small average improvement can hide serious regressions for a specific user group.
Prune, distil, and compress the right parts
Structured pruning removes channels, heads, or filters, producing smaller dense tensors that standard mobile runtimes can usually accelerate. Unstructured pruning creates more zeros but often delivers no real latency benefit unless the target hardware and runtime exploit sparsity. Measure the converted model rather than relying on theoretical FLOP reductions.
Knowledge distillation is often more dependable than extreme pruning. Train a smaller student model against a stronger teacher using both ground-truth labels and teacher outputs. For classification, distil logits or feature representations; for speech and language, preserve task-specific probabilities and evaluate difficult examples separately. Distillation can reduce size while retaining quality, but it does not eliminate the need for device-level profiling.
Choose the runtime and delegate
The best runtime depends on the model format, target devices, and supported operators. LiteRT (formerly TensorFlow Lite) is a practical choice for TensorFlow-derived models and broad Android deployment. ExecuTorch suits teams working in PyTorch that want a lightweight edge runtime. On Apple platforms, Core ML can use the Neural Engine, GPU, and CPU, while MediaPipe remains useful for packaged vision and perception pipelines.
Validate operator coverage during conversion. Unsupported operations can silently force execution back to the CPU, creating large latency and battery penalties. Test the exact delegate path used in production—CPU, GPU, NNAPI, Core ML, Qualcomm acceleration, or another vendor backend—and log fallback operations.
For language models, local runtimes such as llama.cpp and MLC-based stacks can support quantized models, but memory bandwidth, context length, KV-cache growth, and token generation speed become central constraints. Teams exploring fully local deployments can also review how to deploy large language models locally.
Design for India’s device and network mix
Treat device fragmentation as a product requirement. Define support tiers, for example:
- Baseline tier: CPU-first, smaller INT8 model, reduced input resolution, and conservative concurrency.
- Standard tier: GPU or NPU acceleration with the default quality setting.
- Premium tier: Larger model, higher resolution, longer context, or richer real-time effects.
Detect capabilities at runtime rather than assuming that a chipset name guarantees acceleration. Keep a CPU fallback, but make it bounded: reduce frame rate, queue work, or switch to a smaller model when thermals rise. For intermittent connectivity, keep essential inference local and use cloud escalation only for low-confidence or high-complexity cases. This also limits API spend; teams building voice products can apply similar principles from voice AI API cost optimization.
Privacy and model security need explicit decisions. Do not place sensitive inputs in diagnostic logs. Encrypt downloaded model files where appropriate, verify their integrity, and assume that a client-side model can eventually be extracted. On-device inference improves data minimization, but it is not a complete anti-tampering mechanism.
Benchmark the complete experience
Use a repeatable test matrix covering representative Android and iOS devices, OS versions, thermal states, input sizes, and concurrent app activity. Measure:
- cold-start and warm-start latency;
- p50 and p95 inference time;
- peak RSS and allocation spikes;
- sustained throughput and throttling;
- battery drain per task;
- model download and initialization time;
- accuracy, confidence calibration, and failure rates.
Profile traces to identify whether the bottleneck is preprocessing, model execution, memory transfer, post-processing, or application scheduling. A faster model can still feel slow if image decoding or tokenization dominates. Add regression gates to CI for model size, operator support, latency, and quality on a fixed device farm.
A practical release workflow
1. Establish quality and resource baselines on real devices.
2. Select a compact architecture and remove unnecessary pipeline work.
3. Convert to FP16 or INT8 and validate accuracy by user cohort.
4. Apply distillation or structured pruning only where profiling shows value.
5. Integrate the correct runtime delegate and inspect fallbacks.
6. Add tiered models, thermal controls, and offline behavior.
7. Run sustained tests, staged rollout, and crash or ANR monitoring.
8. Revisit the model when new chipsets, OS versions, or user data change the workload.
For computer-vision teams, an efficient training-to-deployment workflow matters as much as the final conversion; this guide to building computer vision models on GitHub offers a useful starting point for reproducible experimentation.
The objective is not the smallest model at any cost. It is the best quality-per-watt, quality-per-megabyte, and quality-per-rupee that users can experience reliably. In 2026, successful mobile AI products will be designed around those trade-offs from the first prototype, not compressed hurriedly before launch.