Small language models can now handle useful mobile workloads: text classification, rewriting, extraction, translation, autocomplete, and compact chat experiences. Running inference on the device can reduce cloud costs, improve privacy, and keep core features working when connectivity is poor. The trade-off is that mobile hardware has limited RAM, battery, thermal headroom, and storage.
This guide explains how to run a small language model on mobile in a way that is practical for Android and iOS teams in 2026. It focuses on inference and deployment rather than training a model from scratch.
Start with the right mobile use case
Do not begin by choosing the largest model that fits in an APK. First define the task, quality threshold, and latency target.
- Classification: sentiment, spam detection, intent routing, or document labels.
- Extraction: names, amounts, dates, addresses, and structured fields from text.
- Generation: short replies, summaries, rewriting, and offline assistance.
- Language support: Indic text input, translation, transliteration, or moderation.
A classifier or small encoder model is usually cheaper and more reliable than a generative model. For Hindi and other Indian languages, review the practical constraints covered in this guide to low-resource Indic natural language processing. If your product needs regional-language generation, evaluate vocabulary coverage and real user text—not only English benchmarks.
Set targets before implementation. For example, you might require a model under 500 MB, first-token latency below two seconds, p95 request latency below five seconds, and an acceptable battery impact during a normal session.
Choose a model and runtime
Common options include compact encoder models, distilled Transformer models, and quantised decoder models. Select a model based on task accuracy, context length, licence, language coverage, and runtime support.
For a new mobile project, consider runtimes with active support for the target operating systems:
- ExecuTorch for PyTorch-based edge deployment.
- ONNX Runtime Mobile for ONNX models and cross-platform applications.
- TensorFlow Lite for compatible TensorFlow graphs and Android deployments.
- Core ML for Apple hardware, with conversion through Apple-supported tooling.
- llama.cpp-based runtimes for compatible small decoder models using GGUF files.
Runtime choice affects more than model loading. It determines available operators, accelerator support, tokenisation options, threading controls, and how easily you can ship updates. Review AI model optimisation for mobile devices before committing to a format.
Check the model licence and redistribution terms. A model that is technically small may still be unsuitable for a commercial app if its weights, tokenizer, or training-data terms do not permit your intended use.
Quantise and package the model
Quantisation lowers weight precision to reduce storage and memory use. Common formats include FP16, INT8, and lower-bit formats such as INT4. Lower precision can improve speed and fit, but it may reduce accuracy or produce less stable generation.
Use a representative calibration set that reflects real Indian user input where relevant: spelling variations, code-mixed text, local names, Romanised Indic text, and short mobile messages. Compare the quantised model with the original on both quality and safety tests.
A practical packaging checklist is:
- Keep the tokenizer and special-token configuration with the model version.
- Store large weights outside the initial app bundle when platform rules allow it.
- Use signed, versioned model downloads and verify checksums.
- Support resumable downloads and Wi-Fi-only installation for large files.
- Remove training-time assets, unused operators, and debug symbols.
- Test cold start, model reload, and low-storage behaviour.
Do not assume that a smaller file always means lower peak RAM. Some runtimes dequantise tensors or allocate temporary buffers during execution. Measure resident memory while the model is generating, not only the file size on disk.
Integrate inference on Android and iOS
On Android, load the model once per process where possible, use Kotlin coroutines or a worker thread, and keep inference off the main UI thread. Configure the number of CPU threads carefully: more threads can reduce latency but increase heat and battery use. If the runtime supports NNAPI or GPU delegates, benchmark them on representative devices rather than enabling them blindly.
On iOS, use Core ML or a compatible runtime and select compute units deliberately. Neural Engine execution can be efficient, but unsupported operations may fall back to the CPU. Profile on older supported iPhones as well as the newest device. A feature that works on a flagship phone may be unusable on a mid-range handset common among Indian users.
A typical inference pipeline is:
1. Normalise input without destroying meaningful script or punctuation.
2. Tokenise using the exact tokenizer version used during evaluation.
3. Enforce a maximum input length and truncate predictably.
4. Run inference on a background executor.
5. Stream or display output incrementally for generation tasks.
6. Cancel work when the user leaves the screen or submits a new request.
7. Release buffers and record latency, errors, and resource usage.
For generative models, limit output tokens and stop generation at defined separators. This controls both latency and cost in battery terms.
Test on real devices
Desktop benchmarks are not mobile benchmarks. Build a test matrix covering budget Android phones, mid-range devices, recent flagships, and supported iPhones. Include different chipsets, RAM levels, OS versions, and thermal conditions.
Track at least:
- Cold-start and warm-start time.
- Time to first token and tokens per second.
- Peak RAM and app size.
- Battery drain during repeated inference.
- Temperature and throttling after sustained use.
- Accuracy, hallucination rate, and failure cases.
- Behaviour with airplane mode, interruptions, and low storage.
Run long-duration tests, not just one successful prompt. Mobile performance often falls after several minutes as the device heats up. Log model version, runtime version, device model, input length, output length, and delegate used. Avoid collecting raw user text unless you have a clear consent and retention policy.
Design for privacy and reliability
On-device inference can keep sensitive messages, customer records, and voice transcripts off your servers, but the app still needs careful data handling. Do not write prompts or generated text to analytics logs by default. Encrypt local caches, restrict model access where appropriate, and provide a clear deletion path.
Use cloud fallback only when the user understands it and the product has a safe failure mode. Make offline behaviour explicit: queue work, return a deterministic error, or switch to a smaller rules-based feature. For customer-facing products, pair the model with validation and constrained outputs rather than trusting free-form text.
Teams building multilingual products can also study open-source small language models for Hindi and fine-tuning Llama for Indian regional languages. These resources are useful for comparing language coverage, adaptation strategies, and evaluation data before deployment.
A practical launch checklist
Before releasing an on-device language feature, confirm that:
- The model licence permits distribution and commercial use.
- Quantised accuracy meets the product threshold.
- Tokenisation works for every supported script and input method.
- Inference remains responsive on the lowest supported device.
- Memory, thermal, and battery limits have been measured.
- Model downloads are authenticated, resumable, and rollback-capable.
- Prompts and outputs are excluded from telemetry unless explicitly approved.
- The app handles corrupted weights, interrupted downloads, and unsupported operators.
- You have a model update and emergency disable mechanism.
FAQ
Can a small language model run fully offline?
Yes. Package the weights and tokenizer with the app or download them securely after installation. Offline operation depends on model size, runtime compatibility, and available device memory.
What model size should I target?
There is no universal limit. Start with the smallest model that meets quality requirements, then benchmark its peak RAM and thermal behaviour. A model file that fits storage may still exceed runtime memory limits.
Should I use INT8 or INT4 quantisation?
Use INT8 when quality and operator compatibility are priorities. Test INT4 for generative models where storage and memory are more constrained, but validate language quality and output stability on real inputs.
Is mobile inference suitable for Indian-language applications?
It can be, provided the tokenizer, training data, and evaluation set represent the target languages and code-mixed usage. Test Devanagari, Romanised text, spelling variation, and short conversational inputs separately.
How can I improve generation speed?
Reduce context and output length, use an efficient quantised model, reuse the KV cache where supported, tune threads, and select a hardware delegate only after measuring it across target devices.
Apply for AI Grants India
If you are building an on-device AI product for Indian users, AI Grants India can help you explore funding and support for prototyping, evaluation, and deployment.