On-device language models are now viable for focused mobile features: offline assistants, private summarisation, Indic-language keyboards, document search, and structured extraction. The engineering challenge is not simply fitting a model into an APK or IPA. You must balance quality, memory, latency, battery, download size, licensing, and device coverage.
For Indian products, this trade-off is especially important. Users may be on 4GB or 6GB Android phones, have intermittent connectivity, or switch between English and Indic languages. A reliable small model with predictable offline behaviour can be more valuable than a larger cloud model that fails when the network is weak.
Decide what should run on-device
Start with the product task, not the model. Local inference is a strong fit when the feature needs privacy, low latency, offline access, or predictable per-user cost:
- Rewrite, classify, or extract data from short text.
- Provide autocomplete, summarisation, or form-filling assistance.
- Search a user’s private notes or downloaded documents.
- Translate or transliterate short messages, including Indic scripts.
- Power narrow conversational flows with a controlled response format.
Cloud inference may remain preferable for long documents, complex multi-step reasoning, multimodal workloads, or occasional tasks where download size and battery cost are unacceptable. A hybrid design can use a local model by default and request user consent before sending selected tasks to a server.
If your app needs Indian-language support, define language and script targets early. The considerations in this low-resource Indic NLP guide are relevant when evaluating tokenisation, training data, transliteration, and quality across languages rather than assuming English benchmarks will transfer.
Choose a model by task and device tier
Small language models are usually the right starting point. Consider models in the roughly 1B–4B range for broad Android coverage, then test larger quantised models only on capable devices. Model names and releases change quickly, so evaluate the current checkpoint, licence, tokenizer, and supported languages rather than relying on an old leaderboard.
A practical shortlist might include:
- 1B–2B models: suitable for classification, extraction, short rewriting, and entry-level devices.
- 3B–4B models: a useful middle ground for chat, summarisation, and lightweight reasoning.
- 7B–8B models: potentially strong on premium phones, but demanding in memory, thermals, and storage.
- Specialised small models: often better than general chat models for translation, speech-adjacent tasks, or structured output.
Check the model licence for commercial distribution, attribution, acceptable-use requirements, and redistribution of weights. Also test the tokenizer: a model that splits Indic text inefficiently may consume its context window rapidly and perform poorly despite good English scores.
Fine-tuning is not always necessary. Prompting, constrained decoding, and retrieval may solve the problem at lower cost. When adaptation is justified, follow a disciplined data and evaluation process using best practices for fine-tuning LLMs on custom data.
Plan memory, storage, and context
A rough weight-memory estimate is:
parameters × bits per parameter ÷ 8, plus runtime overhead, temporary buffers, and the KV cache.
For example, a 3B model at 4-bit precision needs about 1.5GB for raw weights, but the real working set will be higher. A 7B model at 4-bit may require around 4GB or more once overhead and context are included. Never size a deployment from weight files alone.
The KV cache grows with context length and can become the largest variable during long conversations. Set a deliberate context budget—often 2,048 or 4,096 tokens for mobile—and trim, summarise, or retrieve older messages. Keep prompts compact, use structured templates, and stop generation as soon as the required JSON or answer is complete.
Treat the model as a separately managed asset. A large model can make app-store distribution impractical, so consider:
- Downloading weights after installation with clear consent.
- Hosting versioned model files with resumable downloads.
- Offering a smaller default model and an optional high-quality pack.
- Verifying checksums and encrypting sensitive local files where appropriate.
- Removing unused tokenizer, training, and framework files from the bundle.
Quantise for the target hardware
Quantisation reduces memory and often improves speed, but it can reduce accuracy. Compare FP16, 8-bit, 6-bit, and 4-bit variants on your actual tasks. A 4-bit model is not automatically the best choice: some models lose instruction-following or Indic-language quality sharply at aggressive bit widths.
Common deployment formats and paths include GGUF for llama.cpp-style CPU and GPU inference, vendor-specific formats for mobile accelerators, and converted graphs for frameworks such as ExecuTorch, MediaPipe, or Core ML. AWQ and GPTQ are useful quantisation approaches, but compatibility with the selected runtime matters more than the label.
Measure:
- First-token latency and sustained tokens per second.
- Peak RAM during prompt processing and generation.
- Battery drain and device temperature over a realistic session.
- Output quality, refusal behaviour, and structured-output validity.
- Performance after the device has thermally throttled.
Select the mobile runtime
On Android, evaluate MediaPipe LLM Inference, llama.cpp-based integrations, MLC, ExecuTorch, and hardware-vendor SDKs. Vulkan, GPU delegates, and newer Android acceleration paths can help, but support varies across chipsets. On iOS, Core ML, Metal-backed runtimes, MLC, and ExecuTorch offer different trade-offs in conversion effort, operator coverage, and access to Apple silicon acceleration.
Keep the inference layer behind your own interface. Your app should be able to swap a model or backend without changing product logic. Expose cancellation, streaming tokens, maximum output length, and a clear out-of-memory fallback. Test cold start separately from warm generation; loading several gigabytes can dominate the user experience.
For more general guidance on runtime choices, profiling, and production architecture, see this guide to building high-performance AI applications with open-source tools.
Build for India’s device distribution
Create capability tiers instead of assuming every phone can run the same model:
- Entry tier: remote inference or deterministic features, with optional tiny local models.
- Standard tier: 1B–3B quantised models, short context, CPU or modest GPU execution.
- Premium tier: larger models, accelerator support, longer context, and richer offline features.
Detect available RAM, free storage, operating-system version, accelerator support, and thermal state. Keep a conservative fallback for 4GB devices. Benchmark on popular mid-range Android hardware, not only developer laptops and flagship phones.
For multilingual products, evaluate real user text: code-switching, spelling variation, romanised Indic input, names, addresses, and local formats. Open-source work from Indian builders can provide useful testing ideas; explore Indian open-source AI developer projects for relevant patterns and datasets.
Privacy, safety, and product reliability
Local inference reduces data transmission, but it does not automatically make an app private. Explain what is stored, protect prompts and generated content, and avoid placing secrets in logs or crash reports. If the app uses on-device retrieval, encrypt the local index and define deletion controls.
Add safeguards for hallucinations and prompt injection. Restrict tools and file access, validate model-generated JSON with a schema, and show uncertainty when the output affects money, health, identity, or legal decisions. For document retrieval, return citations or source snippets instead of presenting unsupported answers as facts.
A production rollout checklist
Before release, complete the following:
- Define supported device tiers and minimum free storage.
- Benchmark at least three model sizes and two quantisation levels.
- Test cold start, long sessions, airplane mode, low battery, and thermal throttling.
- Measure quality on representative Indian languages and code-switched inputs.
- Add cancellation, timeouts, crash recovery, and model-download resume.
- Version model files separately from the application and support rollback.
- Monitor opt-in, anonymised performance metrics without collecting private prompts.
- Document licences, model provenance, and third-party notices.
A good mobile LLM deployment is not the largest model that fits. It is the smallest model that meets the product’s quality bar across the devices your users actually own. Start with a narrow offline feature, establish measurements, and expand only when the battery, privacy, and reliability trade-offs remain acceptable.