0tokens

Apply for AI Grants India

Financial support for innovators building the future of AI in India.

Apply now

Chat · which quantized models can run on mobile

Which Quantized Models Can Run on Mobile? A 2026 Guide

  1. aigi

    Mobile deployment is no longer limited to tiny image classifiers. In 2026, developers can run quantized vision, speech, embedding, and small language models directly on Android and iOS devices—provided the model, runtime, and hardware are matched carefully. The right choice depends less on a model’s headline parameter count than on memory bandwidth, accelerator support, context length, thermal limits, and latency targets.

    This guide answers which quantized models can run on mobile, what precision to choose, and how to validate performance on Indian devices and networks.

    What quantization changes

    Quantization represents model weights and sometimes activations with fewer bits than the usual FP32 format. Common deployment formats include:

    • FP16: roughly half the storage of FP32; widely supported by mobile GPUs and neural processing units (NPUs), with modest accuracy change.
    • INT8: a strong default for convolutional vision, speech, and embedding models. It reduces memory and often improves CPU or accelerator throughput.
    • INT4: useful for small language models and some transformer workloads, but more sensitive to calibration and runtime support.
    • Mixed precision: keeps sensitive layers at FP16 or INT8 while compressing less sensitive layers more aggressively.

    Quantization does not automatically make every model fast. Unsupported operators can force execution back to the CPU, and poor calibration can reduce accuracy on Indian accents, scripts, lighting conditions, or camera hardware. Treat it as a deployment decision, not merely a file-compression step.

    Model families that work well on phones

    MobileNet and EfficientNet-Lite for image classification

    MobileNetV2, MobileNetV3, and EfficientNet-Lite remain dependable choices for classification, defect detection, crop diagnosis, document routing, and offline visual search. INT8 versions are typically small enough for broad Android coverage and can use TensorFlow Lite delegates or equivalent vendor accelerators.

    Choose MobileNet when latency and broad compatibility matter most. Choose EfficientNet-Lite when you can spend more compute for improved accuracy. For a wider development workflow, see this guide to building computer vision models on GitHub.

    YOLO variants for object detection

    Small YOLO variants—including YOLOv5n, YOLOv8n, YOLOv10n, and newer nano-scale releases—can run on mobile after export and quantization. They are suitable for counting, safety alerts, retail inventory, traffic analysis, and camera-assisted field applications. INT8 is usually the practical starting point, although post-training quantization may hurt small-object recall.

    Measure performance at the camera’s actual input resolution. A detector that benchmarks well at 320×320 may become unusable at 640×640 once thermal throttling and camera preprocessing are included. For segmentation rather than bounding boxes, lightweight DeepLab or MobileSAM-style variants may work, but they require substantially more memory and testing.

    Whisper and compact speech models

    Whisper Tiny and other compact automatic speech-recognition models can run on newer phones with FP16 or carefully calibrated INT8 weights. They are useful for offline transcription, voice commands, call notes, and multilingual interfaces. Speech workloads are especially sensitive to sustained compute, so benchmark a 10–15 minute recording rather than a short clip.

    For Indian deployments, test Hindi, Tamil, Telugu, Marathi, Bengali, and code-switched speech separately. Accent coverage, background noise, and microphone quality can matter more than a small difference in model size.

    Embedding and reranking models

    Small sentence-transformer and multilingual embedding models are among the easiest transformer workloads to deploy locally. Quantized embeddings support semantic search, FAQ retrieval, duplicate detection, and personal knowledge bases without sending user text to a server. INT8 generally offers a good balance; INT4 can be viable where the runtime supports it reliably.

    Keep the embedding dimension and index size under control. A compact model with a local SQLite or approximate-nearest-neighbour index often produces a better user experience than a larger model that repeatedly pages memory.

    Small language models for local chat and extraction

    Models in the roughly 0.5B–4B parameter range can run on selected modern phones when converted to mobile-friendly formats such as GGUF or vendor-specific equivalents. Examples include quantized versions of Phi, Gemma, Qwen, SmolLM, and other small instruction-tuned models. A 4-bit model may fit in memory, but the application still needs space for the key-value cache, tokenizer, runtime, operating system, and conversation context.

    For practical deployments, keep prompts short, cap generation length, stream tokens, and prefer structured extraction over unrestricted chat. A 1B–3B model is often more useful than a larger model that overheats the device or delivers unacceptable first-token latency. If your use case requires Hindi, review open-source small language models for Hindi and test the exact tokenizer and quantized checkpoint rather than relying on English benchmarks.

    How much model can a phone handle?

    A rough planning guide is more useful than a universal device limit:

    • Under 50 MB: image classifiers, keyword spotters, and narrow audio models.
    • 50–250 MB: object detectors, compact segmenters, embeddings, and speech models.
    • 250 MB–1 GB: small language models and richer multimodal components on newer mid-range and flagship devices.
    • Above 1 GB: possible on premium hardware, but risky for broad consumer distribution because of installation size, RAM pressure, and thermal behaviour.

    These figures describe the model file only. Runtime buffers, activations, camera frames, tokenizer data, and LLM KV cache can add hundreds of megabytes. Android devices with 4 GB RAM need a much more conservative design than flagship phones with 12 GB or more.

    Choosing the deployment stack

    Use the runtime that matches the model and target hardware:

    • TensorFlow Lite: mature option for mobile vision and audio, with CPU, GPU, and NNAPI paths.
    • ONNX Runtime Mobile: useful for portable ONNX models and selective operator builds.
    • Core ML: the natural iOS path, especially when conversion enables Apple Neural Engine execution.
    • ExecuTorch: a promising PyTorch-oriented route for mobile and edge targets.
    • llama.cpp and compatible runtimes: common for local small language models, especially GGUF checkpoints.

    Before committing, verify operator coverage, threading controls, delegate support, and licensing. A model may technically load while silently running unsupported layers on the CPU.

    A practical benchmarking checklist

    Benchmark on the least powerful device you intend to support, not only on a developer flagship. Record:

    • Cold-start and warm-start load time
    • Peak RAM and package size
    • First-inference and steady-state latency
    • Tokens per second for language models
    • Battery drain and device temperature
    • Accuracy, recall, transcription error rate, or retrieval quality
    • Behaviour after five to fifteen minutes of continuous use

    Compare FP16, INT8, and INT4 using the same preprocessing and test set. Include poor network conditions if the app falls back to cloud inference. For a broader deployment plan, use this AI model optimization guide for mobile devices.

    Quantization workflow

    1. Establish an FP32 or FP16 accuracy baseline.
    2. Export to the target format and remove unnecessary operators.
    3. Apply dynamic or weight-only quantization for transformers; use representative-data calibration for INT8 activations.
    4. Re-test on data that reflects real users, including Indian languages, accents, lighting, and device cameras.
    5. Profile delegates and inspect whether any layer falls back to CPU.
    6. Package a smaller model for low-memory devices and use remote configuration for staged rollout.

    Quantization-aware training is worth the effort when post-training quantization causes a meaningful accuracy drop, particularly in detection, speech, and multilingual NLP.

    Bottom line

    The safest mobile choices are INT8 MobileNet or EfficientNet-Lite for vision, nano-scale YOLO for detection, compact speech models for offline audio, small multilingual embedding models for search, and 4-bit small language models for carefully scoped local generation. Select by end-to-end latency, memory, thermals, and task accuracy—not parameter count alone.

    For Indian builders, offline capability can reduce cloud costs and improve privacy, but device fragmentation is real. Ship a measured fallback strategy: local inference on capable phones, a smaller model on constrained devices, and optional server inference when the user permits it.

    Last updated 23 September 2026

AIGI may be inaccurate. Replies seeded from the guide above.