Edge AI is no longer limited to simple image classifiers. In 2026, developers can run object detection, speech recognition, translation, anomaly detection, and compact language models on phones, gateways, cameras, Raspberry Pi-class computers, and specialised NPUs. The deciding factor is not simply whether a model is “small”. It is whether its quantized format, operators, runtime, and hardware accelerator work together.
This guide explains which quantized models can run on edge devices, what to measure, and how to select a deployment path for Indian products that may need to work offline, on intermittent connectivity, or under strict data-residency requirements.
What quantization changes
Quantization represents model weights and sometimes activations with fewer bits. A typical FP32 model uses 32-bit floating-point values; an INT8 model uses 8-bit integers. Weight-only 4-bit formats can reduce memory further, although they do not always deliver the same speed or energy benefit as fully integer inference.
The main options are:
- FP16: Usually offers a modest size reduction with limited accuracy risk. Suitable for GPUs and mobile processors that support half-precision arithmetic.
- INT8 post-training quantization: A practical default for CNNs and many production pipelines. It requires representative calibration data and can use integer-friendly NPUs.
- Quantization-aware training (QAT): Simulates low-precision arithmetic during training. It often preserves accuracy better when post-training quantization causes a material drop.
- INT4 or lower-bit weight quantization: Useful for language models where memory is the main constraint. Runtime support and token-generation speed vary significantly.
Quantization does not automatically make a model faster. Unsupported operators may fall back to the CPU, conversions between data types can add overhead, and memory bandwidth may remain the bottleneck. Always benchmark the complete exported model on the target board.
Quantized vision models for edge deployment
For image classification, MobileNetV3, MobileNetV4 variants, EfficientNet-Lite, and ShuffleNet are strong starting points. Their depthwise separable or hardware-conscious architectures reduce computation before quantization. INT8 versions are commonly suitable for Android devices, embedded Linux boards, and camera gateways, provided the selected runtime supports their operators.
For object detection, consider YOLOv5n, YOLOv8n, YOLO11n, and other nano-sized YOLO variants, along with SSD-MobileNet and NanoDet. A quantized nano detector can provide useful real-time performance on an edge GPU, NPU, or modern CPU, but accuracy depends heavily on input resolution, camera quality, and the local training data. For Indian deployments, test on real scenes: low light, crowded roads, local scripts on signage, monsoon conditions, and camera angles are often more important than a leaderboard score.
For segmentation, Mobile-UNet, Fast-SCNN, and lightweight DeepLab variants can work well when pixel-level output is required. These models are relevant to crop monitoring, industrial inspection, road analysis, and medical pre-screening. Because segmentation produces large output tensors, measure RAM usage and post-processing time—not only the neural network’s MAC count. Teams building custom vision pipelines can use this computer vision model development guide before selecting a quantization strategy.
Quantized speech and language models
Keyword spotting models such as DS-CNN, MobileNet-based audio classifiers, and compact Conformer variants are well suited to always-on microphones. INT8 inference can reduce energy use, which matters for battery-powered devices and local voice controls. For full speech-to-text, small Whisper variants, Vosk, and hardware-optimised streaming models may run at the edge, but latency depends on audio length, beam search, and whether decoding is accelerated.
For text generation, compact models such as SmolLM, Qwen2.5 small variants, Gemma 3 small variants, and Llama 3.2 small variants can be deployed in 4-bit or 8-bit formats on capable phones, laptops, gateways, and embedded computers. These models are not equivalent to cloud-scale systems: context length, multilingual quality, tool use, and reasoning are constrained. For Hindi and other Indian languages, evaluate language-specific quality rather than assuming that a smaller English benchmark transfers. Compare local models using the guidance in open-source small language models for Hindi.
A practical edge language stack often combines a small quantized model with retrieval, deterministic business rules, and cloud escalation for difficult requests. This keeps routine queries local while preserving a fallback for complex tasks. If the product requires fully offline operation, test prompt templates, token throughput, first-token latency, and peak memory under the longest supported input.
Match the model to the hardware
The same ONNX, TFLite, or GGUF file can perform very differently across devices. Identify the target processor first:
- Android phones and tablets: TensorFlow Lite, LiteRT, ONNX Runtime Mobile, and vendor delegates can use CPU, GPU, or NNAPI/NPU paths.
- Apple devices: Core ML models using FP16 or INT8 may use the Neural Engine, GPU, or CPU.
- Raspberry Pi and ARM Linux: ONNX Runtime, TFLite/LiteRT, NCNN, and vendor-specific runtimes are common; an accelerator can change the result substantially.
- NVIDIA Jetson: TensorRT, often with FP16 or INT8 calibration, is a strong option for vision and multimodal workloads.
- Qualcomm, MediaTek, and other NPUs: Use the manufacturer’s converter and delegate where possible, then verify that every critical operator remains on the accelerator.
- Microcontrollers: CMSIS-NN, TFLite Micro, and vendor SDKs typically favour very small INT8 models, especially for keyword spotting and sensor classification.
For a broader deployment checklist, see this AI model optimization guide for mobile devices. It covers packaging, profiling, and the operational details that are often missed after model training.
A practical selection and benchmarking process
Start with an accuracy baseline in FP32, then export the exact architecture supported by your chosen runtime. Build a representative calibration set: for an Indian retail camera, that may include regional products, scripts, lighting, skin tones, clothing, and crowded scenes. Compare FP16, INT8, and—where appropriate—INT4 versions.
Record at least:
- Model file size and peak resident memory
- Cold-start time and steady-state latency
- Frames per second or tokens per second
- Energy per inference and thermal throttling
- Accelerator utilisation and CPU fallback percentage
- Accuracy by important demographic, language, location, and lighting slices
- Failure behaviour when the network is unavailable
Use real device measurements rather than desktop emulation. A model that meets latency targets for one camera stream may fail when four streams share the same gateway. Similarly, a language model that fits in RAM may become unusable once the operating system, vector index, and application are loaded.
Common failure modes
Accuracy loss is the most visible problem. Improve calibration data, exclude sensitive layers from aggressive quantization, or use QAT. Unsupported operators are another frequent issue; simplify the graph or select a model family designed for the target accelerator. Memory spikes can occur during conversion or loading, even when the final file is small. Profile peak memory and use streaming or memory-mapped loading where supported.
Treat privacy and reliability as design requirements. Local inference can keep voice, health, agricultural, or customer data on the device, but model updates still need signing, rollback, access control, and monitoring. For regulated use cases, quantized predictions should be validated against the same safety and audit requirements as cloud inference.
Bottom line
The best quantized model is the smallest model that meets your accuracy, latency, memory, energy, and reliability targets on the actual edge hardware. For vision, begin with MobileNet, EfficientNet-Lite, SSD-MobileNet, or a nano-sized YOLO. For audio, test DS-CNN, compact Conformer, or a small streaming recogniser. For language, evaluate 4-bit or 8-bit small language models with Indian-language data and a realistic context window.
Choose the runtime and accelerator alongside the model, benchmark end to end, and keep a cloud fallback only where the product genuinely needs it. This approach produces faster, more private, and more deployable edge AI than selecting a model by parameter count alone.