Quantized models make it possible to run useful AI directly in a browser tab instead of sending every input to a cloud API. For Indian product teams, this can reduce recurring inference costs, improve responsiveness on uneven networks, and keep sensitive inputs—such as documents, voices, and camera frames—on the user’s device.
The practical answer to which quantized models can run in browser depends on four variables: model size, operator support, device memory, and the browser runtime. A small INT8 vision model may run smoothly on a mid-range Android phone, while even a 4-bit language model can struggle if the device lacks enough RAM or WebGPU support.
What quantization changes
Quantization represents model weights and, sometimes, activations with fewer bits. FP32 uses 32-bit floating-point values; FP16 uses 16 bits, while INT8 and INT4 use eight and four bits respectively. Lower precision generally means smaller downloads, lower memory use, and faster inference—but it can also reduce accuracy or create compatibility issues.
Common approaches include:
- Post-training quantization: Convert a trained model after development. It is the fastest route for deployment and often works well for classification and embedding models.
- Quantization-aware training: Train while simulating reduced precision. This usually preserves accuracy better, especially for sensitive vision or speech tasks.
- Weight-only quantization: Compresses weights while leaving activations at higher precision. This is common for small language models and can be easier to support in browser runtimes.
Quantization is not the same as pruning or distillation. Distillation creates a smaller student model, while quantization changes numerical representation. Combining both can deliver better browser performance than quantizing a large model alone.
Model families that work well in browsers
Vision models: the safest starting point
Compact convolutional networks remain the most reliable browser deployments. MobileNetV3, EfficientNet-Lite, SqueezeNet, and small ResNet variants are suitable for image classification, quality checks, and on-device tagging. For detection, lightweight YOLO variants, SSD MobileNet, and NanoDet can work when exported with supported operators.
These models are useful for camera-based inspection, retail cataloguing, crop or plant screening, and document pre-processing. Teams building their own pipeline can use this guide to build computer vision models on GitHub, then export and benchmark the model in a browser rather than assuming server performance will translate to mobile hardware.
For pose estimation and segmentation, reduced versions of MoveNet, BlazePose, and MediaPipe models are practical choices. Keep input dimensions modest—224px or 320px is often a better starting point than 640px—and process every second or third video frame if real-time inference is not essential.
Speech and audio models
Browser audio inference is feasible, but microphone permissions, model download size, and background noise matter as much as raw model accuracy. Small keyword-spotting networks, VAD models, and compact speech encoders can run using WebAssembly or WebGPU. Whisper-style automatic speech recognition models are possible, but browser deployments should favour tiny or base-sized variants, aggressive quantization, and chunked transcription.
For Indian applications, test Hindi, Tamil, Telugu, Marathi, Bengali, and mixed English speech separately. A model that performs well on clean English audio may fail on code-switching, regional accents, or noisy phone recordings. If the product needs a conversational interface, first decide whether a voice agent or chatbot is better for the business; a browser speech layer can handle capture and transcription while more demanding reasoning remains server-side.
Text classifiers and embeddings
Small transformer encoders are more browser-friendly than generative language models. MiniLM, DistilBERT, ALBERT, and compact multilingual BERT variants can support sentiment analysis, intent classification, language identification, spam detection, and semantic search. INT8 or mixed-precision versions are generally easier to deploy than large 4-bit decoder models.
A useful architecture is to run classification, redaction, or retrieval locally and call a server only when generation is required. For Indian-language products, evaluate tokenisation carefully: a model may be small but still inefficient if it fragments Devanagari or another script excessively. Compare results against benchmarking NLP models for Telugu and Sanskrit before selecting a multilingual encoder.
Small generative language models
Small language models can run in browsers through WebGPU-enabled runtimes, but the experience varies sharply by device. Decoder models in the roughly 0.5B–3B parameter range are the most realistic targets. 4-bit weight quantization reduces download and memory requirements, while WebGPU can accelerate matrix operations. However, a 1B model may still require substantial memory once the runtime, key-value cache, tokenizer, and browser overhead are included.
Use compact models for autocomplete, structured extraction, short rewriting, offline FAQ search, and constrained classification—not unrestricted long-form reasoning. Limit context length, stream tokens, and provide a server fallback. Developers choosing Hindi or other Indian-language models can start with this overview of open-source small language models for Hindi, then verify licence, tokenizer coverage, and performance on actual target phones.
Browser runtimes and formats
The model format often determines whether deployment succeeds. Current options include:
- TensorFlow.js: A mature option for JavaScript and converted TensorFlow models, with WebGL, WebGPU, and CPU backends.
- ONNX Runtime Web: Useful for ONNX models, with WebAssembly, WebGL, and WebGPU execution depending on operator support.
- Transformers.js: Convenient for many encoder and smaller generative models, using ONNX-based browser execution.
- MediaPipe Tasks: Strong for supported vision, audio, and gesture workloads with ready-made browser APIs.
- WebLLM and similar WebGPU runtimes: Designed for running compatible language models with browser-side GPU acceleration.
WebAssembly is usually the compatibility baseline; WebGPU is the performance target. Do not assume that a model converted successfully will run efficiently. Unsupported operators may trigger CPU fallbacks, and one fallback layer can dominate total latency.
How to choose and benchmark a model
Start with a deployment budget, not a model leaderboard. Define maximum download size, first-result latency, sustained tokens or frames per second, peak memory, and acceptable accuracy. Then test on the actual devices used by Indian customers, including budget Android phones and low-bandwidth connections.
A practical benchmark should measure:
- Compressed model size and cache behaviour
- Cold-start time after a fresh page load
- Warm inference latency and throughput
- Peak RAM and GPU memory where available
- Battery drain during a realistic session
- Accuracy before and after quantization
- Behaviour when WebGPU is unavailable
Serve model files with compression, cache immutable versions, load them after critical UI content, and show progress for large downloads. For privacy-sensitive use cases, explain clearly that inputs remain on-device; browser storage and permissions still require careful handling.
Common failure modes
The largest model that fits is rarely the best product choice. A quantized model can still be too slow because of a long context window, inefficient tokenizer, unsupported operation, or repeated tensor copies between JavaScript and the accelerator. Accuracy can also degrade disproportionately for minority languages, rare names, medical terms, and noisy images.
Use progressive enhancement: ship a compact local model first, detect browser capabilities, and fall back to a server endpoint or a simpler heuristic when necessary. For workloads that genuinely need larger models, compare this approach with deploying large language models locally or a controlled cloud deployment.
Recommended decision rule
Choose quantized MobileNet, MediaPipe, or small YOLO models for browser vision; MiniLM or DistilBERT-class encoders for text classification and embeddings; tiny or base speech models for constrained transcription; and sub-3B, 4-bit language models only when WebGPU coverage and device testing justify them.
The winning browser model is not simply the smallest file. It is the model that meets accuracy, latency, privacy, and battery requirements on the devices your users actually own. Benchmark that complete experience before committing to a browser-first architecture.