Running a small language model in a browser is now practical for focused tasks such as text classification, summarisation, autocomplete, translation, and local chat. The key is to choose a model that fits the browser’s memory and compute limits, then load it through a runtime designed for JavaScript rather than trying to serve a desktop-sized model unchanged.
For Indian builders, browser inference can reduce hosting costs and improve privacy in tools that handle customer messages, internal documents, or Indic-language text. It is particularly useful for prototypes, offline-first products, education tools, and applications that need a fast response after the model has been downloaded.
What “small” means in a browser
A browser model is constrained by more than parameter count. You must account for:
- Weights: A quantised model may be tens or hundreds of megabytes, while an unquantised copy can be several times larger.
- Memory overhead: Runtime buffers, tokenisation, intermediate tensors, and multiple browser tabs consume additional RAM.
- Compute: WebGPU is generally preferable to CPU execution, but support and performance vary by device and browser.
- Task complexity: Classification and extraction are far easier to run locally than long-context text generation.
- Language coverage: Test Hindi and other Indic languages rather than assuming an English-centric model will transfer well. For background, see this guide to low-resource Indic natural language processing.
A small encoder model such as a compact BERT variant may work well for classification, while a decoder model is needed for generation. Do not select a model solely because it has a low parameter count: tokenizer quality, context length, quantisation support, and licensing matter just as much.
Choose the browser runtime
In 2026, the most straightforward route for many projects is Transformers.js, which runs compatible models through ONNX Runtime Web and can use WebGPU where available. It supports common NLP pipelines and avoids writing tensor-management code from scratch.
Other options are useful in specific cases:
- ONNX Runtime Web: A lower-level choice when you need control over graph execution, providers, and model packaging.
- TensorFlow.js: Suitable when your model is already in TensorFlow.js format or when you are building a custom JavaScript ML pipeline.
- WebLLM or WebGPU-native runtimes: Useful for supported generative language models and chat-style interfaces.
- WASM fallback: Important for devices without WebGPU, though usually slower for generation.
If your application includes mobile inference, review AI model optimisation for mobile devices because the same principles—quantisation, memory limits, batching, and thermal constraints—apply in the browser.
A practical Transformers.js setup
Create a small Vite project and install the runtime:
npm create vite@latest browser-model -- --template vanilla
cd browser-model
npm install
npm install @huggingface/transformers
npm run devThe following example performs sentiment classification. Replace the model with one that supports your language and task, and verify its licence before shipping it:
import { pipeline, env } from '@huggingface/transformers';
// Use a public model host during development. Configure a local model path for production if required.
env.allowLocalModels = false;
env.backends.onnx.wasm.numThreads = 1;
const classify = await pipeline(
'sentiment-analysis',
'Xenova/distilbert-base-uncased-finetuned-sst-2-english',
{ device: 'webgpu' }
);
document.querySelector('#run').addEventListener('click', async () => {
const text = document.querySelector('#input').value.trim();
if (!text) return;
const output = await classify(text);
document.querySelector('#output').textContent = JSON.stringify(output, null, 2);
});Add a minimal interface in index.html:
<textarea id="input" rows="5" placeholder="Enter text"></textarea>
<button id="run">Analyse</button>
<pre id="output" aria-live="polite"></pre>
<script type="module" src="/src/main.js"></script>For a Hindi or multilingual application, use a model trained for that purpose rather than the English sentiment model above. The open-source small language models for Hindi guide is a useful starting point for comparing options.
Make model loading reliable
The first inference is often slow because the browser must download, cache, compile, and initialise the model. Treat loading as a product state, not an implementation detail.
- Show download and initialisation progress.
- Disable the submit button until the pipeline is ready.
- Cache model assets using the browser Cache API or a service worker.
- Serve files over HTTPS; WebGPU and service workers require secure contexts in most deployments.
- Pin model revisions instead of silently pulling a changing
mainbranch. - Keep model files on a CDN with correct CORS headers and long-lived cache headers.
- Consider an explicit “Download for offline use” action for large models.
Do not reload the model on every button click. Initialise it once, retain the pipeline, and reuse it for subsequent inputs. On low-memory phones, release unused objects and avoid keeping long conversation histories in memory.
WebGPU, WASM, and graceful fallback
Feature-detect WebGPU rather than assuming it exists:
const supportsWebGPU = 'gpu' in navigator;
const device = supportsWebGPU ? 'webgpu' : 'wasm';
const model = await pipeline('sentiment-analysis', MODEL_ID, { device });A fallback is essential for Indian users on varied hardware and networks. WebGPU may be unavailable, blocked by policy, or unstable on an older browser. WASM can still support classification and short inputs, but generation may be too slow for an acceptable experience.
Set realistic limits: cap input length, truncate safely at token boundaries, and avoid promising real-time generation on entry-level devices. Measure time to first token, tokens per second, peak memory, and model download size on representative Android phones—not only on a developer laptop.
Privacy, security, and responsible deployment
Local inference reduces the need to send prompts to your server, but it does not automatically make an application private. A webpage can still collect inputs through analytics, logs, error reporting, or third-party scripts. Document what is processed locally and audit network requests.
Also consider:
- Model licences and restrictions on commercial use.
- Prompt injection when model output is inserted into HTML or commands.
- XSS risks: render generated text as plain text unless sanitisation is deliberate.
- Sensitive data remaining in browser storage or crash reports.
- Model extraction: public model files can be downloaded by users.
- Quality and safety failures in Hindi, Hinglish, and regional languages.
If the model must be adapted to a regional language, browser inference is only the final delivery layer. Training and evaluation may require server-side infrastructure; see fine-tuning Llama for Indian regional languages for the broader workflow.
Testing checklist before launch
Test at least three browser classes: desktop Chrome or Edge with WebGPU, Safari where supported, and an Android device using the expected network conditions. Record:
- Initial download size and repeat-visit load time.
- Time to first result and steady-state latency.
- Peak RAM and behaviour under multiple tabs.
- Accuracy on real, noisy user inputs.
- Hindi, English, Hinglish, and relevant regional-language examples.
- Offline behaviour after the model is cached.
- Fallback behaviour when WebGPU fails.
- Accessibility of progress, errors, and generated output.
Start with a narrow task and a compact quantised model. Browser-based AI works best when it is designed around the device, language, and job—not when a large server model is compressed at the last minute.