0tokens

Apply for AI Grants India

Financial support for innovators building the future of AI in India.

Apply now

Chat · deploying large language models on edge devices

Deploying Large Language Models on Edge Devices

  1. aigi

    Edge inference is no longer limited to demos. Phones, laptops, point-of-sale terminals, factory gateways, vehicles, and developer boards can now run capable small language models without sending every prompt to a cloud API. For Indian builders, this matters where connectivity is inconsistent, sensitive data cannot leave the device, and per-token cloud costs make high-volume products difficult to sustain.

    Deploying large language models on edge devices is not simply a matter of downloading a quantised checkpoint. You must match model size, context length, runtime, accelerator, memory, thermal limits, and product requirements. A good deployment is one that remains useful under real battery, network, and device conditions—not merely one that produces a strong benchmark score on a workstation.

    Start with the product constraint

    Define the workload before choosing a model. A local assistant that extracts fields from invoices has very different requirements from an offline Indic-language tutor or a voice interface for field workers.

    Document these constraints:

    • Offline requirement: Must the product work with no network, or can it use cloud fallback?
    • Latency target: Measure time to first token and sustained tokens per second separately.
    • Privacy boundary: Identify which prompts, documents, logs, and generated outputs may leave the device.
    • Memory budget: Account for model weights, runtime overhead, tokenizer, KV cache, application code, and operating-system pressure.
    • Languages and tasks: Test Hindi and other Indian languages directly; English-centric evaluations are not enough.
    • Update model: Decide how weights, prompts, safety rules, and tokenizers will be updated after installation.

    For multilingual products, model capability often matters more than parameter count. A smaller model with strong Hindi or regional-language coverage may outperform a larger general model that was not trained or evaluated adequately on the target language. Compare this with guidance on small language models for Hindi and use representative, locally collected test data where governance permits.

    Choose a feasible model size

    Model memory begins with the weights, but weights are only the first cost. A rough calculation is:

    Weight memory ≈ parameter count × bytes per parameter

    An 8-billion-parameter model requires about 16 GB at FP16, 8 GB at 8-bit, and roughly 4 GB at 4-bit before runtime overhead and the KV cache. Actual files vary because of metadata, embedding formats, and quantisation schemes. A device with 8 GB of RAM therefore cannot safely be treated as an 8 GB model host.

    As a starting point:

    • 1B–3B models: suitable for summarisation, extraction, classification, and tightly scoped assistants.
    • 7B–8B models: useful on high-end phones, laptops, and edge gateways when quantised and carefully configured.
    • 13B and above: generally better suited to machines with substantial shared memory or a split edge-cloud architecture.

    Do not use a larger model to compensate for weak retrieval, unclear prompts, or poor domain data. A smaller model with constrained output schemas and local retrieval can be faster, cheaper, and easier to operate.

    Quantisation is the primary deployment lever

    Quantisation reduces numerical precision so weights occupy less memory and move faster through the hardware. Common choices include 8-bit, 6-bit, 5-bit, and 4-bit formats. Four-bit quantisation is often the practical balance for general-purpose edge inference, but the best choice depends on the model and task.

    Evaluate quality after quantisation rather than assuming that a format is universally safe. Test:

    • Exact-match extraction and structured JSON validity
    • Factuality on domain documents
    • Hindi, English, and code-switching behaviour
    • Long-context retrieval
    • Refusal and safety behaviour
    • Repetition, truncation, and hallucination rates

    Weight-only quantisation is common, while activation-aware approaches such as AWQ can preserve quality on selected workloads. GGUF is widely used with llama.cpp; other runtimes may prefer ONNX, safetensors, or vendor-specific formats. For a broader treatment of memory planning, kernels, and mobile packaging, see this guide to AI model optimisation for mobile devices.

    Pruning and distillation can reduce the model itself, but they usually require more experimentation than quantisation. Distillation is attractive when a team owns a narrow workflow and can generate high-quality teacher outputs. Structured pruning is easier to operate when the target runtime supports the resulting architecture efficiently.

    Select the runtime around the hardware

    The runtime determines whether your compressed model actually uses the device effectively.

    • llama.cpp: A strong choice for portable C/C++ inference, GGUF models, CPU execution, and GPU backends such as Metal, CUDA, and Vulkan.
    • MLC LLM: Useful when compiling models across Android, iOS, browsers, and desktop targets through TVM-backed execution.
    • ONNX Runtime: Suitable for teams already using ONNX graphs and Windows, Android, or vendor execution providers.
    • MediaPipe LLM Inference: Practical for supported mobile application workflows, especially where Android integration is central.
    • Vendor SDKs: Qualcomm, Apple, MediaTek, NVIDIA, and Intel stacks can unlock NPUs or specialised accelerators, but increase portability and maintenance costs.

    Treat NPU support carefully. A device may advertise AI TOPS without allowing your exact model, operator set, quantisation format, or context workload to run fully on the NPU. Confirm execution-provider coverage and inspect logs for CPU fallbacks. A partially accelerated graph can be slower than a well-optimised CPU or GPU path.

    Design the inference pipeline, not just the model

    A production edge system normally includes a tokenizer, prompt template, conversation store, retrieval layer, safety checks, streaming UI, and telemetry. Each component consumes memory or battery.

    Keep context bounded. The KV cache grows with sequence length and can become the dominant memory cost during long conversations. Use rolling summaries, retrieval of only relevant passages, and explicit maximum-token limits. For offline document assistants, local embeddings and a compact vector index can keep source material on-device; avoid loading entire documents into every prompt.

    Use constrained decoding where possible. JSON schemas, grammars, stop sequences, and task-specific prompts reduce retries and make smaller models more reliable. If the product supports voice, separate speech recognition, language-model inference, and speech synthesis budgets rather than assuming one accelerator will handle the complete pipeline.

    Benchmark real devices

    Desktop benchmarks are poor predictors of mobile performance. Build a device matrix covering your actual users: flagship Android phones, mid-range phones, Apple devices, Windows AI PCs, and the gateways or boards used in the field.

    Record:

    • Time to first token and tokens per second
    • Peak and steady-state RAM usage
    • Battery drain and energy per request
    • Temperature and performance after 5–15 minutes
    • Crash, timeout, and out-of-memory rates
    • Quality by language, task, and prompt length
    • Cold-start time and model download size

    Test with screen-off conditions, low battery, background applications, weak storage, and intermittent connectivity. Thermal throttling can make a fast first response unusable for a long session. Cache the model securely, support resumable downloads, and consider separate packages for different hardware tiers.

    Build a hybrid edge-cloud fallback

    Fully offline inference is valuable, but it need not be the only mode. A robust Indian deployment can use three tiers: a small local model for routine tasks, a larger local model on capable hardware, and a cloud endpoint for requests that exceed local quality or latency limits.

    Route based on explicit signals such as device memory, battery, network availability, language, and task sensitivity. Never silently upload a private prompt as a fallback. Show the user what is processed locally, obtain appropriate consent for cloud escalation, encrypt transport, and minimise retained logs. These practices support privacy obligations, including responsibilities relevant to the Digital Personal Data Protection framework, without treating compliance as a substitute for sound engineering.

    A practical launch checklist

    Before release, confirm that you can:

    • Reproduce the model and quantisation build from pinned artefacts
    • Verify checksums and sign model packages
    • Delete or rotate sensitive prompts and local indexes
    • Update models without breaking the application schema
    • Detect unsupported hardware and select a safe fallback
    • Measure quality on Indian languages and domain workflows
    • Explain offline, local, and cloud processing to users
    • Roll back a model that causes regressions

    For teams building broader local inference systems, the related guide on deploying large language models locally covers workstation and server patterns that complement edge deployments. If your application includes multimodal inputs, evaluate whether a compact vision-language model is necessary rather than adding a large model by default; open-source vision-language models for Indian languages is a useful starting point.

    The winning edge deployment is rarely the one with the biggest model. It is the one that meets a defined task, stays responsive under thermal and memory pressure, protects user data, and has a measured fallback when the hardware cannot deliver. Start with a narrow workflow, benchmark on devices your users actually own, and expand only after quality and operations are stable.

    Last updated 23 September 2026

AIGI may be inaccurate. Replies seeded from the guide above.