0tokens

Apply for AI Grants India

Financial support for innovators building the future of AI in India.

Apply now

Chat · how to deploy large language models locally

How to Deploy Large Language Models Locally

  1. aigi

    Local LLM deployment is no longer limited to research labs. A capable workstation can run small and medium open-weight models for chat, coding, document search, translation, and internal automation—without sending every prompt to a third-party API. For Indian startups, this can improve privacy, predictable costs, offline resilience, and latency while supporting applications that handle financial, health, legal, or proprietary data.

    The right setup depends on the workload, not only the model’s parameter count. A developer testing an 8B model needs a very different stack from a team serving thousands of concurrent requests or building a multilingual voice product. This guide explains how to deploy large language models locally in 2026, with practical choices for hardware, runtimes, quantisation, APIs, evaluation, and operations.

    Start with the workload

    Define these constraints before downloading a model:

    • Task: chat, extraction, coding, summarisation, RAG, translation, or agent workflows.
    • Quality target: accuracy, reasoning, tool use, structured JSON, and language coverage.
    • Latency: interactive response time versus batch processing.
    • Concurrency: one developer, a small internal team, or a public product.
    • Data boundary: fully offline, private VPC, or approved external APIs.
    • Context length: short prompts, document-scale context, or long-running sessions.

    A smaller model with a strong prompt, retrieval pipeline, and structured output can outperform a larger model used without evaluation. If your product serves Indian-language users, test the actual languages and scripts rather than assuming that English benchmarks transfer. For background on compact Hindi models, compare this practical guide to open-source small language models for Hindi.

    Choose hardware by memory and throughput

    VRAM is usually the first constraint. In approximate terms, model weights alone require 2 bytes per parameter at FP16, 1 byte at 8-bit, and 0.5 bytes at 4-bit. Runtime overhead, the KV cache, framework buffers, and the operating system require additional memory.

    A practical starting point is:

    • 7B–9B models: 8–12 GB VRAM at 4-bit; 16 GB is more comfortable.
    • 14B–作 models: typically 12–24 GB VRAM at 4-bit, depending on context and runtime.
    • 30B–34B models: commonly 20–32 GB or more at 4-bit.
    • 70B models: generally require multiple GPUs, substantial system RAM, or careful CPU offload.

    The unusual character above should not be present; use 14B–20B models instead. Consumer GPUs such as 12–16 GB cards are suitable for small models, while 24 GB cards offer much more flexibility. CPU inference is viable with sufficient RAM and a GGUF model, but expect lower throughput. For production, account for power, cooling, storage, and replacement logistics—not just GPU price. In India, a UPS and adequate airflow matter during voltage fluctuations and hot seasons.

    Keep at least 1.5–2 times the model file size available in fast storage for downloads, conversions, caches, and updated versions. NVMe SSDs make model loading and container startup considerably less painful.

    Pick the right runtime

    Ollama for local development

    Ollama is the simplest route for a developer or small team. It manages model files and exposes a local HTTP endpoint. After installation, a typical test looks like:

    ollama pull llama3.1:8b
    ollama run llama3.1:8b

    Its OpenAI-compatible integrations make it useful for prototypes, RAG experiments, and internal tools. Do not treat the default configuration as a production control plane: add authentication, network restrictions, logging, and resource limits before exposing it beyond localhost.

    llama.cpp for efficient CPU and hybrid inference

    llama.cpp runs GGUF models across CPU, CUDA, Metal, and other backends. It is a strong choice for offline applications, edge devices, and machines with limited VRAM. It also provides a server mode and granular control over GPU layers, context size, batching, and threads.

    vLLM for GPU serving

    Use vLLM when throughput, batching, and an OpenAI-compatible API matter. Its paged-attention design helps serve multiple requests efficiently on NVIDIA GPUs. A representative launch command is:

    python -m vllm.entrypoints.openai.api_server \
      --model Qwen/Qwen2.5-7B-Instruct \
      --dtype auto \
      --max-model-len 8192

    The exact entry point and flags can change between releases, so pin a tested version and follow its current documentation. For high-volume deployments, also evaluate SGLang or Hugging Face TGI rather than selecting a server solely by popularity.

    Desktop interfaces

    LM Studio and similar applications are useful for model comparison, prompt testing, and demonstrations. They are not substitutes for a reproducible deployment process. Export the chosen model, prompt templates, settings, and evaluation results into version-controlled configuration.

    Understand quantisation and model formats

    Quantisation reduces weight precision to lower memory use and often improve speed. Common choices include:

    • GGUF: broadly supported by llama.cpp and desktop runtimes; suitable for CPU or mixed CPU/GPU inference.
    • GPTQ and AWQ: GPU-oriented quantisation formats used by several serving stacks.
    • bitsandbytes 8-bit or 4-bit: convenient when loading Transformers models directly.
    • FP16 or BF16: higher memory use but often preferable when quality, fine-tuning, or maximum compatibility matters.

    Do not choose the smallest file automatically. Quantisation can affect tool calling, multilingual output, factuality, and long-context behaviour. Test representative prompts at the target context length. Save the model revision, quantisation method, calibration details, tokenizer, and license alongside your application so another engineer can reproduce the result.

    A production-ready deployment path

    A reliable local deployment usually follows this sequence:

    1. Select an instruction-tuned model with a license compatible with your product and geography.
    2. Download from a trusted registry and record the revision or checksum.
    3. Benchmark locally using real prompts, expected JSON schemas, Indian languages, and failure cases.
    4. Serve behind an internal API using vLLM, llama.cpp, or another pinned runtime.
    5. Add an application layer for authentication, rate limits, prompt templates, retrieval, validation, and retries.
    6. Containerise the stack where practical, pin CUDA and driver compatibility, and separate model storage from ephemeral containers.
    7. Monitor quality and operations: latency, tokens per second, queue depth, GPU memory, errors, refusal patterns, and sensitive-data exposure.
    8. Create rollback procedures for model, tokenizer, runtime, and prompt changes.

    A minimal OpenAI-compatible client can point at a local endpoint:

    from openai import OpenAI
    
    client = OpenAI(
        base_url="http://localhost:8000/v1",
        api_key="local-only"
    )
    
    response = client.chat.completions.create(
        model="Qwen/Qwen2.5-7B-Instruct",
        messages=[{"role": "user", "content": "Summarise this text in five bullets."}],
        temperature=0.2,
    )
    print(response.choices[0].message.content)

    For agent products, local model serving is only one component. Tool permissions, sandboxing, state management, and audit logs are equally important; see this guide to deploying open-source AI agents in production.

    Retrieval, multilingual use, and safety

    For private documents, prefer a RAG pipeline over putting an entire corpus into a prompt. Keep ingestion, embeddings, retrieval, generation, and access control as separate components. Apply document-level permissions before retrieval, redact secrets where possible, and log references rather than raw sensitive text.

    Indian-language applications require dedicated evaluation for code-switching, transliteration, named entities, and regional terminology. Consider language-specific fine-tuning only after measuring whether prompting and retrieval are insufficient. This guide to fine-tuning Llama for Indian regional languages covers the main trade-offs.

    Never assume local means automatically safe. Protect model endpoints with network policies and authentication, encrypt disks and backups, restrict shell access, scan uploaded files, and prevent untrusted text from controlling tools. Establish retention rules for prompts, outputs, traces, and user identifiers.

    Performance tuning checklist

    • Reduce the context window to the smallest value that meets the task.
    • Use continuous batching for concurrent GPU workloads.
    • Enable FlashAttention or the runtime’s equivalent when supported.
    • Measure prompt processing and generation speed separately.
    • Use deterministic settings for extraction and structured outputs.
    • Reserve memory for the KV cache instead of filling VRAM with weights alone.
    • Test CPU offload only after measuring its latency impact.
    • Batch offline jobs rather than optimising an interactive server for them.
    • Use tensor parallelism only when one GPU cannot hold the model or throughput requires it.

    For voice interfaces, streaming tokens and interruption handling are as important as raw generation speed. A local LLM can sit behind a voice pipeline; the voice-agent architecture guide explains the surrounding design.

    Costs and operating decisions in India

    Local inference replaces per-token bills with capital expenditure and operations. Compare the total cost of ownership: GPU depreciation, electricity, cooling, storage, monitoring, engineering time, support, and downtime. For sporadic workloads, a managed GPU instance may be cheaper. For predictable, sustained traffic or strict data residency, owned hardware can make sense.

    Start with one measurable workload, not an oversized cluster. Record baseline cloud API cost, latency, quality, and monthly request volume. Then compare a local pilot against the same evaluation set. If demand grows, add queueing and replicas only after profiling the bottleneck.

    FAQ

    Can I run a language model without a GPU?
    Yes. Use a quantised GGUF model with llama.cpp and enough system RAM. It is practical for batch tasks and lightweight assistants, but usually slower for interactive production traffic.

    What is a sensible first model?
    Choose a current 7B–9B instruction model for general experimentation, then compare larger models or specialised coding and multilingual models against your own evaluation set. Model names and releases change quickly, so avoid treating one benchmark winner as permanent.

    Is local deployment automatically compliant?
    No. It reduces external data transfer but does not remove obligations around consent, access control, retention, security, licensing, or sector-specific regulation. Document where data is stored and who can access it.

    Should I fine-tune immediately?
    Usually not. Begin with prompting, retrieval, structured output validation, and evaluation. Fine-tune when you have a stable dataset, a measurable gap, and a clear deployment plan.

    Local LLMs are most valuable when treated as an engineering system rather than a downloaded model. Define the workload, test quality on Indian use cases, secure the endpoint, and measure total cost before scaling. For teams building Llama-based workflows, this Llama 3 agents deployment guide is a useful next step.

    AI Grants India supports Indian builders working on privacy-preserving and sovereign AI infrastructure. Explore the AI Grants India community for opportunities, resources, and support.

    Last updated 23 September 2026

AIGI may be inaccurate. Replies seeded from the guide above.