0tokens

Apply for AI Grants India

Financial support for innovators building the future of AI in India.

Apply now

Chat · deploying open source llms on local hardware

Deploying Open-Source LLMs on Local Hardware

  1. aigi

    Running a language model on your own workstation is no longer limited to research labs. In 2026, an Indian developer can deploy capable open-source models on a laptop, desktop GPU, or small office server and expose them through a private API. The right setup can reduce recurring inference costs, keep sensitive data inside your network, and make experimentation faster.

    The trade-off is operational responsibility. You must select a model that fits available memory, understand its licence, measure quality on your data, and secure the endpoint. This guide explains how to move from model selection to a reliable local service.

    Start with the workload, not the model

    Define the job before downloading weights. A model used for document search, customer-support drafts, coding assistance, or Hindi summarisation may have very different requirements.

    Record four constraints:

    • Inputs: text only, images, audio, or long documents.
    • Quality: factual accuracy, tool use, coding ability, or fluency in Indian languages.
    • Latency: interactive chat may need a fast first token; batch extraction can tolerate slower processing.
    • Concurrency: one developer, a classroom, an internal team, or a customer-facing product.

    For Indic-language products, test the exact languages and scripts you need rather than relying on an English benchmark. Resources such as low-resource Indic NLP guidance can help identify tokenisation, evaluation, and dataset risks early.

    Also check the model card and licence. “Open source” is used loosely in AI: some releases provide weights but impose restrictions on commercial use, redistribution, or high-risk applications. Keep the licence, model revision, and checksum in your deployment records.

    Hardware: plan around memory

    For inference, memory capacity is usually more important than raw CPU speed. GPU VRAM holds model weights, the KV cache for conversation history, and runtime overhead. Longer prompts and more simultaneous users require additional memory.

    A rough starting point for 4-bit inference is:

    • 3B–4B models: 6–8 GB VRAM, or modern CPU-only systems for slower workloads.
    • 7B–8B models: 8–12 GB VRAM for short contexts; 16 GB gives more headroom.
    • 12B–14B models: generally 16–24 GB VRAM, depending on context length.
    • 30B–34B models: around 24–48 GB VRAM, often with quantisation or multiple GPUs.
    • 70B models: 48–96 GB or more, commonly across several GPUs or with CPU offload.

    These are planning estimates, not guarantees. A 4-bit file advertised as 5 GB still needs space for runtime buffers and the KV cache. Leave 10–20% capacity unused where possible.

    NVIDIA GPUs remain the easiest route on Linux because CUDA support is broad. AMD GPUs can work with ROCm-compatible stacks, while Apple Silicon uses unified memory and Metal acceleration. For CPU inference, prioritise modern cores, adequate RAM, and an NVMe SSD. System RAM should normally exceed the model file size by a comfortable margin; 32 GB is a practical baseline for experimenting with 7B–14B models.

    Quantisation and model formats

    Quantisation stores weights at lower numerical precision, reducing memory use and often improving practicality on consumer hardware. For most developers, 4-bit quantisation is the first option to test, not an automatic guarantee of unchanged quality.

    • GGUF: A flexible format for llama.cpp-based tools, including CPU inference and partial GPU offload.
    • EXL2: Designed for efficient NVIDIA GPU inference, especially when the model fits in VRAM.
    • AWQ and GPTQ: Common choices for GPU serving stacks such as vLLM and other CUDA runtimes.
    • FP16 or BF16: Higher memory use, but useful when quality, fine-tuning, or maximum throughput matters.

    Compare quantised variants on representative prompts. Test factuality, JSON validity, code execution, refusal behaviour, and Indic-language output. A smaller model that answers reliably may be more useful than a larger model that constantly spills into RAM.

    Choose an inference engine

    Ollama is a strong starting point for a single developer or private application. It manages model downloads, exposes a local API, and works across common desktop platforms. A basic setup is:

    curl -fsSL https://ollama.com/install.sh | sh
    ollama run qwen2.5:7b

    Use the API from Python or any HTTP client:

    import requests
    
    payload = {
        "model": "qwen2.5:7b",
        "prompt": "Summarise this policy in Hindi and English.",
        "stream": False,
    }
    response = requests.post(
        "http://localhost:11434/api/generate",
        json=payload,
        timeout=120,
    )
    response.raise_for_status()
    print(response.json()["response"])

    llama.cpp gives more control over GGUF files, CPU threads, GPU layers, context size, and batching. It is useful for constrained hardware and reproducible experiments.

    vLLM is better suited to a Linux server, multiple users, and OpenAI-compatible serving. Its paged KV-cache management and batching can deliver substantially higher throughput, but it requires more attention to CUDA versions, model compatibility, memory limits, and monitoring. If you are building a larger service, pair deployment knowledge with high-performance open-source AI application practices.

    LM Studio is convenient for visual model discovery and local API testing. Treat desktop interfaces as development tools; production services still need authentication, logging controls, process supervision, and updates.

    Make local RAG useful

    A local model alone does not know your latest policies or documents. For internal search, add a retrieval-augmented generation pipeline: parse documents, create embeddings, store chunks in a vector index, retrieve relevant passages, and instruct the model to answer only from the supplied evidence.

    Keep retrieval and generation separate so you can evaluate them independently. Store document version, source, access permissions, and timestamps with each chunk. Never assume that local inference automatically makes a system private: documents may still be exposed through logs, browser interfaces, backups, or an unsecured network port.

    Security, privacy, and operations

    Bind development services to localhost by default. If another device must connect, place the service behind a private network, VPN, or authenticated reverse proxy. Do not expose Ollama or a raw vLLM port directly to the public internet.

    Use these controls:

    • Redact secrets and personal data from prompts and logs.
    • Restrict model and document directories by operating-system permissions.
    • Pin model versions and record configuration changes.
    • Set request, context, and output-token limits.
    • Monitor GPU temperature, VRAM use, latency, queue depth, and error rates.
    • Add a health check and restart policy for unattended servers.
    • Maintain a fallback path for outages or workloads beyond local capacity.

    For agentic systems, local model serving is only one layer. Review the production deployment checklist for open-source AI agents before granting tools access to files, databases, payments, or external APIs.

    Performance tuning and troubleshooting

    Benchmark with your real prompt lengths and concurrency. Measure time to first token, tokens per second, total latency, memory use, and answer quality. A short benchmark prompt can hide failures caused by long context windows.

    If generation is slow, check whether layers have spilled from VRAM to RAM, whether the context window is excessive, and whether thermal throttling is occurring. Reduce context length, use a smaller quantisation, or choose a smaller model before buying hardware. If CUDA errors appear, verify driver and runtime compatibility and use a pinned container where practical.

    For Indian teams, calculate total cost rather than comparing only token prices. Include GPU depreciation, electricity, cooling, storage, maintenance, engineering time, and the value of keeping data on-premises. Local inference is especially compelling for predictable internal workloads, offline environments, and repeated high-volume requests; hosted APIs may still win for occasional use or frontier reasoning.

    A practical deployment path

    1. Start with a 7B–8B instruct model in GGUF or an Ollama-supported format.
    2. Evaluate it on 50–100 real, anonymised examples.
    3. Add retrieval if answers depend on private or changing information.
    4. Measure quality and latency at the expected context length.
    5. Move to vLLM or a managed service only when concurrency justifies the complexity.
    6. Document licences, model versions, hardware, prompts, and evaluation results.

    Teams building locally can also learn from Indian open-source AI developer projects and adapt proven practices instead of treating every deployment as a one-off experiment. With disciplined evaluation and basic operations, local hardware becomes a practical foundation for private AI—not merely a demo environment.

    Last updated 23 September 2026

AIGI may be inaccurate. Replies seeded from the guide above.