0tokens

Apply for AI Grants India

Financial support for innovators building the future of AI in India.

Apply now

Chat · how to deploy small language models locally

How to Deploy Small Language Models Locally

  1. aigi

    Small language models (SLMs) make useful AI possible without sending every prompt to a cloud API. A 1B–8B model can power document search, classification, extraction, coding assistance, voice workflows, and internal copilots on a laptop, workstation, or edge server. For Indian startups and institutions handling customer, health, financial, or government data, local inference can also simplify privacy and connectivity decisions.

    This guide explains how to deploy small language models locally in 2026: how to size hardware, select model formats, use Ollama or llama.cpp, expose a local API, add retrieval, and evaluate quality before putting a system in front of users.

    Start with the workload, not the model

    Define the task before downloading weights. Local models are particularly effective when the output is constrained and the knowledge source is known:

    • Extract fields from invoices, forms, or support tickets.
    • Classify messages by intent, language, urgency, or risk.
    • Summarise internal documents and meeting notes.
    • Draft responses using a controlled knowledge base.
    • Run lightweight coding, translation, or transcription post-processing.

    A local SLM is not automatically the best choice for open-ended reasoning. Benchmark the exact task using representative Indian English and Indic-language examples. If your product needs a voice interface, treat speech recognition, the language model, and text-to-speech as separate components; the architecture principles in this voice agent deployment guide are useful here.

    Choose hardware realistically

    Model size is only one part of memory use. You also need space for the runtime, the operating system, the context window, and temporary buffers. Longer prompts and larger batch sizes increase memory requirements.

    As a practical starting point:

    • 1B–3B models: 8GB system RAM is workable for quantised CPU inference; 16GB is more comfortable.
    • 7B–8B models: Plan for 16GB system RAM on CPU, or roughly 6–10GB of available VRAM for many 4-bit GPU builds.
    • 13B-class models: Usually require a workstation GPU, substantial unified memory, or CPU offloading with reduced speed.
    • Storage: Use an SSD. Keep at least 10–20GB free for model variants, caches, and logs.

    On Apple Silicon, unified memory can work well for local inference. On Linux or Windows, an NVIDIA GPU with correctly installed CUDA support generally provides higher throughput. CPU inference remains viable for low-volume batch jobs and privacy-sensitive deployments, but measure tokens per second rather than assuming performance.

    Select the model and format

    Compare models on your real evaluation set, not only public leaderboards. Check the licence, commercial-use terms, supported languages, instruction-following behaviour, context length, and tool-calling support. For Indian deployments, test Hindi, Tamil, Telugu, Bengali, Marathi, and code-switched queries directly. A model that performs well in English may produce poor spelling, script handling, or factuality in an Indic language. The guide to low-resource Indic NLP provides useful background for this evaluation.

    For general local inference, GGUF is the most flexible format. It works well with llama.cpp and many desktop tools. Quantisation reduces weight precision so a model occupies less memory:

    • Q8: Higher quality and larger files; useful when memory allows.
    • Q6 or Q5: A sensible quality-to-size compromise.
    • Q4: Often the practical default for 7B–8B models.
    • Q3 and below: Smaller, but quality loss can become visible on reasoning and multilingual tasks.

    A 7B model stored in FP16 may need around 14GB just for weights. A 4-bit version may fit in roughly 4–6GB, but the exact footprint depends on the quantiser, architecture, context length, and runtime. Always leave headroom rather than filling all available memory.

    Deploy quickly with Ollama

    Ollama is a convenient starting point for developers who want a local model and an HTTP API with minimal setup. Install it for your operating system, then pull and run a supported model:

    ollama pull llama3.2:3b
    ollama run llama3.2:3b

    The service commonly listens on http://localhost:11434. A basic API request looks like this:

    curl http://localhost:11434/api/generate \\
      -d '{"model":"llama3.2:3b","prompt":"Summarise this text in Hindi","stream":false}'

    For production-like use, set model parameters explicitly, cap the context window, and stream responses where the interface benefits from progressive output. Do not expose the Ollama port directly to the public internet. Put authentication, network controls, request limits, and an application-level audit trail in front of it.

    Use llama.cpp for control and efficiency

    Choose llama.cpp when you need a small portable binary, CPU optimisation, Apple Metal support, CUDA offload, or precise control over threads and GPU layers. Download a compatible GGUF file, then start its server:

    ./llama-server -m ./models/model-q4_k_m.gguf \\
      -c 4096 -ngl 999 --host 127.0.0.1 --port 8080

    The exact flags vary by build and hardware. Tune thread count, batch size, GPU layers, and context length independently. Benchmark several prompts and watch memory usage. A faster setting that causes swapping is not a faster deployment.

    Desktop users may prefer LM Studio for model discovery, quantisation comparison, and a local OpenAI-compatible server. It is useful for prototyping, but automate repeatable deployments with pinned model files, configuration, and startup scripts.

    Add private knowledge with retrieval

    An SLM’s weights do not contain your latest policies, catalogues, or internal documents. Use retrieval-augmented generation (RAG) when answers must cite changing business information. A typical local pipeline is:

    1. Extract and clean documents.
    2. Split them into meaningful chunks with metadata.
    3. Generate embeddings locally.
    4. Store vectors in a local database such as SQLite-backed storage or Chroma.
    5. Retrieve a small number of relevant chunks.
    6. Ask the SLM to answer only from those chunks and show citations.

    Keep retrieval separate from generation so you can test failures. If the system retrieves the wrong document, changing the prompt will not solve the underlying problem. For agentic workflows, see this guide to deploying open-source AI agents, and for Llama-based orchestration, review how to deploy Llama 3 agents.

    Secure and evaluate the deployment

    Local does not mean automatically secure. Encrypt disks, restrict model and document directories, patch runtimes, and prevent sensitive prompts from appearing in unrestricted logs. Separate tenants if multiple organisations share one machine. Define retention rules for prompts, retrieved passages, and generated outputs.

    Create an evaluation set before launch. Track:

    • Exact-match or F1 scores for extraction and classification.
    • Groundedness and citation accuracy for RAG.
    • Response latency, tokens per second, and peak memory.
    • Failure rates across languages, scripts, and code-switching.
    • Prompt-injection, data-leakage, and unsafe-output behaviour.

    Use low temperature for deterministic extraction, but do not expect temperature alone to prevent hallucinations. Enforce JSON schemas, validate fields in application code, and return a clear “not found” result when evidence is missing.

    A practical rollout plan

    Begin with a 3B or 7B instruct model and a Q4 or Q5 GGUF build. Run it on a developer laptop, then test the same workload on the intended office server or edge device. Pin the model digest, runtime version, prompt template, and generation settings. Measure quality and cost against a cloud baseline before deciding whether local-only, hybrid, or cloud fallback is appropriate.

    For Indian builders, the strongest deployment is often hybrid: local inference for sensitive or offline tasks, and a stronger hosted model only for explicitly approved complex requests. This keeps latency and data exposure manageable without forcing one model to handle every job.

    Last updated 23 September 2026

AIGI may be inaccurate. Replies seeded from the guide above.