0tokens

Apply for AI Grants India

Financial support for innovators building the future of AI in India.

Apply now

Chat · open source indic llm hosting guide

Open-Source Indic LLM Hosting Guide: Build and Deploy

  1. aigi

    Open-source Indic LLM hosting is no longer just a model-download exercise. A useful deployment must handle script diversity, uneven tokenisation, bursty traffic, privacy requirements, and the economics of serving users in India. The right architecture depends as much on your target languages and latency budget as on parameter count.

    This guide covers the decisions that matter in 2026: model evaluation, GPU sizing, quantisation, serving, observability, and production safeguards. For background on the data and evaluation challenges behind Indian-language systems, see this low-resource Indic NLP builder’s guide.

    Start with the workload, not the model

    Define the product before choosing weights. A customer-support chatbot, document translator, voice assistant, and batch classification service have different infrastructure needs.

    Record these requirements:

    • Languages and scripts: List the languages users will actually send, including code-mixed Hindi-English or Tamil-English input.
    • Tasks: Separate chat, extraction, translation, summarisation, classification, and structured JSON generation.
    • Context length: Long legal or government documents can increase KV-cache memory far more than the model weights.
    • Traffic: Estimate requests per second, concurrent conversations, peak bursts, and average output tokens.
    • Latency: Set targets for time to first token and tokens per second, rather than relying on a single average response time.
    • Data controls: Decide whether prompts may leave India, whether logs can contain personal data, and how long traces are retained.

    A smaller model with a strong Indic tokenizer can outperform a larger English-centric model on cost, speed, and answer quality. Benchmark on representative user inputs before committing to a checkpoint.

    Select and verify an Indic model

    Model names and checkpoints change quickly, so treat every repository as an unverified dependency. Inspect the model card, licence, tokenizer files, supported languages, base model, context limit, quantised variants, and known limitations. Confirm that commercial use and redistribution match your product plans.

    Candidate families may include multilingual checkpoints, Hindi-focused models such as Airavata or OpenHathi, Tamil-specialised models, and newer Indic instruction-tuned releases. Do not assume that a model claiming support for 22 languages performs equally well across all of them.

    Build a small evaluation set containing:

    • Native-script prompts for each priority language
    • Code-mixed and transliterated input
    • Names, places, dates, currency, and Indian phone-number formats
    • Long passages with punctuation and numerals
    • Safety-sensitive and adversarial requests
    • JSON or tool-calling tasks if your application needs structured output

    Score factuality, instruction following, script fidelity, refusal behaviour, formatting, and latency. Include native speakers in review; perplexity alone will not reveal broken grammar or culturally inappropriate answers.

    Size GPUs using memory, concurrency, and context

    A rough FP16 estimate is two bytes per parameter, before runtime overhead. An 8B model therefore needs about 16 GB for weights, but serving also requires space for the KV cache, CUDA graphs, activations, and framework overhead. Quantisation reduces weight memory but does not eliminate the cost of long contexts or concurrent requests.

    Practical starting points:

    • Local development: An RTX 3090 or 4090 with 24 GB can run many 7B–8B models in 4-bit form.
    • Small production workloads: L4, A10, or comparable GPUs can suit moderate traffic when batching is controlled.
    • High concurrency and long context: A100 80 GB or H100-class hardware provides more memory bandwidth and headroom.
    • CPU or edge deployment: GGUF with llama.cpp is viable for offline, low-volume, or privacy-sensitive use, but interactive speed may be limited.

    Keep at least 20–30% memory headroom for traffic spikes and longer prompts. Measure GPU utilisation, KV-cache occupancy, queue time, and out-of-memory events under realistic concurrency. A GPU that is cheap per hour can be expensive if it forces aggressive limits on context or throughput.

    Choose a serving stack

    For an OpenAI-compatible API and continuous batching, vLLM is usually the strongest first choice for Llama-, Gemma-, and Mistral-derived Indic models. It offers PagedAttention, streaming responses, tensor parallelism, and production-friendly metrics. Validate architecture and quantisation support against the exact checkpoint; compatibility is not guaranteed merely because the base model is supported.

    TGI is a reasonable alternative when your team already operates Hugging Face infrastructure. llama.cpp is well suited to CPU, Apple Silicon, and edge deployments. Ollama is convenient for local prototyping, but production teams should still manage authentication, quotas, observability, and model lifecycle explicitly.

    If the LLM is one component of a larger workflow, separate inference from orchestration. A multilingual voice-agent architecture and deployment plan may require dedicated speech recognition, translation, retrieval, and text-to-speech services rather than one oversized LLM endpoint.

    Deploy vLLM with a reproducible baseline

    Pin the vLLM version, CUDA image, model revision, and tokenizer revision. Avoid using latest in production. A basic NVIDIA-container deployment looks like this:

    docker run --rm --gpus all \\
      -p 8000:8000 \\
      -v $HOME/.cache/huggingface:/root/.cache/huggingface \\
      vllm/vllm-openai:<pinned-version> \\
      --model <org>/<model> \\
      --dtype bfloat16 \\
      --max-model-len 4096 \\
      --gpu-memory-utilization 0.90

    Use a quantisation flag only when the checkpoint was prepared for that method. AWQ, GPTQ, bitsandbytes, and GGUF are not interchangeable. Start with a conservative context limit, then raise it after measuring KV-cache pressure. Expose the service privately behind an API gateway, require authentication, enforce request and token quotas, and set timeouts for both queued and generated requests.

    For multi-GPU serving, test tensor parallelism with your model and interconnect. A distributed setup can increase capacity, but communication overhead may reduce performance for small requests. Keep a separate staging endpoint for model upgrades and tokenizer changes.

    Quantise without damaging Indic quality

    Quantisation is a quality decision, not only a cost decision. AWQ and GPTQ can provide strong GPU performance, while GGUF is practical for llama.cpp and mixed CPU/GPU execution. Compare FP16 or BF16 against the compressed version on every priority language.

    Track:

    • Exact-match and structured-output success
    • Translation adequacy and terminology preservation
    • Native-speaker ratings for grammar and fluency
    • Repetition, truncation, and script corruption
    • Time to first token, generation speed, and cost per million tokens

    Do not infer quality from an English benchmark. A four-bit model can appear healthy on English prompts while losing distinctions in Bengali, Kannada, or Marathi. Keep the original model available for regression testing and document the acceptable quality trade-off.

    Fix tokenisation and retrieval bottlenecks

    English-optimised tokenizers may represent Indic text inefficiently, increasing input cost and reducing effective context. Measure tokens per character and tokens per sentence for each target language. Compare native script, transliteration, and code-mixed text; production traffic rarely matches clean benchmark prompts.

    Retrieval systems need the same care. Chunk by language and document structure, preserve Unicode normalisation, and use multilingual embeddings validated on Indian-language search. Store the original text alongside normalised text so citations and user-visible output remain faithful. If you are building a broader open-source stack, review Indian open-source AI developer projects for adjacent implementation ideas.

    Operate securely in India

    Run sensitive workloads in an Indian region or private VPC when contractual and regulatory requirements demand it. The DPDP Act does not automatically require every workload to remain in India, but data minimisation, purpose limitation, access control, retention, and processor governance still matter.

    At minimum:

    • Redact personal data from application and inference logs.
    • Encrypt traffic in transit and disks at rest.
    • Keep model files, adapters, prompts, and evaluation data under access control.
    • Record model and tokenizer versions for every response.
    • Provide deletion and retention controls for user conversations.
    • Scan downloaded weights and containers; verify licences before commercial deployment.

    Do not expose a raw vLLM port to the public internet. Put rate limiting, authentication, request validation, and abuse monitoring in front of it.

    Monitor what users experience

    Operational dashboards should distinguish queue delay, prefill time, decode speed, error rate, GPU memory, and tokens generated. Break metrics down by language and route. A healthy aggregate latency can hide severe degradation for long Telugu prompts or low-volume languages.

    Log correlation IDs and safe metadata, not raw prompts by default. Alert on rising queue time, repeated out-of-memory errors, abnormal token counts, and quality-regression test failures. Use canary traffic for new weights and roll back quickly when a tokenizer or quantisation change causes failures.

    For systems that use planning, retrieval, or multiple model calls, treat the inference service as a dependency with budgets and fallbacks. Patterns from distributed AI-agent systems are useful for isolating retries, queues, and failure domains.

    A practical launch checklist

    Before production, confirm that you can:

    • Reproduce the environment from a pinned container and infrastructure definition
    • Restore model files without relying on an untracked local cache
    • Serve every priority language with acceptable quality and latency
    • Survive peak concurrency without exhausting KV-cache memory
    • Authenticate, rate-limit, and audit API access
    • Roll back model, tokenizer, and quantisation changes independently
    • Remove sensitive logs and honour retention policies
    • Estimate cost per request using real input and output token distributions

    The best open-source Indic deployment is not necessarily the largest model. It is the smallest, well-evaluated system that serves your languages reliably, protects user data, and leaves enough operational headroom for growth. Start with a measurable workload, benchmark native-language traffic, and expand hardware only when evidence—not model size—requires it.

    Last updated 23 September 2026

AIGI may be inaccurate. Replies seeded from the guide above.