What self-hosting an open-source LLM means
A self-hosted open-source LLM runs on infrastructure you control rather than sending every prompt to a managed API. That infrastructure may be a developer workstation, an office server, an Indian cloud region, or a private Kubernetes cluster. The model’s weights, prompts, documents, logs, and inference service remain under your operational control—subject to the model’s licence and the policies of your hosting provider.
Self-hosting is not automatically cheaper or easier. It becomes compelling when you need predictable latency, offline operation, data residency, custom networking, high request volume, or control over model behaviour. For a small prototype, a hosted API may still be the sensible choice. Compare the options against your workload, compliance requirements, engineering capacity, and peak traffic rather than choosing on model size alone.
If you are new to the ecosystem, start with best open-source AI projects for beginners before committing to a production deployment.
Choose the model before buying hardware
Start by writing down the task and its acceptance criteria: answer questions over internal documents, extract fields from invoices, generate code, translate Indian languages, or power an agent. A smaller instruct model with good retrieval and evaluation can outperform a larger general model on a narrow business workflow.
Evaluate these characteristics:
- Parameter size: 3B–8B models are practical for local development; 7B–14B models often offer a useful quality and cost balance; larger models need substantial GPU memory or multi-GPU serving.
- Context window: A long context does not replace retrieval. Test the model with the amount of text your application will actually provide.
- Language coverage: Check performance on Hindi, Tamil, Bengali, Marathi, Telugu, and code-switched queries rather than relying on English benchmarks. Work on low-resource Indic natural language processing can inform dataset and evaluation choices.
- Licence: “Open source” is used loosely in AI. Read the exact licence, acceptable-use terms, attribution requirements, commercial restrictions, and redistribution rules for both weights and code.
- Quantised variants: GGUF, GPTQ, AWQ, and other formats trade some quality for lower memory use. Confirm that your serving engine supports the format.
Prefer a current instruct-tuned model from a maintained project. GPT-Neo and BERT are useful historically or for specific encoder tasks, but they are not default choices for a modern chat application. Review model cards, known limitations, safety notes, training-data statements, and independent evaluations before deployment.
Estimate hardware and operating cost
Memory planning is more useful than a generic “you need a GPU” recommendation. Inference memory includes model weights, the key-value cache for the context, runtime overhead, and any batching. As a rough starting point, weight storage is approximately:
- 4-bit quantisation: about 0.5 bytes per parameter, plus overhead
- 8-bit quantisation: about 1 byte per parameter, plus overhead
- 16-bit weights: about 2 bytes per parameter, plus overhead
An 8B model at 4-bit precision may fit on a modern 8–12 GB GPU for short contexts, while longer contexts and concurrent users require more headroom. CPU inference is possible, especially with quantised models, but latency may be unsuitable for interactive applications. Measure tokens per second, time to first token, maximum concurrency, and tail latency on your actual prompts.
For Indian teams, compare the full monthly cost: GPU rental, persistent disks, bandwidth, electricity, observability, backups, and engineering time. A local workstation can be economical for development; production workloads may suit a GPU cloud or dedicated server. Keep sensitive workloads in a region and network design that match your contractual and regulatory obligations.
Deploy a reproducible inference stack
Separate the model from the application. A common architecture is:
1. Model storage: Download approved weights into versioned, access-controlled storage.
2. Inference server: Use a runtime such as llama.cpp, Ollama, vLLM, or Hugging Face TGI according to hardware, model format, batching needs, and API compatibility.
3. Application layer: Expose an internal service that handles authentication, prompt templates, retrieval, rate limits, and structured outputs.
4. Data layer: Keep document indexes, user data, and conversation history separate from model files.
5. Observability: Record latency, token usage, errors, queue depth, and evaluation outcomes without storing sensitive prompts unnecessarily.
Containerise the service and pin versions. A minimal workflow might look like this after installing Docker and selecting a supported runtime:
docker run --gpus all --name llm-server \\
-p 8000:8000 \\
-v /srv/models:/models \\
your-approved-runtime:version \\
--model /models/your-modelTreat this as a deployment pattern, not a universal command. Check the runtime’s documentation for GPU drivers, CUDA compatibility, model formats, tensor parallelism, and health-check endpoints. Keep model downloads verifiable and maintain a rollback copy of the last known-good image and weights.
Build quality with retrieval before fine-tuning
For company knowledge, begin with retrieval-augmented generation (RAG): clean and chunk documents, create embeddings, retrieve relevant passages, and require the model to answer from cited evidence. Test chunk size, metadata filters, multilingual embeddings, and retrieval recall. This is usually faster and easier to update than fine-tuning.
Fine-tune only when you need consistent style, classification behaviour, structured output, or domain-specific task performance that prompting and retrieval cannot deliver. Use a representative, permissioned dataset; remove secrets and personal data; split training and evaluation examples; and track every dataset and hyperparameter change. Parameter-efficient methods such as LoRA can reduce training cost, but they do not remove the need for careful evaluation.
For language products serving Indian users, include transliteration, spelling variation, mixed-language prompts, numerals, names, and regional terminology in your test set. Projects focused on open-source vision-language models for Indian languages are also useful references when your application combines text with images or documents.
Secure the system before exposing it
Never place an unauthenticated inference endpoint directly on the public internet. At minimum:
- Put the service behind a private network, reverse proxy, VPN, or zero-trust gateway.
- Use strong authentication, per-user authorisation, quotas, and request-size limits.
- Restrict outbound network access from the model container to reduce data-exfiltration risk.
- Scan uploaded files and defend against prompt injection, malicious tool instructions, and data leakage.
- Encrypt traffic and disks; manage secrets outside images and source control.
- Log administrative actions and redact prompts, documents, and personal data where possible.
- Patch the operating system, drivers, runtime, dependencies, and model-serving image.
- Back up configuration, indexes, evaluations, and essential data—not necessarily every disposable cache.
If the LLM can call tools, enforce permissions in application code. The model should propose an action; a policy layer should decide whether that action is allowed.
Evaluate, monitor, and scale
Create a private evaluation set before launch. Score factuality, citation correctness, refusal behaviour, language quality, structured-output validity, latency, and cost. Include adversarial prompts and failure cases, and review a sample of real interactions with appropriate privacy controls.
Start with one GPU or one local node, then scale based on measurements. Continuous batching improves utilisation for concurrent workloads; quantisation reduces memory pressure; caching can reduce repeated computation; and smaller specialist models may handle routing, classification, or extraction more efficiently. Production systems may benefit from a queue, autoscaling, multi-GPU serving, and a fallback model, but each adds operational complexity.
For broader architecture patterns, see this guide to building high-performance AI applications with open-source tools. If the model will operate as part of a workflow, apply the production controls described in how to deploy open-source AI agents.
A practical launch checklist
Before serving real users, confirm that you have:
- A documented model, licence, version, quantisation, and limitations
- A measured hardware profile for normal and peak traffic
- Versioned prompts, retrieval settings, datasets, and evaluation results
- Authentication, network controls, secret management, and backup procedures
- Clear retention and deletion rules for prompts and generated outputs
- A human escalation path for high-impact or uncertain responses
- Dashboards for latency, errors, throughput, utilisation, and quality regressions
- A rollback plan for model, runtime, prompt, and index changes
Self-hosting succeeds when it is treated as an engineering system, not simply a model download. Choose the smallest model that meets your quality bar, keep data flows explicit, and make every production decision measurable.