Open-source models give Indian startups more control over cost, latency, data handling, and language behaviour than a purely API-led architecture. But self-hosting is not automatically cheaper or safer. The business case depends on traffic, response-time targets, model size, operational maturity, and the sensitivity of user data.
This guide covers the decisions that matter when moving from a prototype to production: model selection, quantisation, inference, Indian cloud infrastructure, Indic-language quality, compliance, observability, and unit economics.
Start with the workload, not the model
Define the product requirement before comparing checkpoints. A customer-support assistant, document extractor, voice agent, and coding copilot have different latency and hardware profiles.
Write down:
- Inputs and outputs: text only, PDFs, images, audio, or tool calls.
- Traffic: average requests per minute, peak concurrency, and expected growth.
- Latency: time to first token and complete-response targets.
- Context size: typical and maximum prompt length.
- Quality threshold: factual accuracy, structured output, language fluency, and refusal behaviour.
- Data classification: public, internal, personal, financial, health, or regulated information.
For low traffic or uncertain product-market fit, a managed API may remain the sensible choice. Self-hosting becomes more attractive when request volume is predictable, prompts contain sensitive information, you need custom weights or adapters, or latency must be controlled within India.
Select a model and licence carefully
In 2026, the practical shortlist includes small and medium instruction-tuned models from the Llama, Mistral, Gemma, Qwen, and other open-weight families. “Open source” is often used loosely: some releases provide weights but impose usage restrictions, while others publish code, training details, and permissive licences. Review the exact licence before building a commercial product.
Choose by measured performance on your own test set rather than parameter count alone. A smaller 7B–9B model may outperform a larger model on a narrow workflow when paired with retrieval, good prompts, and constrained decoding. Consider:
- 8B-class models for chat, classification, extraction, and lightweight agents.
- 14B–32B models when reasoning quality or long-form generation justifies higher serving cost.
- Mixture-of-experts models when supported infrastructure can handle their memory and routing characteristics.
- Indic-focused or multilingual models when Hindi, Tamil, Telugu, Bengali, Marathi, code-switching, or transliterated input is central to the product.
For teams learning the ecosystem, Indian open-source AI developer projects can provide useful examples of local datasets, tooling, and deployment patterns.
Size the hardware with quantisation
FP16 inference is memory-intensive. As a rough planning rule, model weights alone require about two bytes per parameter in FP16, before adding the KV cache, runtime overhead, and batching capacity. Quantisation can reduce memory use substantially, but it may affect quality and latency.
Common options include:
- INT8 or 8-bit quantisation: a conservative starting point for quality-sensitive workloads.
- 4-bit quantisation: useful for fitting 7B–14B models on more affordable GPUs.
- GGUF with llama.cpp: practical for CPU, Apple Silicon, and hybrid edge deployments.
- AWQ, GPTQ, or related GPU formats: useful when your serving stack supports them efficiently.
A 24GB GPU can be suitable for development and some 7B–9B production workloads, but concurrency, context length, and output tokens can change the requirement quickly. Benchmark with realistic prompts. Do not size infrastructure from VRAM alone.
Build a production inference layer
For GPU serving, vLLM is a strong default because it supports continuous batching, efficient attention, OpenAI-compatible endpoints, and common quantised models. Hugging Face TGI remains useful where its ecosystem and operational features fit the team. llama.cpp is often better for local, CPU, or edge scenarios, while Ollama is convenient for experimentation rather than a complete high-scale serving strategy.
Your inference layer should include:
- Request authentication, quotas, and tenant isolation.
- Streaming responses where the user experience benefits from them.
- Maximum input and output token limits.
- Timeouts, retries, cancellation, and back-pressure.
- Model versioning and a rollback path.
- Separate development, evaluation, and production endpoints.
Keep the application layer model-agnostic where possible. An internal gateway should allow you to switch between a self-hosted model and a managed fallback without rewriting every product integration.
Choose Indian infrastructure deliberately
AWS, Azure, and Google Cloud have Indian regions, while specialised providers such as E2E Networks and domestic data-centre operators may offer different GPU availability and pricing. Compare the complete bill, not only the hourly GPU rate:
- GPU rental and attached CPU/RAM.
- Persistent storage for weights and logs.
- Egress, load balancing, and backup charges.
- Managed Kubernetes or platform fees.
- Support, uptime commitments, and replacement time for failed GPUs.
Use on-demand capacity for predictable production traffic and spot or interruptible capacity for evaluation, batch inference, and fine-tuning. Keep model artefacts in encrypted storage and restrict access through private networking, identity policies, and short-lived credentials.
Improve Indic-language performance systematically
A model’s ability to generate Hindi or another Indian language is not enough. Test spelling, script handling, transliteration, code-switching, names, currency formats, dates, honorifics, and regional terminology. Include noisy user input and speech-transcribed text if your product is voice-first.
Start with retrieval and prompt improvements before fine-tuning. When adaptation is necessary, use LoRA or QLoRA on carefully reviewed examples. A smaller, representative dataset is more valuable than a large collection of duplicated or synthetic examples. Track performance separately by language, script, domain, and user segment.
Teams building multilingual systems should also review this guide to low-resource Indic NLP. For voice products, language quality depends on the speech stack as well as the LLM; the voice-agent architecture guide covers the broader deployment path.
Treat DPDP compliance as an engineering requirement
Self-hosting can reduce data exposure, but it does not by itself make a product compliant. Map the data flow from client to API gateway, retrieval store, inference server, logs, analytics tools, and support systems.
Implement:
- Data minimisation and purpose limitation.
- Clear retention periods for prompts, outputs, and traces.
- Redaction or tokenisation of personal data before logging.
- Encryption in transit and at rest.
- Role-based access, audit trails, and incident procedures.
- Vendor and subprocesser reviews when using external GPUs or observability tools.
- Human review paths for high-impact decisions.
For financial, health, education, or employment use cases, add sector-specific controls and avoid presenting model output as verified advice without appropriate review.
Evaluate quality, safety, and cost before launch
Create a fixed evaluation set from real, consented, and suitably anonymised product examples. Measure factuality, groundedness, structured-output validity, refusal quality, toxicity, prompt-injection resistance, latency, throughput, and cost per successful task.
Monitor production signals such as token usage, queue time, GPU utilisation, error rates, fallback frequency, and user corrections. Log enough metadata to debug failures without retaining unnecessary personal content. Re-run the evaluation set after every model, prompt, retrieval, or quantisation change.
The relevant comparison is not “open source versus API” in the abstract. Calculate cost per completed task and include engineering time, support, idle capacity, monitoring, and incident response. A hybrid architecture often wins: self-host a stable, high-volume workflow and route complex or infrequent requests to a managed model.
A practical rollout plan
1. Prototype: benchmark two or three models on a representative test set.
2. Pilot: deploy one quantised model behind an authenticated gateway.
3. Harden: add rate limits, private networking, redaction, monitoring, and rollback.
4. Validate: run load tests and language-specific safety evaluations.
5. Scale: add replicas, autoscaling, caching, or a model router only after measuring the bottleneck.
Start narrow, publish internal quality and cost targets, and keep a fallback path. For most Indian startups, disciplined evaluation and data controls matter more than chasing the largest available model.