Open-source AI becomes valuable in production when it is treated as a software system, not simply a model loaded onto a GPU. Teams must make decisions about licensing, data handling, inference capacity, evaluation, reliability, and unit economics before serving their first real users.
For Indian startups and enterprises, deployment adds practical constraints: GPU availability, regional latency, DPDP Act obligations, multilingual workloads, and limited platform engineering capacity. This guide explains how to implement open source AI in production environments with a path from proof of concept to a maintainable service.
Start with the workload, not the model
Define the production job before comparing model leaderboards. Write down:
- The inputs and outputs the system must handle.
- Required accuracy, latency, throughput, and uptime.
- Whether responses must be grounded in company data.
- The languages, scripts, and domain terminology involved.
- Data residency, retention, and access-control requirements.
- A maximum cost per request or per completed task.
A customer-support assistant, document extractor, coding copilot, and voice agent need different architectures. For Indic-language or mixed-language applications, benchmark the model on representative Hindi, Tamil, Bengali, and code-switched samples rather than relying only on English benchmarks. Teams exploring language-specific systems can also review this builder’s guide to low-resource Indic NLP.
Choose a model and licence you can operate
Model selection is a trade-off among quality, context length, speed, hardware requirements, and licence terms. Compare at least one small model for low-cost routing and one stronger model for difficult requests. A smaller model may be the better choice when prompts are narrow, retrieval is strong, or response formats are tightly constrained.
Check the licence, acceptable-use policy, attribution requirements, and restrictions on redistribution or hosted services. Record the exact model revision, tokenizer, quantisation method, and downloaded artefacts in your internal registry. Do not assume that a model described as “open source” has the same permissions as an Apache-2.0 software library.
For inference, use a serving engine rather than an ad hoc PyTorch process:
- vLLM provides continuous batching, paged attention, streaming, and OpenAI-compatible endpoints.
- Hugging Face TGI is useful for teams already standardised on the Hugging Face ecosystem.
- TensorRT-LLM can deliver strong NVIDIA performance when the additional optimisation work is justified.
- llama.cpp or similar runtimes are practical for CPU, edge, and smaller quantised deployments.
Quantisation reduces memory and often lowers cost. Test 4-bit and 8-bit variants against a fixed evaluation set; a small quality loss may be acceptable for classification or extraction, but not for safety-critical or high-value reasoning. Reserve GPU memory for the KV cache, concurrent sequences, and longer context windows rather than sizing only for model weights.
Design the production architecture
A robust deployment normally separates the application layer, model gateway, inference workers, retrieval services, and observability stack. The application should not connect directly to a GPU worker. Put an authenticated gateway in front of the model to enforce quotas, route requests, redact sensitive fields, and provide consistent telemetry.
Use an OpenAI-compatible internal schema where practical. It lets the product team change models without rewriting every client and makes controlled fallback easier. Implement streaming with Server-Sent Events when users benefit from early tokens, but measure complete-response latency as well as time to first token.
For variable traffic, continuous batching and queue-based admission control are more important than simply adding GPUs. Define behaviour for overload: queue, degrade to a smaller model, return a structured retry response, or use a managed fallback. Kubernetes with the NVIDIA GPU Operator can work for multi-service clusters, while a simpler VM or container deployment is often safer for an early production workload. Consider Kubernetes only when the team can operate it reliably.
Build RAG as a data product
Most enterprise deployments need retrieval-augmented generation rather than model weights alone. A production RAG pipeline should version its source documents, parsing rules, chunking strategy, embedding model, index, and access permissions.
The minimum workflow is:
- Ingest documents from approved sources.
- Extract text, tables, metadata, and document-level permissions.
- Remove duplicates and stale versions.
- Create chunks that preserve headings and useful surrounding context.
- Generate embeddings and store them with filters such as tenant, department, language, and date.
- Retrieve candidates, rerank where necessary, and pass only relevant context to the model.
- Cite source documents or return “insufficient evidence” when retrieval is weak.
Milvus, Weaviate, pgvector, and other vector stores can all work; operational fit matters more than brand. Keep the original document identifier with every chunk so users and auditors can trace an answer. For legal, finance, healthcare, or government workloads, retrieval permissions must be enforced before generation, not merely described in the prompt. The same discipline is useful in AI legal document automation in India.
Secure data, models, and dependencies
Self-hosting improves control but does not remove security risk. Treat model files, adapters, prompts, tools, and container images as production dependencies.
- Scan images and Python packages for known vulnerabilities.
- Download models from trusted sources and verify checksums where available.
- Keep model storage private and restrict who can replace artefacts.
- Separate development, staging, and production credentials.
- Encrypt data in transit and at rest; define retention and deletion rules.
- Detect and mask unnecessary PII before logging or indexing.
- Apply tenant-level authorisation to retrieval and tool calls.
- Rate-limit public endpoints and validate file uploads.
- Test prompt injection, data exfiltration, unsafe tool use, and denial-of-service scenarios.
Under India’s DPDP framework, document the purpose and handling of personal data, minimise collection, and provide appropriate safeguards. A data-centre location alone is not a complete compliance strategy. Your contracts, processors, access controls, retention policy, and incident process matter too.
Evaluate before and after launch
Create a test set from real, permissioned examples before tuning the system. Include normal requests, ambiguous inputs, adversarial prompts, long documents, multilingual queries, and cases where the correct response is to refuse or ask for clarification.
Track separate metrics for retrieval and generation:
- Retrieval hit rate, ranking quality, and citation correctness.
- Task accuracy, structured-output validity, and refusal quality.
- Hallucination or unsupported-claim rate.
- Time to first token, end-to-end latency, and tokens per second.
- Error rate, queue time, GPU utilisation, and availability.
- Cost per request, per document, or per successfully completed task.
Use a small human-reviewed evaluation set as a control. LLM-as-a-judge can help with scale, but it should not be your only quality signal. Log prompt and response metadata safely, sample traces for review, and monitor quality drift after model, prompt, retrieval, or corpus changes. Teams building agentic systems should also follow the more focused guide to deploying open-source AI agents in production.
Control GPU cost and capacity
Estimate demand using tokens, concurrent requests, context length, and peak traffic—not just monthly active users. Benchmark the full serving stack on the target hardware. L4 and A10-class GPUs may be suitable for many inference workloads; larger models, long contexts, and high concurrency may require more memory or tensor parallelism. Availability and pricing vary by provider and region, so compare Mumbai, Hyderabad, Chennai, and other accessible locations where relevant.
Use continuous batching, prompt-prefix caching, quantisation, response limits, and smaller routing models. Reserve dedicated capacity for latency-sensitive traffic and use interruptible instances for replay jobs, embedding batches, and offline evaluation. Distillation can reduce recurring cost, but only after proving that the smaller model preserves the required task quality.
Launch with operational controls
Begin with a narrow, observable release. Use shadow traffic or an internal pilot before exposing the system broadly. Define rollback triggers for latency, error rate, unsupported answers, and cost. Keep the previous model and index available until the new version has passed evaluation and a real traffic comparison.
Maintain a model card and deployment record containing the model revision, licence, data sources, known limitations, hardware, quantisation, prompt version, evaluation results, and owner. This turns a fragile prototype into an auditable service. For further implementation patterns, see building high-performance AI applications with open-source tools.
A practical rollout sequence
1. Week 1: Define the task, risk level, SLOs, budget, and evaluation set.
2. Weeks 2–3: Benchmark two or three models and serving configurations on representative hardware.
3. Weeks 3–5: Build the gateway, authentication, retrieval pipeline, structured logging, and offline tests.
4. Weeks 5–7: Run staging load tests, security tests, failure drills, and a controlled user pilot.
5. After launch: Review quality, cost, and incidents weekly; promote model and data changes through versioned releases.
Open-source AI is production-ready when the surrounding system is production-ready. Choose the smallest model that meets the task, make evidence and permissions explicit, measure the real user outcome, and scale infrastructure only after the workload justifies it.