India’s LLM application stack should optimise for more than model quality. Latency, token cost, data governance, Indic-language performance, GPU availability, and operational simplicity all affect whether a prototype becomes a dependable product.
The right approach in 2026 is usually model-agnostic: start with hosted models to validate demand, keep your application portable, and move selected workloads to open models or local inference when volume, privacy, or latency justifies the effort. Avoid assembling a complicated agent platform before you understand the workload. A well-designed retrieval pipeline with clear evaluation often beats an autonomous agent with many tools.
Start with a simple reference architecture
For most Indian startups, a production LLM application can be divided into eight layers:
- Client: React or Next.js for web, with streaming responses and clear loading states.
- Application API: Python with FastAPI, or TypeScript with Node.js, for authentication, business logic, and rate limits.
- Model gateway: LiteLLM, Portkey, or an equivalent abstraction for provider switching, retries, routing, and spend tracking.
- Model layer: One primary model, one lower-cost fallback, and an open-weight option for sensitive or high-volume tasks.
- Knowledge layer: Document extraction, chunking, embeddings, retrieval, reranking, and citations.
- Data layer: PostgreSQL for transactional data, Redis for caching and queues, and pgvector or a dedicated vector database for retrieval.
- Operations: Tracing, evaluation, prompt/version management, alerts, and audit logs.
- Infrastructure: Indian cloud regions for application services, with GPU infrastructure added only when the economics support it.
This separation makes it easier to replace a model or retrieval provider without rewriting the product.
Choose models by workload, not reputation
Use frontier APIs when you need strong reasoning, structured output, tool use, or rapid iteration. OpenAI, Anthropic, and Google remain practical choices, but compare them on your own test set rather than relying on benchmark scores. Measure answer quality, time to first token, total latency, context-window behaviour, rate limits, and the price of the complete request—including input, output, retries, and tool calls.
Use open-weight models when data control, predictable unit economics, offline operation, or custom fine-tuning matters. Llama, Mistral, Qwen, and specialised Indic models can be served privately, but hosting is not automatically cheaper. Include GPU rental, engineering time, monitoring, idle capacity, upgrades, and fallback infrastructure in the calculation.
A practical routing policy is:
- Small or deterministic models for classification, extraction, moderation, and intent detection.
- A mid-tier model for everyday customer support and document questions.
- A stronger model for difficult cases, escalation, and high-value workflows.
- An open model for workloads requiring controlled data residency or very high request volume.
Keep prompts, model settings, and routing rules in configuration rather than hard-coding them into business logic.
RAG: prioritise retrieval quality
Most enterprise LLM applications in India need retrieval-augmented generation (RAG), not fine-tuning. RAG lets the model answer from policies, product documents, contracts, support records, or internal knowledge while preserving a refreshable source of truth.
A robust pipeline should include:
1. File type detection and malware checks.
2. OCR for scanned PDFs and image-heavy documents.
3. Layout-aware parsing for tables, headings, and footnotes.
4. Language detection and normalisation across English and Indic scripts.
5. Chunking based on document structure, not an arbitrary character count.
6. Hybrid search combining keyword and vector retrieval.
7. Reranking before the context is sent to the model.
8. Source citations and an explicit “I don’t know” path.
LlamaIndex and Haystack are useful for data-heavy pipelines; LangChain is useful when the application has many tools or multi-step workflows. Choose one framework and keep core retrieval logic understandable. For messy Indian PDFs, document processing deserves as much attention as prompt design.
Select a database that matches your maturity
pgvector is the sensible starting point when PostgreSQL already stores users, permissions, documents, and application data. It reduces operational overhead and simplifies backups and access control.
Choose Qdrant or Weaviate when you need dedicated vector search features, independent scaling, or more advanced filtering. Managed services can accelerate launch but may increase recurring costs and create a data-residency review. Pinecone is convenient for teams that want a fully managed vector layer; compare its total cost with a self-hosted option before committing.
Do not store embeddings without metadata. Every chunk should carry document version, tenant, access policy, language, page reference, and creation date. Retrieval must enforce authorisation before context reaches the model.
Deploy for Indian users and real costs
Host APIs, queues, databases, and observability services close to your users where possible. AWS Mumbai, Azure India regions, and Google Cloud regions in India can reduce application latency, although the model endpoint may still be outside the country. Streaming hides some perceived delay, but it cannot fix slow retrieval, oversized prompts, or repeated tool calls.
For self-hosted inference, vLLM is a strong default for high-throughput serving; Text Generation Inference and other compatible servers may suit different model or deployment requirements. Indian GPU providers such as E2E Networks, Neysa, and Tata Communications can be evaluated alongside hyperscalers. Compare GPU type, availability, storage, egress, support, utilisation, and contractual uptime—not just hourly rental price.
Before adding GPUs, implement:
- Prompt and response caching where safe.
- Semantic caching for repeatable questions.
- Context trimming and token budgets.
- Batch processing for offline jobs.
- Model routing by task complexity.
- Timeouts, retries, circuit breakers, and provider fallbacks.
Teams expecting significant traffic should also review guidance on scaling backend infrastructure for AI applications.
Build for multilingual and voice use cases
India’s language requirements are product requirements, not a late-stage translation task. Test retrieval and generation separately across the languages, scripts, accents, and code-mixed inputs your users actually submit. Track transliteration, spelling variation, named entities, numerals, and regional terminology.
For voice products, treat speech recognition, turn-taking, text-to-speech, and the LLM as separate components. Measure end-to-end time to first audio, interruption handling, transcription accuracy, and fallback behaviour. The engineering considerations are different from a text chatbot; the guide to building a voice agent covers the relevant architecture, while real-time voice agents with fast barge-in focuses on conversational latency.
Security, privacy, and compliance
Classify data before sending it to any model provider. Redact unnecessary personal data, separate tenant data, encrypt secrets, and retain prompts and outputs according to a documented policy. Use role-based retrieval filters and log who accessed which source.
For regulated products—especially fintech, health, insurance, and government-facing systems—define human review thresholds and escalation paths. Do not present generated text as an authoritative decision where a human or deterministic rule is required. Review vendor terms, data retention, subprocessors, cross-border transfers, and obligations under India’s DPDP framework with qualified legal counsel.
Evaluate before you scale
A useful evaluation set should contain real, anonymised examples: easy questions, ambiguous requests, missing information, adversarial prompts, multilingual inputs, and permission-boundary tests. Track retrieval recall, groundedness, citation accuracy, refusal quality, latency, cost per successful task, and user correction rate.
Use LangSmith, Arize Phoenix, or OpenTelemetry-compatible tracing for request-level visibility. Ragas and DeepEval can automate parts of RAG evaluation, but human review remains necessary for high-impact workflows. Every prompt or model change should run against a versioned regression set before release.
Recommended starter stack
| Layer | Practical default | When to change it |
|---|---|---|
| API | FastAPI or Node.js | Match the team’s strongest ecosystem |
| Models | Hosted frontier model plus low-cost fallback | Add open inference for privacy or volume |
| Gateway | LiteLLM or Portkey | Use provider-native APIs for a very small MVP |
| RAG | LlamaIndex or Haystack | Use LangChain for tool-heavy workflows |
| Database | PostgreSQL + pgvector | Move to Qdrant/Weaviate at search scale |
| Cache/queue | Redis | Add Kafka or managed queues for high throughput |
| Inference | vLLM | Test alternatives for specialised deployments |
| Observability | OpenTelemetry + tracing platform | Add custom dashboards for unit economics |
| Frontend | Next.js with streaming | Use Streamlit or Chainlit for internal prototypes |
Start with the smallest architecture that can measure quality and cost. Then improve the bottleneck you can demonstrate—retrieval, latency, reliability, privacy, or unit economics—rather than adding tools by default. For agent-heavy products, study patterns for building distributed systems with AI agents before introducing multiple autonomous services.
For Indian founders building and validating these systems, AI Grants India offers grants, mentorship, and ecosystem resources to help move from prototype to deployment.