Retrieval-augmented generation (RAG) systems connect a language model to a controlled knowledge base. Instead of asking a model to rely only on its training data, a RAG application retrieves relevant documents at query time and uses them as context for its response. This makes answers more current, traceable, and useful for domain-specific work.
For Indian builders, RAG is often the fastest route from a general-purpose model to a useful product: a policy assistant for a government department, a support copilot for a fintech, a multilingual education tool, or an internal search system for a distributed team. But good RAG is not simply “upload documents and add a chatbot”. It is a data, search, evaluation, and product-design problem.
What a RAG system does
A typical RAG request follows this path:
1. A user submits a question.
2. The system converts the question into a search representation, usually an embedding.
3. A retriever finds relevant passages in a document store or vector database.
4. Optional reranking improves the order and quality of those passages.
5. The application places the retrieved context into a prompt.
6. The language model generates an answer, ideally with citations or source references.
7. Logs and user feedback support evaluation and improvement.
The model is responsible for synthesis and language. The retrieval layer is responsible for finding evidence. Keeping those responsibilities clear makes the system easier to debug.
RAG is especially valuable when information changes frequently, when answers must cite internal material, or when fine-tuning would be expensive and difficult to maintain. It is less suitable when the task requires precise calculations, transactional writes, or knowledge that cannot be safely exposed to the retrieval layer.
Core architecture
1. Source and ingestion layer
Start with authoritative sources: approved PDFs, product documentation, tickets, databases, websites, or structured records. Build an ingestion pipeline that extracts text, preserves headings and tables where possible, removes duplicate content, and attaches metadata such as:
- Document title, owner, version, and publication date
- Department, geography, language, and access permissions
- Section heading, page number, or source URL
- Validity period and review status
Scanned Indian-language documents may require OCR and language-specific quality checks. Do not treat OCR output as trustworthy by default; sample it and measure extraction errors before indexing it.
2. Chunking and indexing
Long documents must be divided into retrievable passages. Fixed-size chunks are a useful baseline, but semantic or heading-aware chunking generally preserves meaning better. Include limited overlap so that a definition split across boundaries remains searchable.
Store both the passage and its metadata. A vector index supports semantic similarity, while keyword search remains valuable for exact terms such as scheme names, case IDs, GST numbers, or legal clauses. In production, hybrid retrieval—combining vector and lexical search—often outperforms either method alone.
3. Retrieval and reranking
Retrieve more candidates than you intend to show the model, then rerank them using a cross-encoder or another relevance model. Apply filters before retrieval where access control, language, product, or date matters.
A multilingual product should test retrieval separately for English, Hindi, and every regional language it supports. Translating every query into English can help in some systems, but it may also lose legal, cultural, or domain-specific meaning. Compare native multilingual embeddings with translation-based approaches using real queries.
4. Generation and citations
The generation prompt should instruct the model to answer only from supplied evidence, distinguish facts from inference, acknowledge missing information, and cite the relevant sources. A confident answer without supporting evidence is a failure, not a success.
For high-stakes use cases, return structured output: answer, citations, confidence or evidence status, and escalation recommendation. Never allow retrieved text to override system instructions. Documents can contain prompt injection, malicious links, or misleading operational directions.
Teams building broader AI products may also benefit from the design patterns covered in building high-performance AI applications with open-source tools.
A practical build process
Define the job before choosing a model
Write down the user, decision, source corpus, acceptable latency, and cost per request. “Answer questions from documents” is too vague. “Help a customer-support agent find the correct refund policy and cite the clause” is testable.
Create a representative evaluation set
Collect real questions, including ambiguous queries, misspellings, short queries, multilingual prompts, adversarial instructions, and questions whose answers are absent from the corpus. Have domain experts label the relevant sources and the expected answer or refusal behavior.
Establish retrieval baselines
Begin with keyword search and a simple embedding retriever. Measure recall@k: how often the required evidence appears in the top k results. Then test chunk sizes, overlap, metadata filters, embedding models, hybrid search, and reranking one change at a time.
Evaluate generation separately
Useful measures include citation correctness, answer faithfulness, completeness, refusal quality, latency, and cost. Automated judges can accelerate iteration, but sample-based human review remains essential—particularly for health, finance, education, public services, and legal workflows.
Add operational safeguards
Implement document versioning, access-control filtering, audit logs, rate limits, PII redaction, prompt-injection detection, and a process for deleting outdated content. Keep tenant data isolated in multi-customer products. A retrieval system must enforce permissions before context reaches the model, not after the answer is generated.
Release in stages
Start with an internal pilot and a narrow corpus. Compare the RAG tool with the existing workflow, track unanswered questions, and give users a way to report incorrect sources. Expand only when retrieval quality and operational controls are stable.
Common failure modes
- Poor source quality: No model can reliably compensate for contradictory or obsolete documents.
- Over-large chunks: Relevant passages become noisy and exceed the model’s useful context.
- Over-small chunks: Definitions and conditions lose the surrounding information needed to interpret them.
- Vector-only search: Exact identifiers and rare terms may be missed.
- No abstention path: The system invents an answer when evidence is absent.
- Unmeasured changes: New embeddings, prompts, or chunking rules silently reduce quality.
- Ignoring latency and cost: Reranking, large contexts, and repeated retrieval can make a product unusable at scale.
Treat RAG as a continuously evaluated system, not a one-time prompt-engineering exercise. For workflows that require multiple specialised steps, compare RAG with agentic orchestration and review patterns from building distributed systems with AI agents.
RAG in Indian products
India-specific deployments need practical attention to language coverage, intermittent connectivity, data residency, consent, and affordability. Retrieval quality may vary sharply between English and regional-language content. Test on the actual scripts, spelling variants, code-mixed queries, and voice-transcribed input your users produce.
For consumer products, keep responses concise and provide escalation to a human or official channel. A multilingual support assistant can combine RAG with the implementation considerations in building multilingual chatbots for Indian startups. For large public or education-facing systems, design for low-bandwidth access, cache safe responses, and avoid exposing confidential documents through broad search.
RAG can also complement voice interfaces, but speech recognition errors should be logged separately from retrieval failures. If you are exploring that interface, see building a voice agent with Whisper and ElevenLabs.
A launch checklist
Before production, verify that:
- Every answer can identify its supporting source or clearly abstain.
- Retrieval has been tested on real, multilingual, and adversarial queries.
- Permissions are enforced before retrieval results enter the prompt.
- Documents have owners, versions, review dates, and deletion procedures.
- Quality, latency, token usage, cost, and failure rates are monitored.
- Users can report errors and reach a human when required.
- Regression tests run whenever the corpus, retriever, prompt, or model changes.
The strongest RAG systems are modest about what they know and rigorous about showing why they know it. Build the retrieval and evaluation foundations first; model selection then becomes an optimisation rather than a gamble.