0tokens

Apply for AI Grants India

Financial support for innovators building the future of AI in India.

Apply now

Chat · rag systems

RAG Systems: Architecture, Evaluation and Production Guide

  1. aigi

    RAG systems—retrieval-augmented generation systems—combine a language model with an external knowledge source. Instead of asking a model to answer only from its training data, a RAG pipeline retrieves relevant documents at query time and supplies them as context for the response.

    This pattern is especially useful for Indian builders working with changing policies, internal documents, multilingual content, product catalogues, support records, and domain-specific research. RAG is not a guarantee of factual answers, but a practical way to ground responses in sources that an organisation can update, inspect, and govern.

    How RAG systems work

    A production RAG system usually has two paths:

    • Ingestion path: collect, clean, split, enrich, embed, and index documents.
    • Query path: understand a user question, retrieve relevant passages, optionally rerank them, and generate an answer with citations or source references.

    A typical query flows through these stages:

    1. A user submits a question in English, Hindi, or another supported language.
    2. The application rewrites or classifies the query when necessary.
    3. An embedding model converts the query into a vector representation.
    4. A vector database, keyword index, or hybrid search engine retrieves candidate passages.
    5. A reranker selects the most useful evidence.
    6. The language model generates a response using the retrieved context.
    7. The application applies safety, access-control, citation, and confidence checks before displaying the result.

    The model is only one component. Retrieval quality, document structure, permissions, latency, and evaluation often determine whether the product is useful.

    Core architecture choices

    Document preparation and chunking

    Poorly prepared documents produce poor retrieval. Extract headings, tables, page numbers, authorship, timestamps, language, department, and access permissions before indexing. Chunk by meaning rather than using a fixed character count wherever possible. A policy clause, product specification, or legal definition should remain intact when it is split.

    Store metadata alongside every chunk. Useful fields include document ID, version, source URL, publication date, jurisdiction, language, and business unit. This enables filters such as “only current Maharashtra policies” or “only documents the user is authorised to view.”

    Embeddings and search

    Vector search is effective when the query and relevant passage express the same idea using different words. Keyword search remains valuable for exact identifiers, invoice numbers, legal sections, medicine names, and part codes. Hybrid retrieval combines both methods and is often a stronger default than vector search alone.

    For Indian deployments, test retrieval across English, Hindi, and the actual language mix used by customers. Translating every query into English may lose names, local terminology, or regulatory nuance. Evaluate models on representative data instead of choosing solely by benchmark rankings.

    Reranking and context assembly

    Initial retrieval may return many loosely related passages. A cross-encoder or specialised reranker can score candidates against the full query and improve precision. Then assemble a compact context window with deduplicated passages, source labels, and document dates.

    More context is not automatically better. Irrelevant passages can distract the model, increase cost, and create contradictory evidence. The application should instruct the model to say when the supplied sources do not support an answer rather than fill gaps from memory.

    Builders planning complex workflows should also understand how to build multi-agent AI orchestration systems, but adding agents does not replace a well-designed retrieval layer. Start with a single, observable pipeline and introduce additional agents only for a measured need.

    Where RAG systems deliver value

    RAG is a strong fit when knowledge changes frequently or is too private, specialised, or extensive to place in a prompt. Common applications include:

    • Internal knowledge assistants for policies, standard operating procedures, and engineering documentation.
    • Customer support grounded in product manuals, tickets, and approved responses.
    • Research tools that summarise papers, government notifications, or market reports with citations.
    • Enterprise search across PDFs, spreadsheets, wikis, and ticketing systems.
    • Compliance workflows that identify relevant clauses and highlight evidence for review.
    • Education products that answer questions from approved course material.

    A RAG system can also reduce repetitive responses in support and operations; however, teams should pair retrieval with response templates and escalation rules. The guidance on reducing repetitive responses in LLM applications is useful when designing that interaction layer.

    RAG is less suitable when the task requires exact computation, transactional updates, or deterministic database queries. Use tools, APIs, SQL, and validation rules for those operations, with RAG supporting explanation rather than replacing the system of record.

    Evaluation: measure the pipeline, not just the answer

    A convincing demo is not evidence of a reliable RAG system. Create a test set from real or carefully anonymised user questions. Include ambiguous queries, multilingual questions, outdated documents, adversarial prompts, and questions with no answer in the corpus.

    Track separate metrics:

    • Retrieval recall: whether the relevant source appears in the retrieved set.
    • Retrieval precision: how much of the retrieved context is actually useful.
    • Faithfulness: whether the answer is supported by the supplied evidence.
    • Answer relevance: whether it addresses the user’s question.
    • Citation accuracy: whether cited sources genuinely support the claims.
    • Abstention quality: whether the system declines or escalates when evidence is insufficient.
    • Operational metrics: latency, token usage, failure rate, and cost per request.

    Use automated evaluation for scale, but manually review high-risk cases. For healthcare, finance, education, and government workflows, establish a human-review path and log the evidence used for each answer.

    Security, privacy, and governance

    RAG introduces a data-access problem: retrieval must respect the permissions attached to the original content. Apply authorisation filters before context reaches the model, not after generation. Separate tenants, encrypt stored documents and embeddings, redact sensitive personal data where possible, and define retention policies.

    Treat retrieved documents as untrusted input. A document may contain prompt-injection instructions designed to manipulate the model. Keep system instructions separate, restrict tool permissions, validate outputs, and monitor suspicious content. For privacy-sensitive deployments, secure local-first operating systems for privacy offers relevant design considerations, although local hosting alone does not solve governance.

    Maintain document versions and deletion workflows. If a source is withdrawn, the corresponding chunks, caches, summaries, and evaluation references should be identifiable and removable. Record model version, retriever version, prompt version, and source IDs for reproducibility.

    Cost and scaling decisions

    Begin with a narrow corpus and a measurable use case. A managed vector database may accelerate an early launch, while PostgreSQL with a vector extension can reduce operational complexity for smaller systems. At scale, separate ingestion workers from query services, cache stable results, batch embedding jobs, and stream document updates instead of rebuilding the entire index.

    Monitor p50 and p95 latency, retrieval failures, context size, model spend, and user feedback. Scaling backend infrastructure for AI applications provides a useful framework for queues, observability, storage, and service boundaries. For high-throughput products, runtime and inference choices also matter; see this guide to a highly performant runtime for AI applications.

    A practical build plan

    1. Define one user problem and the source of truth.
    2. Assemble a small, permission-aware document corpus.
    3. Build ingestion with metadata, versioning, and deletion support.
    4. Implement hybrid retrieval and return source references.
    5. Add reranking only after measuring baseline retrieval.
    6. Create a labelled evaluation set before broad rollout.
    7. Add abstention, human escalation, monitoring, and audit logs.
    8. Pilot with a limited group and review failed queries weekly.
    9. Expand coverage only when quality, cost, and security targets are met.

    The best RAG systems are not the ones with the most elaborate prompts. They are the ones that retrieve the right evidence, respect access boundaries, expose uncertainty, and improve through disciplined evaluation. For Indian startups, that combination can turn a generic chatbot into a dependable product grounded in the organisation’s own knowledge.

    Last updated 24 September 2026

AIGI may be inaccurate. Replies seeded from the guide above.