0tokens

Apply for AI Grants India

Financial support for innovators building the future of AI in India.

Apply now

Chat · rag retrieval pipelines

RAG Retrieval Pipelines: Architecture, Evaluation and Production

  1. aigi

    Retrieval-augmented generation (RAG) is no longer just a way to attach documents to a chatbot. For Indian startups, enterprises and public-service teams, a well-designed RAG retrieval pipeline can turn policies, product records, research, support tickets and regional-language content into a usable interface—without retraining a language model every time the underlying information changes.

    The important shift is to treat RAG as a search and evidence system, not a prompt template. The quality of the final answer depends on every stage: source preparation, indexing, query understanding, retrieval, reranking, context construction, generation and evaluation.

    What a RAG retrieval pipeline does

    A RAG pipeline retrieves relevant evidence before asking a language model to answer. In a typical request, the system:

    1. Accepts a user question.
    2. Rewrites or expands the question when useful.
    3. Searches a structured index using keyword, vector or hybrid retrieval.
    4. Reranks candidate passages for relevance.
    5. Builds a compact context with source metadata.
    6. Generates an answer constrained by the retrieved evidence.
    7. Returns citations, confidence signals or a safe fallback.

    This architecture separates knowledge storage from language generation. A model can remain unchanged while a team updates a government circular, inventory table, clinical protocol or internal handbook. That makes RAG particularly useful where information changes frequently or must be traceable.

    RAG is not automatically factual. If retrieval returns irrelevant, incomplete or outdated passages, the model can produce a fluent but unsupported response. Grounding therefore begins with the index and ends with measurable production checks.

    Core architecture

    1. Source ingestion and preparation

    Start with an inventory of authoritative sources. These may include PDFs, HTML pages, databases, APIs, spreadsheets, tickets and scanned documents. Record ownership, update frequency, access permissions and retention requirements for each source.

    The ingestion layer should:

    • Extract text while preserving headings, tables, page numbers and document identifiers.
    • Apply OCR to scans and validate OCR quality, especially for Indic scripts.
    • Remove boilerplate, duplicate pages and navigation text.
    • Attach metadata such as language, department, date, geography, document type and access level.
    • Preserve the original file or a stable reference for citation and audit.

    For Indian deployments, do not assume that English-only parsing is sufficient. Hindi, Tamil, Bengali, Marathi and mixed-language queries may require language-aware tokenisation, transliteration handling and multilingual embeddings. A system serving district offices may also need to handle inconsistent spelling, abbreviations and code-mixed queries.

    2. Chunking and indexing

    Chunking determines what the retriever can find. Fixed-size chunks are easy to implement but often split definitions from exceptions or tables from their headings. Structure-aware chunks—based on sections, clauses, questions and answers—usually preserve meaning better.

    Useful chunk metadata includes:

    • Parent document and section title
    • Page or paragraph location
    • Effective date and version
    • Language and jurisdiction
    • Product, customer or permission scope

    Create at least two retrieval paths when the corpus justifies it:

    • Lexical retrieval, such as BM25, for exact names, policy numbers, legal terms and model identifiers.
    • Dense retrieval, using embeddings, for semantic similarity and paraphrased questions.

    Hybrid retrieval is often the strongest baseline because it combines exact matching with semantic recall. It is also easier to debug than relying on vector search alone.

    Teams building broader data workflows can compare this design with guidance on building high-performance AI pipelines and implementing scalable ML pipelines.

    3. Query processing and retrieval

    Before searching, classify the query. A question asking for a precise policy clause needs different handling from a request to summarise a long case file. Query processing may include spelling correction, translation, language identification, entity extraction and decomposition into multiple sub-questions.

    Retrieve a generous candidate set first, then rerank it. A cross-encoder or lightweight relevance model can inspect the question and passage together, improving precision at the cost of latency. Apply metadata filters early where possible—for example, restricting results to a user’s organisation, region or current policy version.

    Do not silently mix conflicting versions. If two documents disagree, the application should expose the conflict, prefer the latest approved source where policy allows, or route the question for review.

    4. Context assembly and generation

    The generator should receive only the evidence needed to answer. Large, noisy prompts increase cost and can reduce accuracy. Context assembly should deduplicate overlapping chunks, preserve document order where it matters and label each passage clearly.

    A production prompt should specify that the model must:

    • Answer only from supplied evidence when the task requires grounding.
    • Say when the evidence is insufficient.
    • Distinguish facts from inference.
    • Cite document names, dates or page references.
    • Avoid exposing content the user is not authorised to see.

    For customer-facing systems, this can be paired with the future of voice agents in customer service, where retrieval latency, interruption handling and concise spoken citations become additional design constraints.

    How to evaluate a RAG pipeline

    Evaluation must separate retrieval quality from generation quality. Test with a representative question set covering easy lookups, ambiguous queries, multilingual inputs, outdated documents, adversarial prompts and unanswerable questions.

    Track retrieval metrics such as:

    • Recall@k: whether the required evidence appears in the top-k results.
    • MRR or nDCG: whether relevant evidence is ranked near the top.
    • Context precision: how much retrieved material is actually useful.
    • Citation coverage: whether claims have supporting sources.

    Track answer metrics such as groundedness, correctness, completeness, refusal quality, latency and cost per request. Human review remains essential for high-impact domains. Automated judges can accelerate testing, but they should be calibrated against expert-labelled examples and monitored for bias.

    Use an evaluation set that reflects Indian operating conditions: low-bandwidth users, code-mixed language, scanned forms, local place names, policy changes and permissions across departments. The dedicated guide on evaluating RAG pipelines covers metrics, test design and production checks in greater depth.

    Production risks and controls

    The most common failure modes are operational rather than theoretical:

    • Stale content: schedule incremental ingestion and show effective dates.
    • Poor OCR: retain confidence scores and route low-quality pages for review.
    • Prompt injection in documents: treat retrieved text as untrusted data, not instructions.
    • Access-control leakage: filter documents before context assembly and test with adversarial users.
    • Latency and cost: cache embeddings and frequent queries, limit reranking depth and choose models by task.
    • Unsupported answers: implement abstention, citation checks and escalation paths.
    • Data residency and privacy: minimise sensitive fields, encrypt stores, log access and define retention rules.

    For research-heavy use cases, large language models for scientific knowledge retrieval offers a useful comparison of evidence handling and literature-scale search. For local information systems, RAG can also complement generative AI integration into local information systems, provided data ownership and governance are explicit.

    A practical implementation path

    A reliable first version does not need a complex agent framework. Build a narrow vertical slice:

    1. Select one authoritative corpus and define permitted users.
    2. Create a gold set of real questions with expected evidence.
    3. Implement clean ingestion, structure-aware chunking and hybrid search.
    4. Add reranking, citations and an explicit “not found” response.
    5. Measure retrieval, answer quality, latency and cost.
    6. Pilot with domain users and inspect failed queries weekly.
    7. Add multilingual support, query decomposition or tool use only when evaluation shows a need.

    A modular design also makes it easier to swap embedding models, vector stores or language-model providers without rewriting the application. Keep retrieval, ranking, prompting and evaluation interfaces separate, and version your index alongside source changes.

    What good RAG looks like in 2026

    The strongest RAG systems are not those with the longest context windows or most elaborate agent loops. They are systems with authoritative sources, strong retrieval recall, clear citations, disciplined access control and measurable abstention. For Indian builders, success also means handling multilingual data, uneven document quality, cost-sensitive deployment and real institutional workflows.

    RAG should be chosen when external knowledge changes, must be cited or cannot be safely embedded in model weights. It is not a substitute for clean databases, deterministic business rules or model fine-tuning where behaviour—not knowledge—must change. Used with those tools, it provides a practical route from fragmented organisational data to accountable AI applications.

    Last updated 24 September 2026

AIGI may be inaccurate. Replies seeded from the guide above.