0tokens

Apply for AI Grants India

Financial support for innovators building the future of AI in India.

Apply now

Chat · retrieval systems for ai

Retrieval Systems for AI: Architecture, RAG and Evaluation

  1. aigi

    Retrieval systems for AI are the layer between an organisation’s data and its models. They decide which documents, records, images, events or knowledge-graph entities an application can find—and therefore what an AI system can use in an answer or decision. A capable language model cannot compensate for missing, stale or poorly ranked evidence.

    For Indian builders, retrieval is especially important because production data is distributed across government portals, enterprise software, PDFs, regional-language content, scanned documents and disconnected databases. A useful system must optimise not only for search quality, but also for latency, permissions, auditability and cost.

    What a retrieval system for AI does

    A retrieval system converts a user request or model-generated query into relevant evidence. The basic flow is:

    • Ingest data from files, databases, APIs, websites, event streams or knowledge graphs.
    • Clean and enrich content with metadata such as language, date, department, document type and access level.
    • Index the content using keywords, vectors, structured fields or graph relationships.
    • Retrieve and rank candidate results against the query.
    • Pass grounded context to an AI model, analyst or downstream workflow.
    • Measure outcomes such as relevance, citation accuracy, latency and user feedback.

    This is broader than a vector database. A production retrieval system includes ingestion, parsing, chunking, indexing, query rewriting, ranking, access control, observability and data governance.

    Core architectures to choose from

    Keyword retrieval

    Traditional inverted indexes match terms, phrases and fields. They remain excellent for exact identifiers, legal clauses, product codes, names and dates. Keyword search is fast, explainable and usually inexpensive, but it may miss relevant content when the user’s wording differs from the source.

    Vector retrieval

    An embedding model represents queries and content as vectors, enabling similarity search by meaning. Vector retrieval is useful for support knowledge bases, semantic document search and multilingual discovery. Its results depend heavily on the embedding model, chunking strategy and the quality of the source text.

    Hybrid retrieval

    Hybrid search combines keyword and vector results, then reranks them. It is often the strongest default for enterprise applications because it handles both exact terms and semantic intent. Teams should test reciprocal-rank fusion, weighted scoring and cross-encoder reranking rather than assuming that one method will work for every corpus.

    Knowledge-graph retrieval

    Graphs represent entities and relationships such as a supplier, contract, district, scheme or medical condition. They support relationship-heavy questions and can add structure where documents alone are ambiguous. Graph retrieval is valuable when provenance and multi-step reasoning matter.

    Retrieval-augmented generation

    RAG retrieves evidence before generation. A typical RAG pipeline parses documents, creates chunks and embeddings, retrieves candidates, reranks them, constructs a prompt and requires the model to answer from the supplied context. RAG is preferable to fine-tuning when facts change frequently or answers must cite source material. For model customisation decisions, compare retrieval with the best practices for fine-tuning LLMs on custom data.

    Designing a reliable retrieval pipeline

    Start with the user task, not the database technology. Define what a successful answer or action requires: a correct document, a complete set of records, a current value, a cited clause or a safe refusal when evidence is missing.

    Ingestion and parsing: PDFs, spreadsheets and scans need different treatment. Preserve headings, tables, page numbers, document IDs and publication dates. Use OCR for scanned Indian-language material, then validate extracted text because OCR errors can alter names, figures and legal language.

    Chunking: Fixed token windows are easy to implement but often split definitions from their qualifications. Prefer structure-aware chunks based on headings, clauses, tables or conversation turns. Store parent-document references so the application can expand context when needed.

    Metadata and filters: Add jurisdiction, department, language, effective date, source reliability, sensitivity and tenant identifiers. Metadata filters prevent a system from retrieving a document that is semantically relevant but unauthorised or out of scope.

    Query processing: Rewrite vague questions, detect language, extract filters and expand abbreviations. For complex requests, decompose the query into smaller searches, but retain the original intent and enforce limits on tool calls.

    Ranking and context assembly: Retrieve more candidates than you will show, rerank them with a stronger model, remove duplicates and fit only the most useful evidence into the context window. Include source names, dates and stable citations in the prompt and final response.

    Data quality, security and provenance

    Retrieval quality is constrained by source quality. Establish ownership for each corpus and record when it was last updated. Deduplicate documents, identify conflicting versions and mark whether a source is authoritative, supplementary or unverified. Teams working with regulated or high-stakes data should treat data veracity infrastructure for high-stakes AI as a design requirement, not a post-launch feature.

    Apply security at retrieval time. Enforce row-, document- and field-level permissions before context reaches the model. Encrypt data in transit and at rest, isolate tenants, redact personal information where appropriate and log who accessed which source. Under India’s Digital Personal Data Protection framework and sector-specific rules, define purpose limitation, retention and deletion processes early.

    For healthcare deployments, retrieval must also preserve provenance, validation status and clinical review. Systems handling medical evidence should align their verification process with ICMR-compliant medical AI data verification in India.

    How to evaluate retrieval systems

    Do not evaluate only the final chatbot answer. Separate retrieval metrics from generation metrics:

    • Recall@k: whether the relevant evidence appears in the top k results.
    • Precision@k: how many retrieved results are actually useful.
    • MRR or nDCG: whether the best evidence is ranked near the top.
    • Answer faithfulness: whether claims are supported by retrieved sources.
    • Citation completeness: whether important claims have citations.
    • Freshness: whether current documents outrank expired versions.
    • Latency and cost: whether the system meets service-level targets.
    • Abstention quality: whether it declines when evidence is insufficient.

    Create a test set from real queries, including misspellings, multilingual questions, ambiguous terms, long documents and adversarial prompts. Label relevant passages with domain experts. Segment results by language, department, document type and user role; an average score can conceal serious failures in one group. Low-resource Indian language deployments also benefit from carefully curated low-resource language datasets for AI training in India.

    Practical deployment choices for Indian teams

    A small startup can begin with managed search, an embedding service and a relational database with vector support. This reduces operational overhead while the team validates demand. Move to a dedicated vector or hybrid-search platform when scale, filtering, multitenancy or latency justifies it. Keep an exportable canonical corpus so that changing vendors does not mean losing data or evaluation history.

    For sensitive workloads, consider private networking, regional hosting, self-hosted models or a local-first architecture. Offline and edge retrieval can be valuable for field operations with unreliable connectivity; privacy-conscious teams can review approaches to secure local-first operating systems.

    Build observability from the first release. Track empty-result rates, retrieval scores, source usage, user corrections, model refusals, token consumption and latency by stage. When an answer fails, engineers should be able to determine whether the cause was parsing, chunking, indexing, ranking, permissions or generation.

    Common mistakes to avoid

    • Treating a vector database as a complete RAG system.
    • Indexing noisy documents without ownership, dates or access metadata.
    • Using one chunk size for every document type.
    • Evaluating only fluent answers instead of evidence retrieval.
    • Ignoring exact-match search for IDs, statutes and codes.
    • Sending unauthorised records into the model context.
    • Failing to remove superseded policies and duplicate documents.
    • Launching without an abstention path or human review for high-risk actions.

    A build roadmap

    1. Select one narrow, high-value workflow and define success metrics.
    2. Audit the corpus, permissions, languages and update frequency.
    3. Build a baseline with keyword search, then add vectors and hybrid ranking.
    4. Create a labelled evaluation set from real user questions.
    5. Add citations, access controls, monitoring and feedback capture.
    6. Pilot with domain users, review failures weekly and expand the corpus gradually.

    Retrieval systems for AI are infrastructure for trustworthy applications, not merely a search feature. Teams that invest in clean sources, measurable ranking, strict permissions and transparent provenance can build assistants and decision-support tools that remain useful as models, documents and regulations change.

    Last updated 23 September 2026

AIGI may be inaccurate. Replies seeded from the guide above.