0tokens

Apply for AI Grants India

Financial support for innovators building the future of AI in India.

Apply now

Chat · how to evaluate RAG pipelines

How to Evaluate RAG Pipelines: Metrics, Tests and Production Checks

  1. aigi

    Retrieval-Augmented Generation (RAG) systems fail in more than one way. A retriever can miss the right passage, a generator can misread correct evidence, or a technically good answer can still be too slow or expensive for users. How to evaluate RAG pipelines therefore requires separate tests for retrieval, generation, end-to-end answers and production behaviour.

    Start with a clear evaluation contract

    Before choosing metrics, define what the system is expected to do. Write down:

    • The user groups and languages it serves
    • The document sources it may cite
    • Whether answers must quote, cite or only summarise evidence
    • The acceptable response time and per-query cost
    • When the system should refuse, ask a clarification or say that evidence is unavailable
    • Which errors are unacceptable, such as invented policy details or incorrect financial guidance

    For an Indian deployment, include realistic language and domain variation: English, Hindi and code-mixed queries; Indian names and addresses; rupee amounts; dates in both DD/MM/YYYY and written formats; and documents from government portals, PDFs and scanned regional-language material. A benchmark that contains only clean English questions will hide important failures.

    Build a representative evaluation set

    Create a labelled dataset from real or carefully reconstructed traffic. Keep a private test set that is not used for prompt or retriever tuning. Each test case should contain:

    • The question and relevant metadata, such as user role or language
    • One or more acceptable evidence passages or document IDs
    • A reference answer, if a single answer is appropriate
    • Required facts and prohibited claims
    • The expected action: answer, clarify, cite, or abstain
    • A difficulty label, such as single-hop, multi-hop, temporal or ambiguous

    Include positive, negative and adversarial examples. Add questions whose answer is absent from the corpus, conflicting versions of a policy, stale documents and near-duplicate passages. For operational systems, sample anonymised production queries and have domain reviewers label them. Version the dataset and record changes so that a metric improvement cannot be confused with an easier test set.

    Teams already maintaining repeatable automated eval pipelines for large language models can reuse the same dataset versioning, regression checks and reporting patterns for RAG.

    Evaluate retrieval separately

    Retrieval quality should be measured before judging the generated answer. If the required evidence never reaches the model, changing the prompt or increasing model size will not solve the underlying problem.

    Useful retrieval metrics include:

    • Recall@k: whether at least one relevant passage appears in the top *k* results. This is vital when the generator receives only a small context window.
    • Precision@k: how much of the retrieved context is relevant. Low precision increases distraction and can expose contradictory text.
    • MRR: how early the first relevant result appears, useful when only the top result is trusted.
    • NDCG@k: a graded ranking metric when some passages are more useful than others.
    • Context utilisation: whether the answer actually uses the relevant retrieved passage rather than unrelated context.

    Measure retrieval by query type and language, not only as one overall average. Compare sparse search, dense embeddings and hybrid search. Test chunk size, overlap, metadata filters, reranking and query rewriting independently. A useful experiment changes one component at a time and records recall, latency, index size and infrastructure cost.

    Do not label a document as relevant merely because it shares keywords. Reviewers should identify the smallest passage that supports the required claim. This catches a common issue: a result can look semantically similar while omitting the clause, exception or date that changes the answer.

    Measure generation and grounding

    Once retrieval is sound, assess the generated response against the evidence. Separate the following dimensions:

    • Correctness: Are the claims factually right?
    • Faithfulness or groundedness: Does every material claim follow from the retrieved context?
    • Completeness: Does the answer include the important facts needed to satisfy the question?
    • Relevance: Does it answer the question without unnecessary background?
    • Citation accuracy: Do citations point to passages that genuinely support the associated claims?
    • Refusal quality: Does the system abstain when evidence is missing or contradictory?

    Exact-match and token-level F1 are useful for short factual answers, IDs and structured fields. ROUGE or BLEU can provide limited diagnostic information for summaries, but n-gram overlap is a poor proxy for factuality in open-ended RAG. Prefer claim-level checks: split an answer into verifiable claims, map each to evidence, and mark unsupported, contradicted and supported claims.

    LLM judges can accelerate review, but they are not ground truth. Use a fixed rubric, blind the judge to system names, test judge agreement against human labels and retain a sample for manual review. For high-stakes applications, require a human decision for safety, legal, medical or financial escalation paths.

    Test the complete pipeline

    End-to-end evaluation should combine quality with operational constraints. Track:

    • Answer success rate by intent and language
    • Groundedness and citation coverage
    • Abstention precision and recall
    • p50, p95 and p99 latency
    • Token usage and cost per successful answer
    • Retrieval, reranking and generation failure rates
    • Context-window truncation and timeout rates
    • User feedback, correction rate and escalation rate

    Run controlled ablations: retrieval only, retrieval plus reranker, different chunking strategies, different prompts and different models. Report confidence intervals or bootstrap ranges rather than presenting tiny score differences as meaningful gains. Monitor quality separately for new documents, old documents and sources with frequent updates.

    For teams building broader high-performance AI pipelines, the same discipline applies: define service-level objectives, isolate bottlenecks and treat observability as part of the system rather than an afterthought.

    Create a practical error taxonomy

    A single score cannot tell an engineer what to fix. Tag failures using a taxonomy such as:

    • Retrieval miss: relevant evidence was not returned
    • Ranking error: relevant evidence was returned too low
    • Chunking error: the passage lacked necessary surrounding context
    • Query interpretation error: the question was rewritten incorrectly
    • Evidence conflict: sources disagreed or had different effective dates
    • Grounding failure: the answer introduced unsupported claims
    • Synthesis failure: the model failed to combine multiple passages
    • Citation failure: the citation did not support the claim
    • Abstention failure: the system answered when it should have refused
    • Product failure: latency, formatting, access or language handling failed

    For every failure, record the query, retrieved IDs, scores, prompt version, model version, answer, citations and expected behaviour. This makes regression debugging reproducible and lets teams prioritise fixes by user impact rather than by metric alone.

    Use continuous evaluation in production

    Run a small, stable regression suite on every change to embeddings, chunking, retrieval settings, prompts or models. Run a broader suite before release, and sample live traffic for weekly human review. Alerts should cover both quality and operations: sudden drops in retrieval recall, increased unsupported claims, rising abstentions, citation mismatches, latency spikes and cost growth.

    Keep training, tuning and test data separated. Log document versions and effective dates, especially for policies, schemes and regulations that change over time. Protect personal data in query logs through minimisation, access controls and retention limits.

    A production evaluation pipeline benefits from the same reproducibility practices used in building custom machine learning pipelines on GitHub: pinned dependencies, traceable artefacts, versioned configurations and reviewable changes.

    A compact evaluation checklist

    Before shipping a RAG system, confirm that you can answer yes to these questions:

    • Does the test set represent real users, languages and document formats?
    • Is retrieval evaluated independently from answer generation?
    • Are unsupported claims and citation errors measured at claim level?
    • Does the system abstain appropriately when evidence is absent?
    • Are latency, cost, privacy and access-control failures included?
    • Can every score be traced to a dataset, model, prompt and index version?
    • Are changes tested against a fixed regression suite and a fresh holdout set?

    The goal is not the highest benchmark score. It is a system that retrieves the right evidence, uses it faithfully, communicates uncertainty and remains dependable under real traffic. Evaluate each layer separately, connect metrics to an actionable failure taxonomy, and keep human review in the loop wherever an incorrect answer carries material risk.

    Last updated 23 September 2026

AIGI may be inaccurate. Replies seeded from the guide above.