Why Indian legal benchmarking needs its own design
Benchmarking LLM performance on Indian legal text is not a matter of running a general knowledge test and ranking models by accuracy. A useful benchmark must measure whether a system can retrieve the right authority, understand Indian procedural context, distinguish current law from repealed or amended provisions, and communicate uncertainty when the evidence is incomplete.
This matters more after the transition from the Indian Penal Code, Code of Criminal Procedure, and Indian Evidence Act to the Bharatiya Nyaya Sanhita (BNS), Bharatiya Nagarik Suraksha Sanhita (BNSS), and Bharatiya Sakshya Adhiniyam (BSA). A model can produce a fluent explanation while silently applying the wrong statute, citing a nonexistent judgment, or confusing a ratio with an observation. Those are benchmark failures, not minor quality issues.
For product teams, the evaluation target should be a defined legal workflow—not an abstract score. Examples include finding authorities for a bail memo, extracting obligations from a contract, summarising a Supreme Court judgment, or mapping an older IPC reference to the corresponding current provision.
Define tasks before choosing metrics
Start with a task matrix that specifies the input, expected output, acceptable sources, and risk level. A balanced Indian legal benchmark should include:
- Retrieval: Find relevant judgments, statutory provisions, rules, circulars, or tribunal orders from a dated corpus.
- Extraction: Identify parties, court, bench, dates, sections, relief sought, decision, and cited authorities.
- Summarisation: Produce a structured summary that separates facts, issues, arguments, reasoning, holding, and disposition.
- Question answering: Answer only from supplied authorities, with pinpoint citations and an “insufficient evidence” option.
- Legal classification: Classify matters by issue, jurisdiction, procedural stage, or likely document type.
- Statutory mapping: Link historical provisions to current provisions, while flagging changes in wording and scope.
- Draft assistance: Generate a constrained clause, notice, or pleading section using a supplied template and source pack.
- Multilingual processing: Translate, retrieve, or summarise documents involving Hindi, Tamil, Marathi, Bengali, Kannada, Telugu, Malayalam, Gujarati, or other Indian languages.
A benchmark should contain both ordinary examples and difficult cases: lengthy judgments, scanned PDFs, conflicting authorities, citations with formatting variation, quoted pleadings, and documents containing OCR errors.
Build a defensible Indian legal corpus
Do not treat a large document collection as a benchmark by default. Record provenance and versioning for every item. At minimum, capture the court or authority, date, jurisdiction, language, document type, source URL or identifier, applicable legislation at the time, and whether the text is an official or secondary reproduction.
Separate training, development, and test sets by matter or judgment, not by random paragraphs. Otherwise, nearly identical passages and cited authorities can leak across splits and inflate results. A time-based split is particularly valuable: train on older material and test on later decisions, amendments, or newly notified rules.
Keep a gold answer key with evidence spans. For each question, annotators should mark the passage supporting the answer, identify acceptable alternative answers, and note whether the issue is unsettled. Legal experts should adjudicate disagreements rather than forcing a single answer where authorities genuinely conflict.
Teams building a document pipeline can also use the principles in AI legal document automation in India, especially for OCR quality, metadata capture, access controls, and human review.
Metrics that expose real failure modes
A single aggregate score hides the risks that matter most. Report results by task, court level, language, document length, and legal domain.
- Retrieval: Recall@k, precision@k, nDCG, and authority coverage. Measure whether the system retrieves the controlling or most relevant authority—not merely textually similar documents.
- Extraction: Exact match and token-level F1 for fields such as sections, dates, parties, and outcomes. Add a field-level error rate for critical items.
- Summarisation: Use factual consistency, citation coverage, omission rates, and expert ratings. ROUGE can describe overlap, but it cannot establish legal correctness.
- Question answering: Score answer correctness, evidence entailment, citation precision, citation recall, and abstention quality. Penalise unsupported certainty more heavily than a transparent refusal.
- Statutory mapping: Evaluate provision accuracy, effective-date awareness, and whether the model explains that a mapping is approximate rather than identical.
- Drafting: Check compliance with instructions, inclusion of required clauses, source fidelity, prohibited claims, and review effort—not just stylistic quality.
- Translation and cross-lingual retrieval: Measure meaning preservation, legal-term accuracy, named-entity preservation, and retrieval recall across languages.
For every citation, verify the case exists, the cited court and date are correct, the proposition appears in the authority, and the quotation is faithful. A fabricated SCC, AIR, or neutral citation should be recorded as a critical error.
Evaluate RAG and fine-tuning separately
A legal assistant is usually a system, not just a model. Benchmark the language model, retriever, reranker, prompt, context window, and citation layer independently before measuring the complete application.
For a RAG evaluation, create a closed-book and open-book track. In the closed-book track, the model answers without retrieval; in the open-book track, it receives a controlled source pack. Compare retrieval recall with final answer accuracy to identify whether failures originate in search or reasoning. Test stale, contradictory, and irrelevant documents to measure whether the system respects dates and source hierarchy.
Fine-tuning may improve formatting, classification, or drafting style, but it does not guarantee current law. RAG is generally better for changeable statutes and citation-grounded research, provided the corpus is authoritative and versioned. A hybrid system should be tested for both benefits and regressions, including whether fine-tuning causes the model to rely on memorised but outdated provisions.
Add multilingual, OCR, and long-context tracks
Indian legal access cannot be measured through English judgments alone. Create paired examples where a regional-language FIR, deposition, order, or revenue document must be translated, summarised, or matched to an English authority. Evaluate legal meaning, not literal translation. Penalise dropped negations, altered names, incorrect section numbers, and changes in dates or quantities.
Include native-script PDFs, mixed-script text, handwritten or poor-quality scans, and code-switching. Open-source vision-language models for Indian languages can be useful candidates for these tracks, but compare them against OCR-plus-text baselines rather than assuming a multimodal model is automatically better.
Long-context testing should use realistic document bundles: a judgment, pleadings, annexures, prior orders, and statutory extracts. Ask targeted questions whose answers appear near the beginning, middle, and end. Report performance by context length and measure whether additional irrelevant material increases unsupported claims.
Fairness, security, and human evaluation
Historical case data may encode unequal treatment by gender, caste, religion, disability, region, or socioeconomic status. Test paired scenarios that change a sensitive attribute while keeping legally relevant facts constant. Review recommendations for bail, sentencing, credibility, and risk classification with qualified advocates and domain experts.
Also test prompt injection inside uploaded judgments, poisoned citations, confidential data leakage, and instructions that ask the model to ignore source restrictions. A production benchmark should record latency, cost, refusal behaviour, audit-log completeness, and reviewer correction time alongside accuracy.
Use a double-review protocol for high-risk outputs. Have two independent legal reviewers score correctness, source support, completeness, and harmful overstatement; adjudicate disagreements; and publish confidence intervals where the test set is large enough. Keep a failure catalogue with reproducible inputs, model version, retrieval snapshot, and remediation status.
A practical release gate for 2026
Before deployment, require the system to meet thresholds by workflow rather than one overall score. A sensible release gate includes:
- No critical fabricated citations on a challenge set.
- High precision for statutory sections, dates, parties, and case outcomes.
- Reliable abstention when the source pack does not support an answer.
- Documented performance across relevant Indian languages and courts.
- No material fairness regression on paired tests.
- Human sign-off for outputs used in filings, advice, or decisions.
- Versioned evaluation reports whenever the model, retriever, corpus, or prompt changes.
Founders working on legal AI should treat the benchmark as a product asset. Publish task definitions, data limitations, error categories, and known exclusions so customers can judge whether results apply to their workflow. For broader compliance and operational controls, see how to automate legal compliance with AI in India.
A strong benchmark does not prove that an LLM is a lawyer. It shows where the system is dependable, where it must defer, and what controls are needed before a legal professional relies on it. That distinction is the foundation for responsible Indian legal AI.