Start with the right problem
Fine tuning LLMs for Indian law is useful when a general model needs to produce consistently structured legal work, follow Indian citation conventions, understand procedural vocabulary, or operate in a controlled deployment environment. It is not a shortcut for loading every current statute and judgment into model weights.
A robust legal product usually combines three layers:
- A capable base model for language understanding and reasoning.
- Fine-tuning or adapters for behaviour, formatting, terminology, and task performance.
- Retrieval-augmented generation (RAG) for current, source-linked law and case material.
Builders should first define a narrow workflow: judgment summarisation, provision comparison, case-law retrieval, first-pass contract review, litigation chronology, or multilingual legal assistance. Each workflow needs different data, evaluation criteria, and safeguards. The practical principles in this guide to fine-tuning LLMs on custom data are a useful starting point, but legal systems require stricter provenance and review.
Why Indian legal AI is difficult
Indian legal information is not a single, uniform corpus. It spans the Constitution, Central and State legislation, rules, notifications, tribunal decisions, Supreme Court judgments, High Court judgments, subordinate-court material, and procedural practice. Documents also vary sharply in quality and structure.
Key challenges include:
- Authority and hierarchy: A Supreme Court ratio, a High Court observation, an order without precedential value, and an advocate’s submission must not be treated as equivalent.
- Temporal validity: Provisions change, judgments are overruled, and notifications may apply only for a defined period or jurisdiction.
- Legacy and current terminology: Products must handle IPC, CrPC, and Evidence Act references alongside BNS, BNSS, and BSA terminology.
- Long documents: Judgments often contain facts, submissions, findings, obiter, dissenting opinions, and operative directions in one file.
- Language variation: English legal drafting may sit alongside Hindi or another Indian language, transliteration, abbreviations, and local administrative terms.
- Messy source files: Scanned PDFs, OCR errors, repeated headers, missing paragraphs, and inconsistent citations can silently corrupt training data.
Do not solve these problems by adding more raw documents to a training set. Better metadata, segmentation, retrieval, and evaluation usually create larger gains.
Build a legally defensible dataset
1. Establish source and licensing controls
Create a source register before collecting data. Record the publisher, URL, access date, licence or terms, jurisdiction, document type, court, date, and revision status. Prefer authoritative government and court sources where available, and obtain permission for commercial databases or restricted content.
For every document, preserve the original file and a processed version. Keep hashes, extraction logs, OCR confidence, and transformation history so that a disputed answer can be traced back to its source.
2. Separate corpus types
Do not mix all legal text into one undifferentiated dataset. Maintain separate collections for:
- Constitutional provisions and statutes.
- Rules, regulations, circulars, and notifications.
- Judgments and orders, with court and bench metadata.
- Pleadings, agreements, and anonymised transactional documents.
- Procedure manuals and structured legal workflows.
- High-quality instruction-answer examples created or reviewed by lawyers.
A statute is authoritative text; a judgment explains how a provision was applied; an instruction example teaches the model how to respond. These serve different purposes and should be labelled accordingly.
3. Clean and structure judgments
Parse judgments into meaningful sections such as facts, issues, arguments, statutory provisions, precedents considered, reasoning, holding, directions, and separate opinions. Retain paragraph numbers, citations, dates, judges, court, case number, and links to the original source.
Remove boilerplate and OCR noise, but never remove content merely because it appears repetitive. Run automated checks for broken section numbers, impossible dates, missing pages, duplicated paragraphs, and citation patterns. Sample the output manually with legal reviewers before using it for training.
4. Protect personal data
Anonymise names, addresses, phone numbers, financial details, medical information, and identifiers where they are not essential to the task. Maintain a documented policy for masking versus retaining legally relevant facts. Treat privacy as a dataset and access-control problem, not only as a model-training setting. Assess obligations under the Digital Personal Data Protection framework and applicable professional confidentiality requirements.
Fine-tuning strategy: what to train and what to retrieve
Use supervised fine-tuning (SFT) for repeatable tasks and response behaviour. Examples might ask the model to extract issues from a judgment, produce a structured case brief, distinguish ratio from argument, or draft a source-linked research memo. Each example should specify jurisdiction, date context, document type, and the required uncertainty behaviour.
Use parameter-efficient fine-tuning (PEFT), especially LoRA or QLoRA, for most early experiments. These methods train small adapter weights while leaving the base model largely frozen, reducing compute and making it easier to maintain separate adapters for tasks such as criminal law, tax, contracts, or legal translation.
RAG should supply changing knowledge. Index statutes and judgments with rich metadata, including court, date, section, subject, jurisdiction, language, and whether the text is current. Retrieval should filter by these fields before semantic ranking. The generated answer should cite document title, paragraph or section, court, date, and a stable source link wherever possible.
Fine-tuning is generally the wrong tool for memorising a rapidly changing statute book. It can make a model sound authoritative while preserving outdated information. For a deeper comparison of architecture choices, review best practices for fine-tuning LLMs on custom data.
Evaluate legal AI like a legal product
A high benchmark score is not enough. Build a test set that is never used for training and include adversarial cases. Measure:
- Citation accuracy: Does every cited authority exist and support the claim?
- Authority ranking: Does the system distinguish binding authority from persuasive or irrelevant material?
- Temporal accuracy: Does it identify repealed, amended, or prospective provisions?
- Extraction quality: Are parties, dates, sections, issues, holdings, and directions correct?
- Abstention: Does it say that the record is insufficient instead of inventing an answer?
- Multilingual fidelity: Does translation preserve legal meaning and defined terms?
- Operational performance: Latency, cost, document throughput, and failure recovery.
Have advocates, in-house counsel, law researchers, and domain specialists score outputs using a written rubric. Track severe failures separately from minor style errors. A system that produces polished but unsupported conclusions should fail release review.
Design safeguards before deployment
Add retrieval filters, source display, confidence or evidence indicators, and an explicit “insufficient authority” response. Prevent the model from presenting itself as a lawyer or making unreviewed recommendations in high-stakes matters. Log prompts, retrieved passages, model version, adapter version, and final output under appropriate access controls.
Use a human review queue for criminal, constitutional, immigration, employment, medical, and financial matters. Red-team prompt injection through uploaded pleadings and hostile web content. Test whether confidential material can leak through retrieval, logs, fine-tuning examples, or model outputs.
For multilingual products, evaluate language quality separately rather than assuming an English score transfers to Hindi, Tamil, Bengali, or other languages. Teams building language-aware products can also study this guide to open-source vision-language models for Indian languages, particularly where scanned judgments and mixed-script documents are involved.
A practical 90-day build plan
Days 1–20: scope and data. Select one workflow, define unacceptable errors, map authoritative sources, resolve licensing, and create a small reviewed corpus.
Days 21–45: baseline and retrieval. Test a strong base model without fine-tuning. Build metadata-aware retrieval, source citations, document parsing, and a held-out evaluation set.
Days 46–70: adapters and task data. Create several hundred to a few thousand high-quality examples, run LoRA or QLoRA experiments, and compare against the RAG-only baseline.
Days 71–90: review and pilot. Conduct lawyer-led evaluation, privacy testing, citation audits, adversarial testing, and a limited pilot with full logging. Expand only when the system demonstrates measurable improvement on the target workflow.
Frequently asked questions
Is a larger model always better?
No. A smaller model with reliable retrieval, clear prompts, and a narrow task can outperform a larger model on cost, latency, and factual grounding. Choose based on the evaluation set, not model reputation.
How much training data is required?
There is no universal number. A few hundred carefully reviewed examples can establish a format; broader task coverage may require several thousand. Diversity, correctness, metadata, and held-out testing matter more than raw row count.
Can fine-tuning replace legal research?
No. Fine-tuning improves task behaviour; it does not guarantee current law, authority, or legal correctness. Keep a source-grounded retrieval layer and lawyer review for consequential outputs.
Can a startup run this on Indian infrastructure?
Yes, depending on model size, quantisation, throughput, and security requirements. Open-weight models and adapters can support controlled deployment, but teams must budget for GPU access, monitoring, storage, backup, and secure document handling—not just training.
Indian legal AI will be won by teams that treat provenance, evaluation, and workflow design as core engineering. If you are building a source-grounded legal system for Indian users, explore Indian open-source AI developer projects and consider applying to AI Grants India for support, mentorship, and non-dilutive funding.