0tokens

Apply for AI Grants India

Financial support for innovators building the future of AI in India.

Apply now

Chat · how to fine tune a model using indian legal public data on hugging face

How to Fine-Tune Models on Indian Legal Data with Hugging Face

  1. aigi

    Fine-tuning can make a general language model more useful for Indian legal text—but only when the data, task design, and evaluation are sound. Court judgments, statutes, regulations, and orders have distinctive terminology, long document structures, multilingual references, and citation patterns. A model trained carelessly can reproduce errors, expose personal information, or sound authoritative while giving legally incorrect answers.

    This guide explains how to fine-tune a Hugging Face model using Indian legal public data, with a practical workflow for builders and researchers in 2026. It focuses on task-specific fine-tuning, not training a foundation model from scratch.

    Start with the task, not the dataset

    Choose the exact capability you need before collecting documents. Fine-tuning works best when inputs and expected outputs are clearly defined.

    Common tasks include:

    • Classification: classify documents by court, case type, statute, topic, or procedural stage.
    • Named-entity recognition: identify acts, sections, courts, judges, parties, dates, and locations.
    • Information extraction: extract case outcomes, cited provisions, issues, or hearing details into structured fields.
    • Retrieval and reranking: identify the most relevant judgments for a legal research query.
    • Summarisation: generate a structured summary with facts, issues, reasoning, and disposition.
    • Question answering: answer questions grounded in supplied passages rather than relying on model memory.

    For legal assistants, retrieval-augmented generation is often safer than asking a model to memorise the law. See this practical overview of AI legal document automation in India before deciding whether fine-tuning is necessary.

    Source Indian legal data responsibly

    Potential sources include official Supreme Court and High Court repositories, India Code, legislative department publications, tribunal websites, and datasets released with clear terms. A document being accessible online does not automatically mean it can be redistributed, commercialised, or used without restrictions.

    Create a source register for every file containing:

    • Original URL and access date
    • Court, tribunal, department, or publisher
    • Publication date and document date
    • Licence or terms of use
    • Language and document type
    • Whether personal information appears in the text
    • Hash or identifier for deduplication

    Prefer primary sources over scraped summaries. Legal blogs can help with discovery, but they should not silently become ground truth for a legal model. Check the current terms of each source, remove access credentials and session data, and document your collection process.

    Prepare and de-identify the corpus

    Indian judgments commonly arrive as PDFs, scanned images, HTML pages, or inconsistent text exports. Build a repeatable preprocessing pipeline rather than editing files manually.

    1. Extract text while preserving paragraph boundaries, headings, footnotes, and page references where possible.
    2. Run OCR only on scanned pages and retain the original file for quality checks.
    3. Normalise whitespace, Unicode variants, broken hyphenation, and repeated headers or footers.
    4. Detect duplicates using document identifiers, hashes, and near-duplicate similarity.
    5. Review extraction errors around section numbers, citations, dates, negations, and names.
    6. Redact or pseudonymise unnecessary personal information, especially contact details, addresses, minors’ identities, and sensitive case information.
    7. Store provenance metadata alongside each example.

    Do not train on raw scraped text without a review sample. A small manual audit can reveal OCR failures that would otherwise become learned behaviour.

    Design the dataset format

    Use JSONL or Parquet with explicit fields. For classification, a record might look like this:

    {"text":"...judgment passage...","label":"constitutional_law","court":"Supreme Court of India","year":2022,"source_id":"sc_2022_001"}

    For instruction tuning, separate the instruction, context, and answer:

    {"messages":[{"role":"user","content":"Summarise the supplied passage and cite only information in it.\n\nPassage: ..."},{"role":"assistant","content":"Facts: ...\nIssue: ...\nHolding: ..."}]}

    Keep labels unambiguous and define annotation rules before labelling. If experts disagree, record adjudication rather than hiding disagreement. Split data by case or document, not by random paragraph. Otherwise, passages from the same judgment may appear in both training and test sets, producing misleading scores.

    For broader implementation guidance, use these best practices for fine-tuning LLMs on custom data.

    Choose a suitable Hugging Face model

    Select a model based on language coverage, licence, context length, hardware, and task. For encoder tasks such as classification or entity recognition, a BERT- or RoBERTa-style model is usually more efficient. For instruction following or structured generation, use a compatible causal language model and consider parameter-efficient fine-tuning.

    Before downloading a checkpoint, verify:

    • Commercial and research-use terms
    • Supported Indian languages and scripts
    • Tokeniser behaviour for legal citations and Devanagari or other scripts
    • Context-window limits
    • Existing safety and evaluation documentation
    • Whether the model accepts the intended input format

    A model with fewer parameters and reliable data can outperform a larger model trained on noisy examples.

    Fine-tune with Transformers and Datasets

    Install the core stack:

    pip install transformers datasets evaluate accelerate peft trl

    Load a dataset and tokenizer, then tokenise with a realistic maximum length. Truncation can remove the very paragraph needed to answer a legal question, so inspect length distributions before selecting max_length.

    from datasets import load_dataset
    from transformers import AutoTokenizer
    
    model_id = "your-approved-model"
    tokenizer = AutoTokenizer.from_pretrained(model_id)
    dataset = load_dataset("json", data_files={
        "train": "train.jsonl",
        "validation": "validation.jsonl",
        "test": "test.jsonl"
    })
    
    def tokenize(batch):
        return tokenizer(
            batch["text"],
            truncation=True,
            max_length=2048
        )
    
    tokenised = dataset.map(tokenize, batched=True)

    For large models, use LoRA or another PEFT method instead of updating every parameter. This lowers memory requirements and makes experiments easier to compare. Use gradient accumulation, mixed precision where supported, checkpointing, and a fixed random seed. Keep a run manifest containing the model revision, dataset version, hyperparameters, hardware, and code commit.

    Do not assume the generic Trainer example works unchanged for every task. Classification, token classification, sequence-to-sequence generation, and chat fine-tuning require different data collators, model classes, labels, and evaluation methods.

    Evaluate legal usefulness and failure modes

    Accuracy alone is not enough. Report metrics that match the task:

    • Macro-F1 and per-class recall for imbalanced classification
    • Entity-level precision, recall, and F1 for extraction
    • Exact match and token-level F1 for question answering
    • ROUGE or similar measures only as secondary signals for summarisation
    • Citation precision, groundedness, and unsupported-claim rate for generated answers
    • Latency, memory use, and cost for deployment

    Create a held-out test set from later dates, different courts, or different document sources to measure distribution shift. Include difficult examples: long judgments, OCR noise, bilingual passages, uncommon statutes, conflicting authorities, and questions with insufficient evidence.

    Have legal professionals review a sample using a rubric for factual accuracy, citation validity, completeness, ambiguity, and harmful overconfidence. Evaluate whether the model says “not found in the supplied material” when appropriate. Never present benchmark performance as legal advice quality without task-specific review.

    Publish and deploy with safeguards

    When uploading to Hugging Face, publish a dataset card or model card describing sources, dates, preprocessing, licences, redaction, limitations, intended use, and prohibited uses. Avoid publishing sensitive raw documents or memorised personal information. Restrict private repositories when the licence or risk assessment requires it.

    In production, combine the model with retrieval, source links, confidence thresholds, logging, access controls, and human escalation. Add a clear notice that outputs are informational and require review by a qualified professional. For compliance workflows, pair model output with deterministic checks and an audit trail; this complements guidance on automating legal compliance with AI in India.

    Builders should also plan for updates. Laws change, court interpretations evolve, and an old checkpoint can become misleading. Version the corpus, monitor answer quality, maintain a rollback path, and schedule re-evaluation after major legal or model changes. If you are building an open-source legal AI project, review India-focused open-source AI developer projects for practical collaboration patterns.

    Practical checklist

    Before calling the model ready, confirm that you have:

    • A narrow task definition and documented annotation guide
    • Traceable, permitted sources with provenance metadata
    • Deduplicated, quality-checked, privacy-reviewed text
    • Document-level train, validation, and test splits
    • A model and licence suitable for your intended use
    • Reproducible training records and versioned artefacts
    • Task-specific metrics plus expert error analysis
    • Retrieval or citation grounding for factual legal answers
    • Human review, monitoring, and an update policy

    Fine-tuning is valuable when it teaches a model a stable pattern—such as classification labels, extraction formats, or preferred summaries. It is not a substitute for current legal retrieval, authoritative sources, or professional judgement. In Indian legal technology, data governance and evaluation quality matter as much as model architecture.

    Last updated 23 September 2026

AIGI may be inaccurate. Replies seeded from the guide above.