Start with the task, not the model
Fine-tuning a model using Indian court judgment summaries on Hugging Face works best when the task is narrowly defined. “Understand legal text” is too broad to produce a reliable training target. Choose one measurable outcome first:
- Topic classification: assign a judgment to areas such as criminal, constitutional, tax, labour, or commercial law.
- Outcome or issue classification: predict a defined label, such as whether a petition was allowed, dismissed, or partly allowed—only where the labels are consistently represented.
- Extractive question answering: identify spans containing provisions, parties, dates, or cited authorities.
- Summarisation: generate a short summary from a judgment, preferably with human-written reference summaries.
- Semantic search or retrieval: create embeddings that help users find similar cases, rather than asking a classifier to make legal conclusions.
A classification model is usually the simplest first project. For generation, use the workflow in best practices for fine-tuning LLMs on custom data, especially its guidance on prompt formats, validation, and overfitting.
Build a legally defensible dataset
Indian judgments are not interchangeable documents. A Supreme Court decision, a High Court order, a tribunal ruling, and a case summary may differ sharply in length, language, citation style, and legal authority. Record these fields before training:
case_idand a stable source URL- court, bench, jurisdiction, date, and case type
- original text and summary text as separate fields
- language and whether the document is translated
- task label, label source, and annotator notes
- licence, access conditions, and redaction status
Do not assume that material available online is automatically free to reuse. Check the source’s terms, remove personal information where appropriate, and document provenance. Court records can contain sensitive details about survivors, children, medical histories, addresses, and financial information. A dataset card should explain collection, licensing, redaction, known gaps, and intended use.
Avoid leakage at the case level. If several documents belong to the same matter, keep them in the same split. Otherwise, a near-duplicate can appear in both training and test data and produce an inflated score. A sensible starting split is 80% training, 10% validation, and 10% test, with stratification for labels and separation by case, court, or time where the deployment setting requires it.
Prepare summaries and labels
A summary is useful only if its target is clear. Decide whether it captures facts, legal issues, reasoning, the disposition, or all four. Do not combine summaries written under different guidelines without tracking the difference. For generation tasks, inspect examples for hallucinated citations, omitted exceptions, and conclusions that are stronger than the judgment.
For classification, define labels in a short annotation guide. Include positive and negative examples, rules for ambiguous cases, and a process for adjudicating disagreements. Measure agreement between annotators on a sample. If a label cannot be assigned consistently by trained reviewers, it is unlikely to be a good supervised target.
Normalise text carefully. Preserve section boundaries, paragraph order, citations, and statutory references where they carry meaning. Remove website navigation, repeated headers, OCR artefacts, and boilerplate—but keep a copy of the raw source for auditability. For multilingual datasets, retain the original language and record whether translation was human, machine-generated, or unavailable. Models suited to Indian languages may be more appropriate than an English-only checkpoint; see the landscape of open-source vision-language models for Indian languages if your corpus includes scanned pages or mixed-script documents.
Load and tokenise data with Hugging Face
Install a current environment and pin versions for reproducibility:
pip install -U transformers datasets evaluate accelerate scikit-learnCreate a CSV or JSONL file with a text column and, for classification, a numeric label column. Then load and split the data:
from datasets import load_dataset
raw = load_dataset("json", data_files={"data": "judgments.jsonl"})["data"]
splits = raw.train_test_split(test_size=0.2, seed=42)
valid_test = splits["test"].train_test_split(test_size=0.5, seed=42)
dataset = {
"train": splits["train"],
"validation": valid_test["train"],
"test": valid_test["test"],
}For a classification baseline, select a multilingual or Indian-language-compatible encoder after checking its licence, context length, tokenizer coverage, and benchmark behaviour. distilbert-base-uncased is easy to run but is not automatically a good fit for Indian legal text, regional languages, or long judgments.
from transformers import AutoTokenizer
checkpoint = "xlm-roberta-base"
tokenizer = AutoTokenizer.from_pretrained(checkpoint)
def tokenize(batch):
return tokenizer(
batch["text"],
truncation=True,
max_length=512,
)
encoded = {name: split.map(tokenize, batched=True) for name, split in dataset.items()}Long judgments often exceed the model’s context window. Truncating the beginning may remove the order, while truncating the end may remove the decision. Test alternatives: structured extraction of relevant sections, sliding windows with aggregation, hierarchical models, or retrieval followed by short-context classification. Do not silently treat a truncated document as a complete judgment.
Fine-tune a classification baseline
Use AutoModelForSequenceClassification and an explicit label map. The following example assumes labels are already encoded as integers:
import numpy as np
import evaluate
from transformers import (
AutoModelForSequenceClassification,
TrainingArguments,
Trainer,
)
labels = ["criminal", "constitutional", "commercial"]
model = AutoModelForSequenceClassification.from_pretrained(
checkpoint,
num_labels=len(labels),
id2label=dict(enumerate(labels)),
label2id={label: i for i, label in enumerate(labels)},
)
f1 = evaluate.load("f1")
def compute_metrics(eval_pred):
logits, y_true = eval_pred
y_pred = np.argmax(logits, axis=-1)
return f1.compute(predictions=y_pred, references=y_true, average="macro")
args = TrainingArguments(
output_dir="legal-summary-classifier",
eval_strategy="epoch",
save_strategy="epoch",
learning_rate=2e-5,
per_device_train_batch_size=8,
per_device_eval_batch_size=8,
num_train_epochs=3,
weight_decay=0.01,
load_best_model_at_end=True,
metric_for_best_model="f1",
report_to="none",
)
trainer = Trainer(
model=model,
args=args,
train_dataset=encoded["train"],
eval_dataset=encoded["validation"],
processing_class=tokenizer,
compute_metrics=compute_metrics,
)
trainer.train()Batch size, learning rate, sequence length, and epochs depend on corpus size and hardware. Use gradient accumulation or parameter-efficient methods such as LoRA when GPU memory is limited. The Indian open-source AI developer projects guide can help you identify local tooling and community patterns, but always verify licences and maintenance status before adopting a dependency.
Evaluate for legal use, not just leaderboard scores
Accuracy can conceal serious failures, particularly when labels are imbalanced. Report macro-F1, per-class precision and recall, a confusion matrix, and results by court, language, year, and document length. Compare against a majority-class baseline and a simple keyword or TF-IDF baseline. For summarisation, assess factual consistency, citation preservation, omission of the disposition, and expert-rated usefulness—not only ROUGE.
Run qualitative error analysis on false positives and false negatives. Ask whether the model is learning legal substance or shortcuts such as court names, recurring party names, template language, or publication source. Keep the test set locked until development is complete. For a production search experience, evaluate retrieval separately using relevant-case recall and ranking quality.
Publish and deploy responsibly
Push the model, tokenizer, dataset card, training configuration, evaluation results, and limitations to a private or public Hugging Face repository according to your data rights. Never upload restricted judgments, unredacted personal data, API keys, or documents whose terms prohibit redistribution.
A fine-tuned model is an analysis component, not a lawyer or adjudicator. Do not present predictions as legal advice, a binding interpretation, or a substitute for counsel. Add confidence thresholds, abstention for unfamiliar inputs, citations back to source documents, access controls, audit logs, and human review for consequential decisions. Monitor performance after deployment because court composition, drafting patterns, OCR quality, and label definitions can change.
For teams planning a broader AI product, pair this workflow with a clear AI framework for Indian student entrepreneurs or internal model-risk checklist. The practical goal is not to claim that a model “understands Indian law”; it is to build a bounded, testable system whose evidence, uncertainty, and failure modes remain visible.
Frequently asked questions
Can I fine-tune on summaries alone? Yes, for tasks whose input is a summary. If the deployed system receives full judgments, train and test on inputs that match that setting. Summary-only training will not prove performance on raw judgments.
Should I use a generative LLM for classification? Not necessarily. An encoder classifier is often cheaper, faster, and easier to evaluate. Use a generative model when the required output is a summary, explanation, extraction, or structured response.
How much data is enough? There is no universal threshold. Start with a carefully labelled pilot, plot validation performance against dataset size, and inspect errors. A smaller, consistent dataset can outperform a larger noisy collection.
Can this system decide whether a case will win? It should not be marketed that way. Historical judgments reflect procedural and social biases, and predictive outputs can mislead users. Limit the system to transparent research, retrieval, extraction, or classification tasks with human oversight.