Fine-tuning a language model on Indian compliance documents can improve classification, information extraction, clause retrieval, and drafting workflows—but only when the dataset, task definition, and evaluation process are designed carefully. This guide explains how to fine tune a model using Indian compliance documents on Hugging Face, with a practical workflow for teams handling GST, Companies Act, SEBI, RBI, labour, procurement, privacy, and sector-specific records.
Fine-tuning does not turn a model into a lawyer or make its outputs legally authoritative. It teaches a model to perform a defined task on a defined data distribution. For many compliance applications, retrieval-augmented generation (RAG), structured extraction, or a hybrid approach is safer than training a model to memorise changing rules. Review the broader design choices in best practices for fine-tuning LLMs on custom data before committing compute and annotation effort.
Define the compliance task first
Start with one measurable use case rather than “understand Indian law.” Common fine-tuning targets include:
- Document classification: identify filing type, applicable regulator, risk category, or review queue.
- Named-entity extraction: capture CIN, GSTIN, PAN, dates, monetary values, sections, forms, and filing deadlines.
- Clause classification: label indemnity, audit, termination, data-processing, or disclosure provisions.
- Question answering: answer questions from supplied passages, with citations and an “insufficient evidence” option.
- Sequence-to-sequence drafting: generate a structured checklist or review summary from a document.
For a first production pilot, classification or extraction is usually easier to validate than open-ended legal drafting. Define the input, output schema, acceptable errors, escalation rule, and human review point before collecting data.
Build a legally usable dataset
Indian compliance documents are often PDFs, scanned notices, spreadsheets, circulars, contracts, and filings. Preserve the source and provenance of every record. Your dataset should track:
- Source URL or repository, issuing authority, publication date, and document version.
- Language and script, including English, Hindi, and relevant regional-language content.
- Document type, jurisdiction, regulator, sector, and applicable period.
- OCR status, extraction confidence, and page or paragraph references.
- Copyright, licence, access restrictions, personal-data exposure, and retention policy.
Do not casually upload confidential filings, employee records, customer identifiers, Aadhaar numbers, financial account details, or privileged legal advice to a hosted service. Apply access controls, redaction, encryption, and a documented deletion process. India’s privacy and sectoral obligations may apply differently depending on the organisation and data category; obtain legal and security review before using private material.
PDF extraction is a major failure point. Keep headings, tables, numbered clauses, footnotes, annexures, and page references where possible. OCR errors such as “1” versus “I”, broken rupee amounts, or missing section symbols can change the meaning of a rule. Sample documents manually after extraction and reject low-confidence pages instead of silently training on them.
Prepare examples for Hugging Face
Use a consistent JSONL format. For classification, each row might contain text, label, document_id, and source_page. For extraction, use a schema such as:
{"text":"The entity shall file Form ... by 30 September.","entities":[{"text":"30 September","label":"DEADLINE"}]}For instruction tuning, include the task instruction, relevant passage, expected answer, and citation fields. Keep answers grounded in the provided text and include examples where the correct response is not enough information. This discourages confident fabrication.
Split data by document, issuer, and time period—not by random paragraph alone. Random splitting can place nearly identical clauses from one circular in both training and test sets, producing an inflated score. Hold out newer notifications or entire organisations to test whether the model generalises beyond memorised templates. Maintain a small, carefully reviewed “gold” set for final comparisons.
Select a base model and training method
Use a model that supports your languages, context length, licence, latency target, and hardware budget. English-only BERT models may perform poorly on Hindi-English code-mixed text or Indian names and abbreviations. Compare multilingual and Indic-capable models on a small benchmark rather than selecting by popularity.
For classification, start with AutoModelForSequenceClassification. For token extraction, use AutoModelForTokenClassification. For generative tasks, use a causal or sequence-to-sequence model with supervised examples. Parameter-efficient fine-tuning methods such as LoRA or QLoRA can reduce GPU memory requirements and make it easier to maintain separate adapters for different compliance domains. Quantisation may reduce cost, but test whether it harms dates, numbers, legal references, or multilingual accuracy.
Install a pinned environment rather than relying on an untracked latest version:
pip install "transformers" "datasets" "evaluate" "accelerate" "peft" "bitsandbytes"Load the tokenizer and use the model’s native padding and special-token configuration. Avoid blindly applying padding='max_length' to very long documents; dynamic padding and deliberate chunking are usually more efficient.
Tokenise and train with reproducibility
A simple classification pipeline looks like this:
from datasets import load_dataset
from transformers import AutoTokenizer, AutoModelForSequenceClassification
model_id = "your-approved-base-model"
tokenizer = AutoTokenizer.from_pretrained(model_id)
dataset = load_dataset("json", data_files={"train": "train.jsonl", "test": "test.jsonl"})
encoded = dataset.map(
lambda batch: tokenizer(batch["text"], truncation=True, max_length=1024),
batched=True
)
model = AutoModelForSequenceClassification.from_pretrained(
model_id, num_labels=3
)Use a validation set, fixed random seeds, versioned configuration, and checkpoint retention. Record the model revision, dataset hash, preprocessing code, GPU type, learning rate, batch size, epochs, and stopping criteria. Begin conservatively: low learning rates, few epochs, early stopping, and class weighting or balanced sampling where labels are uneven. Over-training is particularly risky when the corpus is small or repetitive.
Long compliance documents require a strategy: chunk by heading or clause, classify chunks, then aggregate results; use sliding windows where boundary context matters; or use a long-context model after benchmarking memory and latency. Never truncate a deadline, exception, or definition without measuring the effect.
Evaluate for legal risk, not just accuracy
Report precision, recall, and F1 by class, not only overall accuracy. For extraction, measure entity-level exact and partial matches. For generative outputs, check factuality, citation correctness, refusal behaviour, numerical accuracy, and format compliance. Include adversarial cases involving amendments, exceptions, contradictory dates, tables, scanned pages, and outdated rules.
Have compliance professionals review a stratified sample. Track false negatives separately: missing a high-risk filing or deadline may be more serious than sending an item for manual review. Add confidence thresholds and route uncertain results to a human. Compare the fine-tuned model against a strong baseline, a rules-based system, and RAG where appropriate. If the task is mainly “find the current applicable rule,” how to automate legal compliance with AI in India offers useful context on workflow design beyond model training.
Deploy with controls
Publish the model or adapter to a private Hugging Face repository only after security approval. Store dataset and model cards documenting intended use, limitations, languages, known failure modes, training data rights, and evaluation results. Use an internal inference service with authentication, logging, rate limits, and redaction of sensitive prompts. Do not expose raw compliance documents in logs.
Production systems should retrieve the current source text, display citations and page numbers, preserve an audit trail, and make the model’s version visible to reviewers. Add monitoring for data drift, new circulars, OCR changes, language shifts, and performance by regulator or document type. Retrain or refresh the retrieval index when rules change; do not assume fine-tuning alone keeps legal knowledge current.
A practical launch checklist
- Define one task and a measurable error budget.
- Obtain document rights and complete privacy/security review.
- Preserve provenance, versions, page references, and OCR quality signals.
- Create document-level splits and a professionally reviewed gold set.
- Benchmark multilingual models and a RAG or rules-based baseline.
- Train with PEFT where appropriate and log every experiment.
- Evaluate false negatives, citations, dates, numbers, and refusal behaviour.
- Deploy behind human review, audit logging, access controls, and monitoring.
Fine-tuning is worthwhile when you have repeated, labelled examples and a stable task. If your main problem is changing legislation or source-grounded answers, invest first in document retrieval, citation handling, and review workflows. For teams building broader custom-data systems, Indian open-source AI developer projects can also provide practical implementation patterns and reusable tooling.