Small language models can be useful for Indian public-sector workflows without requiring the budget or infrastructure of a large generative model. A classifier can route citizen grievances, a question-answering system can search scheme documents, and a multilingual model can identify topics across Hindi, English, and regional-language submissions. The quality of these systems depends less on model size than on well-defined tasks, clean data, and disciplined evaluation.
This guide explains how to fine-tune a small model on Hugging Face using Indian government data. It focuses on a practical supervised-learning workflow that a builder can run on a capable local GPU, a cloud notebook, or a modest training instance.
Start with the task, not the model
Government datasets are often published as tables, reports, PDFs, dashboards, or APIs. They are not automatically suitable for fine-tuning. First define the output your system must produce:
- Text classification: assign a department, scheme, grievance category, or urgency label.
- Named-entity recognition: extract districts, ministries, dates, crop names, or scheme identifiers.
- Question answering: identify an answer span in an official document.
- Text generation: produce a controlled summary or draft response from source material.
- Semantic search: retrieve the most relevant circular, notification, or scheme rule.
For many public-sector use cases, retrieval and classification are safer starting points than free-form generation. If you need a multilingual system, compare your approach with open-source small language models for Hindi and language-specific models before defaulting to an English checkpoint.
Find and assess Indian government data
The Open Government Data platform is a useful starting point for datasets from ministries, departments, and state agencies. Also check the publishing department, official APIs, annual reports, and terms attached to each download. Record the following before using any file:
- Publisher, dataset title, download date, and source URL
- Licence or reuse terms, including attribution requirements
- Data coverage by state, district, language, and time period
- Whether the fields contain personal, confidential, or sensitive information
- Version changes, missing values, and known limitations
Public availability does not remove privacy obligations. Exclude names, phone numbers, addresses, identification numbers, and free-text details that can identify an individual unless you have a clear lawful basis and a robust de-identification process. Keep a source manifest so every training example can be traced back to its origin.
Prepare a high-quality training set
The original CSV or PDF is rarely ready for a transformer. Convert it into examples with a stable schema. For classification, a simple structure might contain text and label. For question answering, use question, context, and answers. For instruction-style generation, store an input, an expected output, and the source reference.
Clean the data carefully:
- Normalise broken Unicode, whitespace, encoding, and OCR errors.
- Preserve meaningful Indian names, place names, abbreviations, and code-mixed text.
- Remove duplicate rows and near-duplicate documents.
- Standardise labels and document every relabelling decision.
- Keep tables and dates in a consistent representation.
- Flag machine-translated or OCR-derived records for separate quality review.
Avoid blindly removing punctuation or stop words. Legal and administrative language often depends on numbers, dates, section references, and punctuation. For multilingual work, test tokenisation on Devanagari and other scripts rather than assuming an English tokenizer will be efficient. If your project involves regional-language adaptation, see the workflow for fine-tuning Llama for Indian regional languages.
Split the dataset by document, department, or time period—not only by random row. Randomly splitting pages from the same circular can leak nearly identical text into evaluation. A useful starting point is 80% training, 10% validation, and 10% test data, but a smaller dataset may need cross-validation or repeated runs. Reserve a genuinely difficult, recent, or out-of-domain test set.
Choose a compact Hugging Face model
For classification and extractive question answering, encoder models such as DistilBERT, a multilingual BERT variant, or an Indic-focused encoder are often more practical than a generative model. For summarisation or controlled generation, use a small instruction-tuned model that supports your scripts and licence requirements.
Compare checkpoints on:
- Script and language coverage
- Maximum context length
- Tokenisation efficiency for Indian languages
- Licence and redistribution terms
- CPU, GPU, and memory requirements
- Existing benchmark performance on similar tasks
Model selection should follow a baseline. A keyword search, TF-IDF classifier, or zero-shot model gives you a reference point and can reveal whether fine-tuning adds real value. For broader optimisation guidance, use these best practices for fine-tuning LLMs on custom data.
Install the training stack
Use a virtual environment and pin versions for reproducibility:
python -m venv .venv
source .venv/bin/activate
pip install -U torch transformers datasets evaluate accelerate scikit-learn sentencepieceLoad a local CSV with the Datasets library:
from datasets import load_dataset
data = load_dataset("csv", data_files={
"train": "train.csv",
"validation": "validation.csv",
"test": "test.csv",
})A classification setup using a compact checkpoint can look like this:
from transformers import AutoTokenizer, AutoModelForSequenceClassification
checkpoint = "distilbert-base-multilingual-cased"
tokenizer = AutoTokenizer.from_pretrained(checkpoint)
model = AutoModelForSequenceClassification.from_pretrained(
checkpoint,
num_labels=number_of_labels
)
def tokenize(batch):
return tokenizer(batch["text"], truncation=True, max_length=256)
tokenized = data.map(tokenize, batched=True)Use dynamic padding through a data collator when possible. Set max_length after inspecting actual document lengths; an unnecessarily large context increases cost without improving results.
Train with a reproducible baseline
The Hugging Face Trainer API is sufficient for many small supervised tasks:
from transformers import TrainingArguments, Trainer
args = TrainingArguments(
output_dir="./model-output",
eval_strategy="epoch",
save_strategy="epoch",
learning_rate=2e-5,
per_device_train_batch_size=16,
per_device_eval_batch_size=32,
num_train_epochs=3,
weight_decay=0.01,
load_best_model_at_end=True,
metric_for_best_model="f1",
report_to="none",
)
trainer = Trainer(
model=model,
args=args,
train_dataset=tokenized["train"],
eval_dataset=tokenized["validation"],
tokenizer=tokenizer,
)
trainer.train()For limited hardware, reduce batch size and use gradient accumulation, mixed precision, or parameter-efficient methods such as LoRA. Start with a small learning-rate sweep and three fixed random seeds. Track training time, peak memory, model size, and validation metrics—not only loss.
Evaluate for public-sector failure modes
Accuracy alone can hide serious weaknesses. Report precision, recall, and F1 by class, especially for minority districts, languages, departments, and high-risk categories. Inspect a confusion matrix and manually review false positives and false negatives.
Test for:
- Hindi-English code mixing and transliterated text
- Spelling variation in place and person names
- OCR noise and scanned documents
- New policy terminology and unseen districts
- Imbalanced labels and missing categories
- Memorisation of personal or sensitive information
For generated answers, require citations to source documents, measure factual consistency, and include an abstain option when evidence is missing. A human reviewer should remain in the loop for benefits, eligibility, medical, legal, or enforcement-related decisions.
Package and deploy responsibly
Save the model, tokenizer, label mapping, preprocessing code, dataset manifest, and evaluation report together:
trainer.save_model("./model-output/best")
tokenizer.save_pretrained("./model-output/best")Publish to the Hugging Face Hub only after checking data rights, model-card requirements, and whether the training set could expose personal information. For an internal service, wrap inference in an API with authentication, rate limits, logging, and versioned releases. Quantisation or distillation can reduce latency for edge deployments; see AI model optimisation for mobile devices when the target is a field tablet or low-power device.
A practical launch checklist
- Define one measurable task and a safe fallback.
- Confirm dataset provenance, licence, and privacy status.
- Build a clean, deduplicated, multilingual-aware split.
- Establish a non-neural baseline before fine-tuning.
- Evaluate by language, geography, department, and time period.
- Record model, data, hyperparameters, and known limitations.
- Pilot with domain experts before connecting to citizen-facing workflows.
- Monitor drift as schemes, forms, and administrative terminology change.
Fine-tuning a small model is not simply a matter of downloading a government CSV and running trainer.train(). The durable advantage comes from traceable data, language-appropriate modelling, transparent evaluation, and deployment controls. That combination can produce affordable AI systems that are useful in Indian administrative contexts without pretending that a compact model is infallible.