Fine-tuning can make a general-purpose language model more useful for an Indian product—but only when the data, task, and evaluation plan are carefully matched. A Hindi customer-support assistant, a multilingual government-document search tool, and a Tamil speech-to-text pipeline need different models, datasets, and training objectives.
This guide explains how to fine-tune an LLM on Hugging Face using Indian public datasets in a way that is practical for founders, researchers, and student builders in 2026. It focuses on supervised fine-tuning and parameter-efficient methods such as LoRA and QLoRA, which reduce GPU memory requirements.
Start with the task, not the dataset
Define the behaviour you want before downloading data. Fine-tuning is appropriate when a model needs to learn a repeatable format, domain vocabulary, response style, or task-specific mapping. It is usually not the best first solution for adding frequently changing facts; use retrieval-augmented generation for that.
Common India-focused use cases include:
- Question answering over public schemes, policies, or local services.
- Classification of support tickets, reviews, or citizen feedback.
- Summarisation of Indian-language news and government documents.
- Translation and transliteration between English and Indian languages.
- Instruction following in Hindi, Tamil, Telugu, Bengali, Marathi, Kannada, Malayalam, Gujarati, Punjabi, or other languages.
If you are building a production assistant, first compare fine-tuning with retrieval, prompt engineering, and a smaller specialist model. The best practices for fine-tuning LLMs on custom data are especially useful when deciding whether training will solve the actual product problem.
Choose an Indian dataset carefully
Public does not automatically mean clean, legal, or suitable for training. Potential sources include:
- AI4Bharat and Indic-language resources: Useful for translation, language identification, speech, and text tasks. Check the specific dataset card, language coverage, and licence.
- Bhashini ecosystem resources: Helpful for Indian-language applications, but availability, access terms, and task formats vary by resource.
- Indian-language Wikipedia dumps: Suitable for broad factual or language modelling experiments after substantial cleaning.
- Government open data: Data.gov.in and ministry portals can provide domain text, but tables and PDFs often require extraction, normalisation, and manual review.
- Common Crawl or web corpora: Potentially broad, but noisy and difficult to filter for language, location, duplication, personal data, and copyright.
- Open parallel corpora: Useful for translation, provided sentence alignment and licensing are verified.
Prefer a smaller, well-labelled dataset over a large, noisy scrape. Record the source URL, collection date, language, licence, preprocessing steps, and intended use. Remove personal information, credentials, phone numbers, email addresses, and unnecessary identifiers. Do not train on scraped content merely because it is accessible online.
For language coverage, inspect the actual script and language distribution rather than trusting a dataset label. Hindi written in Devanagari, Hinglish in Latin script, and code-mixed Hindi-English are different training signals. Keep them as separate subsets when the product requires controlled behaviour.
Prepare and split the data
Convert the source into a consistent JSONL format. For instruction tuning, each record can contain instruction, input, and output; for classification, use text and label; for translation, use source and target.
{"instruction":"Summarise this notice in simple Hindi","input":"...","output":"..."}Before training:
- Normalise Unicode without destroying meaningful Indic characters.
- Remove HTML, boilerplate, broken OCR, and duplicated records.
- Preserve punctuation, numerals, named entities, and code-mixed text when relevant.
- Filter examples with empty, contradictory, or unsafe outputs.
- Deduplicate near-identical documents to prevent memorisation.
- Create train, validation, and test splits by document or source—not random lines from the same document.
A practical starting point is 80% training, 10% validation, and 10% testing. Keep a held-out test set that represents real users: dialect variation, spelling errors, mixed scripts, short prompts, and domain-specific terms. Never use test examples during prompt formatting or iterative training.
Select a model and training method
Choose a model that supports the languages and context length your application needs. Inspect its model card, licence, tokenizer behaviour, known limitations, and commercial-use conditions. A multilingual base model may be preferable to an English-first model, while an Indic specialist model may perform better on a narrower language task.
For most teams, LoRA or QLoRA is the sensible first experiment. These methods train adapter weights rather than updating every parameter, lowering memory use and making it easier to compare multiple domain or language adapters. Full fine-tuning is more expensive and should be justified by scale, data volume, and measured gains.
Install the core stack:
pip install -U transformers datasets accelerate peft trl bitsandbytes evaluateAuthenticate with Hugging Face only through environment variables or a secret manager. Avoid placing tokens in notebooks or source repositories.
Load and format the dataset
from datasets import load_dataset
raw = load_dataset("your-org/your-indian-dataset")
def format_example(row):
return {
"text": (
"### Instruction:\n" + row["instruction"] +
"\n### Input:\n" + row.get("input", "") +
"\n### Response:\n" + row["output"]
)
}
train = raw["train"].map(format_example)
valid = raw["validation"].map(format_example)Use the model’s recommended chat template when training a chat model. Do not invent a format that differs from the model’s inference template unless you have tested the consequences. Tokenise with truncation, set a sensible maximum sequence length, and measure how many examples are being cut off.
Fine-tune with LoRA
A minimal supervised fine-tuning workflow using TRL looks like this:
from transformers import AutoTokenizer, AutoModelForCausalLM
from peft import LoraConfig
from trl import SFTConfig, SFTTrainer
model_id = "your-base-model"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
model_id,
device_map="auto",
torch_dtype="auto"
)
lora = LoraConfig(
r=16,
lora_alpha=32,
lora_dropout=0.05,
target_modules=["q_proj", "k_proj", "v_proj", "o_proj"],
task_type="CAUSAL_LM"
)
args = SFTConfig(
output_dir="./indic-llm-adapter",
num_train_epochs=2,
per_device_train_batch_size=2,
gradient_accumulation_steps=8,
learning_rate=2e-4,
eval_strategy="steps",
eval_steps=200,
logging_steps=20,
save_steps=200,
max_seq_length=1024,
report_to="none"
)
trainer = SFTTrainer(
model=model,
tokenizer=tokenizer,
train_dataset=train,
eval_dataset=valid,
peft_config=lora,
args=args,
dataset_text_field="text"
)
trainer.train()
trainer.save_model("./indic-llm-adapter")The exact argument names can change across transformers and TRL releases, so pin tested versions in requirements.txt. Start with a short pilot run. Compare the tuned model with the base model using the same prompts, decoding settings, and test set before increasing epochs or rank.
Evaluate Indian-language performance
Loss alone is not enough. Track task metrics appropriate to the use case:
- Classification: macro-F1, per-language F1, precision, and recall.
- Translation: COMET or chrF alongside human review; BLEU alone can miss acceptable variants.
- Summarisation: factuality, coverage, readability, and human preference.
- Generation: exact-format accuracy, groundedness, toxicity, refusal quality, and hallucination rate.
- Multilingual chat: language adherence, script accuracy, code-switching behaviour, and response consistency.
Build a review set with native or highly proficient speakers. Evaluate across major languages separately instead of reporting one pooled score that hides weak performance. Test names, addresses, dates, rupee amounts, caste and community references, regional idioms, and low-resource language prompts carefully.
Check for memorisation by searching distinctive training phrases in model outputs. Also test whether the model reproduces personal data or copyrighted passages. If performance improves only on training-like prompts, the dataset may be too narrow or the model may be overfitting.
Deploy responsibly
Publish the adapter or model with a clear model card covering the base model, datasets, licences, languages, limitations, evaluation results, and safety considerations. If the dataset licence does not permit redistribution, publish training details without uploading the underlying data.
For inference, merge the adapter only when operationally useful; keeping adapters separate can simplify versioning and allow one base model to serve multiple domains. Quantisation may reduce serving costs, but validate quality after quantisation, especially for Indic scripts and longer contexts.
Before exposing the system to users, add input filtering, logging with privacy controls, rate limits, fallback responses, and human escalation. For applications such as education, finance, healthcare, or government services, provide source citations or retrieval evidence rather than presenting generated text as authoritative. Builders working on multilingual products can also review open-source vision-language models for Indian languages when their data includes scanned forms, images, or video.
A practical launch checklist
- Define one measurable task and one target user group.
- Verify dataset licences and remove sensitive information.
- Audit language, script, duplication, and label quality.
- Create document-level train, validation, and test splits.
- Establish a base-model benchmark before fine-tuning.
- Run a small LoRA experiment before full training.
- Evaluate each language and real-world failure mode separately.
- Document the model card, data provenance, and known limitations.
- Monitor drift and user feedback after deployment.
Fine-tuning is most valuable when it produces a measurable improvement over a strong baseline—not when it simply makes a model sound more local. Use Indian public data with discipline, keep the evaluation honest, and optimise for the exact users and languages your product serves.