PEFT (Parameter-Efficient Fine-Tuning) lets you adapt a Hugging Face model by training a small set of additional parameters instead of updating the entire network. For Indian-language and India-specific applications, this can reduce GPU cost while preserving the capabilities of a strong base model.
This guide shows a practical workflow for classification and instruction-tuning projects, including dataset selection, LoRA configuration, multilingual evaluation, and common failure modes. It complements broader best practices for fine-tuning LLMs on custom data, but focuses on implementing PEFT with Indian datasets.
Choose the right PEFT approach
For most builders, LoRA (Low-Rank Adaptation) is the best starting point. It freezes the base model and inserts trainable low-rank matrices into selected layers. QLoRA combines LoRA with 4-bit quantisation, making larger causal language models practical on a single capable GPU.
Use the method that matches the task:
- LoRA: Good default for text classification, extraction, and instruction tuning.
- QLoRA: Useful when GPU memory is limited and the base model is too large for full-precision loading.
- IA3 or prompt tuning: Worth testing for smaller, narrowly defined tasks, but less commonly used in production workflows.
- Full fine-tuning: Consider only when you have substantial data, sufficient compute, and a clear reason adapters cannot meet the target quality.
Select a base model that already supports the languages and script used by your application. IndicBERT, MuRIL, multilingual BERT variants, and multilingual causal language models can be sensible starting points, but benchmark candidates on a representative validation set before training.
Prepare an Indian dataset carefully
The dataset usually matters more than the adapter configuration. Sources may include public data.gov.in records, licensed enterprise data, Hugging Face datasets, customer-support conversations, government documents, or carefully reviewed community contributions. Confirm the licence, consent basis, redistribution rights, and restrictions on personal information before uploading data to a hosted service.
Create a consistent schema. For supervised classification, use fields such as text and label. For instruction tuning, use messages with role and content values. Keep the original source, language, script, date, and annotator information in metadata rather than mixing them into the training prompt.
Pay particular attention to Indian-language issues:
- Normalise Unicode without destroying meaningful script distinctions.
- Decide how to handle transliterated text such as Hinglish or Romanised Tamil.
- Preserve code-mixed examples if they reflect real users.
- Standardise punctuation, whitespace, numerals, and common spelling variants.
- Remove phone numbers, Aadhaar details, addresses, account numbers, and other sensitive data.
- Split duplicates and near-duplicates before creating train, validation, and test sets.
- Stratify by language, label, region, and domain where possible.
Do not randomly split near-identical documents. A source-level or time-based split gives a more honest estimate of production performance. Keep a small, manually reviewed challenge set for dialects, code-switching, spelling variation, and low-resource languages.
Install the training stack
Use a recent Python environment and pin versions for reproducibility. The core packages are:
pip install -U transformers datasets peft accelerate evaluate bitsandbytes sentencepieceFor a new project, record the model revision, dataset version, random seed, GPU type, package versions, and training configuration. This makes it possible to reproduce results and investigate regressions.
Fine-tune a classifier with LoRA
The following example assumes a dataset with text and integer label columns. Replace the model identifier and dataset path with models and data appropriate to your language coverage.
from datasets import load_dataset
from transformers import (
AutoTokenizer, AutoModelForSequenceClassification,
DataCollatorWithPadding, TrainingArguments, Trainer
)
from peft import LoraConfig, TaskType, get_peft_model
model_id = "google/muril-base-cased"
dataset = load_dataset("your-org/indian-text-classification")
tokenizer = AutoTokenizer.from_pretrained(model_id)
def tokenize(batch):
return tokenizer(batch["text"], truncation=True, max_length=256)
encoded = dataset.map(tokenize, batched=True)
model = AutoModelForSequenceClassification.from_pretrained(
model_id, num_labels=3
)
lora_config = LoraConfig(
task_type=TaskType.SEQ_CLS,
r=16,
lora_alpha=32,
lora_dropout=0.1,
target_modules=["query", "value"],
)
model = get_peft_model(model, lora_config)
model.print_trainable_parameters()
args = TrainingArguments(
output_dir="./indian-peft-classifier",
eval_strategy="epoch",
save_strategy="epoch",
learning_rate=2e-4,
per_device_train_batch_size=16,
per_device_eval_batch_size=32,
num_train_epochs=3,
weight_decay=0.01,
load_best_model_at_end=True,
metric_for_best_model="f1",
report_to="none",
)
trainer = Trainer(
model=model,
args=args,
train_dataset=encoded["train"],
eval_dataset=encoded["validation"],
tokenizer=tokenizer,
data_collator=DataCollatorWithPadding(tokenizer),
)
trainer.train()
trainer.save_model("./indian-peft-classifier")target_modules differ across architectures. Check the model implementation rather than copying a configuration blindly. For causal language models, common targets include attention projection layers such as q_proj and v_proj; for encoder models, names may be query and value.
Train an instruction-following adapter
For chat or generation tasks, format every example using the model's chat template. Keep prompts and responses distinct, and calculate loss on the assistant response where the training framework supports response-only masking. This prevents the adapter from learning to reproduce instructions instead of answering them.
Start with conservative settings: a maximum sequence length that covers real examples, a small learning rate, gradient accumulation, and early stopping. QLoRA can reduce memory requirements, but quantisation does not fix poor data or inadequate evaluation. If using bitsandbytes, verify that your CUDA, GPU, and package versions are compatible before launching a long run.
Adapters are small and easy to version. Store the adapter, tokenizer, configuration, dataset commit, and evaluation report separately from the base model. You can merge an adapter into the base model for a simpler serving path, but retain the unmerged adapter so you can update or roll back it independently.
Evaluate beyond one accuracy score
Indian-language systems can appear strong overall while failing badly on a smaller language or region. Report results by:
- Language and script, including code-mixed inputs.
- Label, intent, or document type.
- Geography or customer segment when relevant and lawful.
- Short versus long inputs.
- Clean, noisy, and transliterated text.
- Abstention, refusal, and harmful-content cases.
Use macro-F1 for imbalanced classification, precision and recall for high-cost errors, and exact match or task-specific measures for extraction. For generation, combine automated checks with blind review by native speakers and domain specialists. Measure latency, memory use, and cost on the hardware you will actually deploy.
Build a baseline before PEFT. Compare the untuned model, a simple keyword or retrieval baseline, and the adapted model on the same held-out data. If the adapter improves one language while degrading another, adjust sampling, add targeted examples, or consider language-specific adapters instead of increasing training time.
Common problems and fixes
- Training loss falls but quality does not: inspect labels, duplicates, leakage, and prompt formatting.
- One language dominates: use balanced sampling or language-aware batches.
- Overfitting appears quickly: reduce epochs, lower the adapter rank, add dropout, or expand the dataset.
- Outputs are truncated: increase
max_lengthonly after measuring GPU memory and sequence distribution. - Model memorises personal data: remove sensitive records, deduplicate aggressively, and test with canary strings.
- Results vary between runs: set seeds, pin versions, and save the exact dataset revision.
For applications such as education or public services, involve native-language reviewers early. The same evaluation discipline is useful when building open-source vision-language models for Indian languages or other multimodal systems.
Deploy responsibly
Before serving the adapter, document intended use, unsupported languages, known failure modes, and escalation paths. Add input validation, logging that excludes sensitive content, rate limits, and human review for high-impact decisions. Do not treat a high benchmark score as evidence that the model is safe for medical, legal, credit, employment, or welfare decisions.
For production pilots, monitor quality by language and task, not only aggregate traffic. Track drift as vocabulary, schemes, policies, or user behaviour change. A small, regularly reviewed evaluation suite is more valuable than a one-time leaderboard result. Builders working on local-language products may also find the architecture lessons in Indian open-source AI developer projects useful.
FAQ
Does PEFT require a powerful GPU?
Not always. LoRA on encoder models can run on modest hardware, while QLoRA can make larger causal models feasible. Memory still depends on model size, sequence length, batch size, and quantisation support.
How much Indian-language data is needed?
There is no universal number. A few thousand clean, representative examples can beat a much larger noisy set for a narrow task. Test several data sizes to identify whether quality is still improving.
Should I translate everything into English first?
Usually no. Translation can remove cultural, linguistic, and code-mixing signals. Fine-tune and evaluate in the languages users will actually use; compare translation only as a baseline.
Can I publish the adapter?
Only if the base-model licence, dataset terms, privacy obligations, and any third-party content permissions allow it. Publish a model card with languages, data sources, limitations, metrics, and intended use.
For India-focused AI teams, careful data governance and measurable multilingual quality are as important as the training command. Start with a narrow task, establish a baseline, run a reproducible PEFT experiment, and expand only when the evidence supports it.