LoRA is a practical way to adapt a multilingual language model without updating every parameter. For Indic-language applications, that matters: teams often need to improve performance for a specific language, script, domain, or interaction style while working with limited data and modest GPU budgets. A well-designed adapter can specialise a capable base model for Hindi customer support, Marathi document extraction, Tamil education content, Bengali search, or code-mixed Hinglish without maintaining a separate full model for every use case.
This guide focuses on supervised fine-tuning with LoRA or QLoRA for causal language models and encoder models. The same principles apply to classification, retrieval reranking, summarisation, translation, and conversational systems. For broader dataset and evaluation strategy, start with this builder’s guide to low-resource Indic NLP.
1. Define the adaptation target
Do not begin by choosing a LoRA rank. First specify what the adapter must improve:
- Language coverage: Hindi, Bengali, Tamil, Telugu, Marathi, Kannada, Gujarati, Malayalam, Punjabi, Odia, Assamese, Urdu, or a code-mixed variety.
- Script: Devanagari, Bengali-Assamese, Gurmukhi, Gujarati, Kannada, Malayalam, Odia, Tamil, Telugu, Perso-Arabic, or Romanised text.
- Task: generation, classification, extraction, translation, question answering, speech post-processing, or tool calling.
- Domain: healthcare, finance, agriculture, government services, education, or customer support.
- Quality target: factuality, terminology, instruction following, spelling, transliteration, latency, or cost.
An adapter should solve a measurable gap rather than merely make outputs sound more local. Create a fixed validation set before training, including native-language prompts, code-mixed examples, spelling variants, numerals, named entities, and realistic user errors.
2. Choose a suitable base model
Select a model whose tokenizer and pre-training data already provide reasonable coverage of the target language. A multilingual model with poor tokenisation can remain expensive and weak after adaptation because words are split into too many subword pieces. Compare token counts for representative samples in both native script and Romanised text before committing.
For generation, use a causal language model compatible with the transformers, peft, and bitsandbytes ecosystem. For classification or sequence labelling, an encoder such as XLM-R may be more appropriate. Check the model licence, commercial-use terms, supported context length, quantisation compatibility, and existing safety limitations.
A practical 2026 workflow is to test two or three candidate models on the same small benchmark, then train only the strongest one. Do not assume that the largest checkpoint will win on a narrow Indic-language task.
3. Build a clean, representative dataset
Data quality usually matters more than increasing the LoRA rank. Gather text from sources you are legally allowed to use, such as licensed documents, opt-in product logs, public-domain material, or carefully reviewed synthetic examples. Record provenance and consent where personal or sensitive data is involved.
For instruction tuning, store examples in a consistent structure:
{"messages":[
{"role":"user","content":"कृपया इस आवेदन का संक्षेप दें।"},
{"role":"assistant","content":"यह आवेदन ..."}
]}Prepare separate train, validation, and test splits. Avoid placing near-duplicates, translated copies, or documents from the same customer in different splits. Include:
- Native-script text and, if relevant, Romanised or code-mixed input.
- Dialect, register, and formal/informal variants.
- Domain terminology, abbreviations, dates, currency, addresses, and names.
- Negative examples where the correct response is to ask for clarification or refuse.
- Human-reviewed answers with a consistent style.
Normalise only what is safe to normalise. Unicode composition, punctuation, danda characters, zero-width marks, whitespace, and numeral formats can affect Indic-language meaning or tokenisation. Preserve an untouched copy of the source data and create a documented preprocessing version. For production systems, follow the data controls described in best practices for fine-tuning LLMs on custom data.
4. Configure LoRA or QLoRA
Install the core tooling:
pip install -U transformers datasets peft accelerate bitsandbytes trl evaluateA typical PEFT configuration looks like this:
from peft import LoraConfig, TaskType
lora_config = LoraConfig(
r=16,
lora_alpha=32,
lora_dropout=0.05,
target_modules=["q_proj", "k_proj", "v_proj", "o_proj"],
task_type=TaskType.CAUSAL_LM,
bias="none"
)Start with r=8 or r=16; increase to 32 only when validation results indicate that the task needs more capacity. lora_alpha is commonly set to roughly two times r, while dropout between 0.05 and 0.1 can help on small datasets. These are starting points, not universal settings.
Target module names vary by architecture. Inspect the model rather than copying a configuration blindly. Attention projections are a sensible baseline. For demanding domain adaptation, experiment with output projections and selected MLP layers, but track memory use and overfitting. QLoRA loads the base model in low-bit precision while training LoRA weights, reducing VRAM requirements; it does not remove the need for good data or careful evaluation.
5. Train with controlled experiments
Keep the base model frozen and train only the adapter parameters. Use gradient accumulation, mixed precision, checkpointing, and a learning-rate schedule appropriate to your hardware. A small supervised run might begin with a learning rate around 1e-4 for LoRA parameters, but sweep rather than treating this as a rule.
Track each experiment in a configuration file containing:
- Base-model revision and tokenizer version.
- Dataset commit, language mix, and number of examples.
- LoRA rank, alpha, dropout, and target modules.
- Sequence length, batch size, learning rate, and training steps.
- Quantisation settings and GPU type.
- Validation metrics and representative outputs.
Use early stopping or select the checkpoint with the best validation score. Too many epochs can cause memorisation, especially when the corpus is small or repetitive. For a multilingual adapter, test whether improving one language harms another; a separate adapter may be better than forcing all languages into one checkpoint.
6. Evaluate Indic-language quality properly
Accuracy alone rarely captures language quality. Build a test suite that combines automated and human review:
- Task metrics: F1, exact match, ROUGE, BLEU, chrF, or calibration, depending on the task.
- Language checks: script correctness, spelling, grammar, transliteration quality, and code-mixing behaviour.
- Robustness: noisy text, punctuation changes, dialect variation, short prompts, long context, and unseen names.
- Safety: privacy leakage, harmful instructions, medical or financial overconfidence, and incorrect translation of sensitive content.
- Operational metrics: latency, peak VRAM, tokens per second, adapter load time, and cost per request.
Use native speakers for a blind comparison between the base model and adapter. Ask reviewers to score correctness, fluency, cultural appropriateness, instruction adherence, and harmful omissions. Maintain a challenge set that is never used for training. This is especially important for public-facing products aimed at India’s diverse users, a concern also covered in building AI apps for the next billion users in India.
7. Package and deploy the adapter
Save the adapter separately from the base model so it can be versioned, shared, rolled back, or combined with other compatible adapters. Record the exact base-model revision and tokenizer. Before production, decide whether to load the adapter dynamically or merge it into the base weights. Dynamic loading is useful when serving several domains; merging can simplify inference but reduces flexibility.
Expose language and adapter selection explicitly in your serving layer. Log the model and adapter version for every request, redact personal data, and add fallback behaviour when the input language is uncertain. If the model powers a voice product, evaluate the full speech-to-text and text-to-speech pipeline rather than judging the text adapter alone; architecture choices from how to build a voice agent are relevant here.
Common failure modes
- Training on transliterated text only: the adapter performs poorly in the native script.
- Random data splits: validation scores look strong because duplicates leak across splits.
- Wrong target modules: the configuration silently misses the model’s actual projection names.
- Over-cleaning: punctuation, diacritics, zero-width characters, or code-mixing patterns are removed.
- Synthetic-data feedback loops: generated text amplifies grammar and factual errors.
- Single-language regression: gains in one language reduce performance elsewhere.
- No licence review: the dataset or base model cannot legally be used in the intended product.
A practical launch checklist
Before releasing an Indic-language adapter, confirm that you have:
- A documented use case, data licence, and privacy review.
- Native-speaker validation data and a held-out challenge set.
- Baselines against the unmodified base model and a simpler alternative.
- Reproducible training configurations and versioned artifacts.
- Evaluation across scripts, dialects, code-mixing, and noisy input.
- Monitoring for drift, harmful outputs, and language-specific failures.
- A rollback path and a clear model card for users and developers.
LoRA makes customisation affordable, but it does not replace language expertise. The strongest Indic-language systems combine a capable base model, carefully governed data, script-aware preprocessing, targeted adapter training, and evaluation led by native speakers.