What fine-tuning should solve
Fine-tuning Llama for Indian regional languages is useful when prompting alone cannot deliver reliable fluency, script handling, terminology, or instruction following. The objective is not simply to make a model produce more Hindi or Tamil. A production model should understand native scripts, tolerate Romanised input where users actually type it, preserve meaning across code-switching, and respond safely in the intended cultural and institutional context.
Start by defining the job clearly:
- Language coverage: one language, a language pair, or a multilingual model across Indic languages.
- Script coverage: native script, Roman transliteration, or both.
- Task type: chat, classification, translation, summarisation, extraction, speech transcripts, or domain-specific assistance.
- Quality target: human-rated fluency, factual accuracy, terminology consistency, latency, and cost.
For a broader training plan, use the principles in best practices for fine-tuning LLMs on custom data before choosing a model or GPU configuration.
Choose the base model and adaptation method
Use an instruction-tuned Llama checkpoint when the application needs conversational behaviour. A base pretrained checkpoint may be preferable for continued pretraining on a large, clean corpus, but it requires more data and careful post-training. For most Indian startups, continued pretraining followed by supervised fine-tuning is only justified when the target language or domain is poorly represented. Otherwise, begin with supervised fine-tuning or QLoRA and establish a baseline.
A sensible progression is:
1. Prompt the base or instruction model with representative examples.
2. Measure errors on a held-out, native-speaker-reviewed test set.
3. Apply QLoRA to the smallest model that meets the quality bar.
4. Compare against retrieval, translation pipelines, or a multilingual model.
5. Scale to a larger checkpoint only when evaluation shows a clear benefit.
QLoRA loads the model in 4-bit precision, freezes the base weights, and trains compact LoRA adapters. It reduces memory requirements and makes experiments practical on rented GPUs or Indian cloud infrastructure. Adapter training is usually easier to version, roll back, and maintain than repeatedly publishing full model copies. Teams building Indian open-source AI developer projects can also release adapters separately, subject to the base model’s licence.
Audit tokenisation before changing the tokenizer
Tokenisation is a measurable engineering issue, not an assumption. Take a sample of real user text in Devanagari, Bengali, Gujarati, Kannada, Malayalam, Marathi, Odia, Punjabi, Tamil, Telugu and Romanised variants. Compare:
- Tokens per word and tokens per character.
- Truncation rates at the planned context length.
- Memory and latency for typical prompts.
- Fragmentation of common suffixes, names, numbers and punctuation.
Indic scripts can be inefficiently represented by a tokenizer trained on a corpus dominated by other languages. That raises inference cost and can weaken learning from limited data. However, adding tokens to the vocabulary is not automatically the right fix. It changes the embedding matrix and may complicate compatibility with existing checkpoints. First test whether better data, sequence packing, a language-specific model, or an adapter provides enough improvement.
If you do extend the tokenizer, add frequent, linguistically meaningful units rather than arbitrary characters. Retrain or initialise new embeddings carefully, then evaluate the modified model against the unchanged-tokenizer baseline. Keep native-script and Romanised text in separate evaluation slices; a model that performs well on one may fail on the other.
Build a dataset that reflects Indian usage
Data quality matters more than raw volume. Useful sources can include AI4Bharat corpora, openly licensed government material, domain documents, public educational content, and consented product interactions. Check the licence and permitted use for every source. Do not treat scraped news, books, social posts, or translated material as automatically reusable for commercial training.
Prepare the corpus with language-aware processing:
- Deduplicate near-identical documents and translated copies.
- Remove boilerplate, broken markup, spam, and machine-generated repetition.
- Preserve punctuation, sentence boundaries, diacritics and meaningful code-switching.
- Normalise Unicode without erasing distinctions users rely on.
- Detect language at the sentence or turn level, not only per document.
- Filter personal data, credentials, phone numbers and sensitive case details.
- Keep provenance, licence, source date and transformation history.
Instruction data should resemble the product. Include short queries, long context, spelling variation, Romanised input, mixed Hindi-English or Tamil-English turns, names, numerals, local units, and requests that require refusal. Synthetic translations can expand coverage, but native speakers must review them for unnatural phrasing, gender agreement, honorifics, idioms and regional terminology. Never allow synthetic examples to dominate the validation set.
Design the training recipe
Format examples with the exact chat template expected by the Llama checkpoint. Incorrect role markers or duplicated system prompts can reduce performance. Keep the system instruction stable and put task variation in user turns. For supervised fine-tuning, mask loss on user instructions when appropriate so training focuses on the assistant response.
Practical starting points for QLoRA include:
- Rank: 16 or 32; test 8, 16 and 32 rather than assuming a larger rank is better.
- Target modules: attention projections, with MLP projections added if language adaptation is weak.
- Learning rate: begin conservatively and run a small sweep; high rates can damage general capabilities.
- Sequence length: use the longest length supported by actual product traffic, with packed examples where possible.
- Epochs: stop based on held-out loss and human quality, not a fixed epoch count.
- Precision and memory: use 4-bit loading, gradient checkpointing, accumulation and efficient attention where supported.
Keep separate train, validation and test sets by source and, where possible, by topic. Randomly splitting near-duplicate translations can produce misleading scores. Log the base model revision, tokenizer, data hash, adapter settings, random seed and evaluation results for every run.
Evaluate language quality and safety
BLEU or ROUGE can help with narrow translation or summarisation tasks, but they are poor proxies for conversational naturalness. Build a multilingual test suite with native reviewers. Score fluency, meaning preservation, instruction following, factuality, terminology, script handling, code-switching and politeness. Report results per language and input mode instead of publishing one blended score.
Include adversarial tests for:
- False premises and hallucinated government schemes.
- Unsafe medical, financial and legal advice.
- Caste, religion, gender and regional slurs.
- Prompt injection in mixed scripts.
- Personally identifiable information and requests to expose it.
- Dialect variation and ambiguous transliteration.
For applications such as education, test the model with curriculum-aligned examples rather than generic benchmarks. A classroom assistant may need different controls from an AI tutor for Indian competitive exams, especially around answer explanations, source citation and uncertainty.
Deploy with a realistic Indian product architecture
Fine-tuning does not replace retrieval, translation, speech recognition or application controls. For changing information, use retrieval with dated, authoritative sources. For voice products, evaluate the complete chain—speech recognition, language model and text-to-speech—because errors often enter before generation. This matters when adapting systems for voice agent services for Indian businesses.
Serve the adapter separately when your inference stack supports it, and merge only when it simplifies latency or compatibility. Quantise the final model after measuring quality, not before. Track per-language latency, token usage, failure rates, escalation rates and user corrections. Provide a fallback to a stronger multilingual model or human support when confidence is low.
A practical launch checklist
Before release, confirm that you have:
- A language-by-script coverage matrix and tokenizer audit.
- Licensed, deduplicated and documented training data.
- Native-speaker validation for every supported language.
- Separate tests for native script, Romanisation and code-switching.
- Safety evaluations covering local terms and sensitive contexts.
- Reproducible adapter, tokenizer and dataset versions.
- Monitoring for drift, user corrections and harmful outputs.
- A rollback path and a process for removing problematic data.
The strongest Indian-language Llama systems are built through disciplined measurement rather than a single large training run. Start with a narrow use case, improve the data and evaluation loop, and expand language coverage only after the model is dependable for the users it serves.