0tokens

Apply for AI Grants India

Financial support for innovators building the future of AI in India.

Apply now

Chat · how to train llms on indian datasets

How to Train LLMs on Indian Datasets

  1. aigi

    Indian-language LLM work is not a simple translation exercise. A useful model must handle multiple scripts, rich morphology, spelling variation, Romanised text, code-mixed prompts, regional terminology, and uneven data availability. It must also respect consent, privacy, copyright, and the social context in which Indian-language systems are deployed.

    For most teams, the right goal is language adaptation or domain adaptation, not training a foundation model from zero. Start with a strong open-weight base model, define the languages and use cases you actually need, build a defensible dataset pipeline, and measure performance with native-speaker review. This guide lays out that workflow for 2026.

    1. Define the model’s job before collecting data

    “Indian languages” is too broad to be a training specification. Decide which combination of language, script, domain, and task matters to your product.

    Write down:

    • Target languages: Hindi, Bengali, Tamil, Telugu, Marathi, Kannada, Malayalam, Gujarati, Punjabi, Odia, Assamese, Urdu, or a smaller low-resource set.
    • Script requirements: Native scripts, Romanised input, transliteration output, or all three.
    • Tasks: Chat, retrieval-augmented generation, translation, summarisation, classification, speech transcripts, or tool use.
    • Domains: Education, healthcare, agriculture, finance, government services, customer support, or general conversation.
    • Quality targets: Factuality, instruction following, terminology accuracy, latency, and acceptable hallucination rates.

    A model for a Kannada education assistant needs different data from a multilingual call-centre model. If voice is central, pair language-model work with a carefully designed speech and transcript pipeline; the principles in this technical guide to building voice agents are relevant for the downstream system.

    2. Choose the right training strategy

    Start with an open-weight base model

    Training from scratch is justified only when you have a large, legally usable corpus, sustained GPU access, tokenizer expertise, and a clear reason existing models cannot meet your requirements. Otherwise, begin with an established multilingual or Indian-language-capable model.

    Use continued pre-training when the base model needs stronger fluency in a language or domain. Feed it high-quality, language-balanced text with the causal-language-modelling objective. Use supervised fine-tuning (SFT) when the model already understands the language but needs better instruction following, formatting, or task behaviour. Use LoRA or another PEFT method when budget, iteration speed, or serving multiple adapters matters. A practical overview of the trade-offs is available in this guide to fine-tuning LLMs on custom data.

    Do not assume that more training always improves results. Continued pre-training can cause catastrophic forgetting, reduce English performance, or amplify noisy patterns. Maintain a held-out multilingual evaluation set and compare every run against the original checkpoint.

    3. Build a lawful, representative data mixture

    Useful sources may include:

    • Government and public-sector text: Laws, schemes, public notices, education material, and official translations, subject to the source’s licence and terms.
    • Open research datasets: Indic language corpora, translation sets, labelled NLP datasets, and speech transcripts from reputable research projects.
    • Licensed commercial data: Professionally translated or domain-specific material where training rights are explicit.
    • Synthetic data: Generated instructions, translations, and paraphrases, always filtered against source contamination and factual errors.
    • Product data: User interactions only when consent, notice, retention limits, and applicable contractual terms permit their use.

    BHASHINI and AI4Bharat are important places to investigate, but neither should be treated as a blanket guarantee that every asset can be used for every training purpose. Record the dataset version, licence, provenance, collection method, language, domain, and known limitations in a data card.

    Balance the mixture by tokens and quality, not only by document count. A million short Hindi social posts do not provide the same coverage as a carefully curated legal or agricultural corpus. For low-resource languages, retain valuable examples rather than aggressively filtering them to match high-resource languages.

    4. Clean scripts, language labels, and duplicates

    Indian-language data commonly contains Unicode inconsistencies, OCR mistakes, spelling variation, mixed scripts, boilerplate, and unreliable language labels. Build preprocessing as a reproducible pipeline rather than a one-off notebook.

    Recommended stages include:

    • Unicode normalisation, including consistent handling of combining marks and punctuation.
    • Script detection and language identification at document and segment level.
    • Removal of navigation text, spam, boilerplate, malformed HTML, and repeated templates.
    • Exact and near-duplicate removal using hashes, MinHash, or locality-sensitive hashing.
    • Quality scoring for fluency, completeness, translation alignment, and domain relevance.
    • PII detection and redaction for names, phone numbers, addresses, IDs, medical details, and financial information.
    • Toxicity and abuse review using multilingual classifiers plus native-speaker audits.

    Preserve useful variation. Do not “correct” dialects, informal spelling, or code-mixed usage simply because they differ from textbook language. Store both the original and normalised forms when possible, and document which version enters each training stage.

    5. Treat tokenization as an engineering metric

    A tokenizer trained mostly on English can split Indic text inefficiently. That increases sequence length, memory use, inference latency, and cost. Measure token fertility—the number of tokens per character or word—by language, script, and domain before changing the vocabulary.

    Possible approaches are:

    • Use the base tokenizer if efficiency and quality are already acceptable.
    • Extend the vocabulary with frequent Indic subwords, then resize embeddings and carefully retrain or adapt the model.
    • Train a new BPE or Unigram tokenizer on a balanced, cleaned corpus when the existing vocabulary is severely unsuitable.

    Evaluate tokenization on native scripts, Romanised text, numerals, punctuation, names, and code-mixed sentences. A tokenizer that performs well on formal Hindi may still fail on Romanised Marathi or Tamil-English customer messages. Vocabulary expansion also changes checkpoint compatibility and can require substantial embedding training, so benchmark it against a no-vocabulary-change baseline.

    6. Train for code-mixing and transliteration explicitly

    Real users may write “kal meeting hai,” “நாளைக்கு call பண்ணலாம்,” or switch scripts within one sentence. Include these patterns deliberately rather than expecting them to emerge from monolingual text.

    Build training slices for:

    • Native-script conversation.
    • Romanised Indian languages with spelling variation.
    • English-plus-Indic code-mixing.
    • Transliteration in both directions.
    • Regional names, places, products, and abbreviations.
    • Speech-like disfluencies if the model will process transcripts.

    Keep language tags or metadata during evaluation so you can identify whether a model fails because of language confusion, script handling, or factual weakness. For high-stakes systems, allow users to select a preferred language and script instead of relying entirely on automatic detection.

    7. Design an evaluation set that reflects India

    Do not use English benchmarks as a proxy for Indian-language quality. Create a balanced test suite with native speakers and task-specific examples.

    Measure:

    • Instruction following and refusal behaviour in each target language.
    • Translation adequacy and terminology preservation.
    • Factual question answering against verified references.
    • Summarisation of government, educational, legal, and health text.
    • Robustness to spelling errors, dialect variation, Romanisation, and code-mixing.
    • Toxicity, caste and religious stereotyping, gender bias, and region-specific harms.
    • Token efficiency, throughput, memory use, and latency.

    Use blind human evaluation with at least two reviewers per item where feasible, adjudication for disagreements, and separate scores for fluency, meaning preservation, cultural appropriateness, and factuality. Track performance by language rather than reporting one aggregate score that hides weak languages.

    8. Keep compute and deployment economics realistic

    A small, well-curated adaptation can outperform a much larger but noisy training run for a narrow product. Use gradient checkpointing, mixed precision, packed sequences, and distributed training only when they solve a measured bottleneck. Start with a pilot on one or two languages, establish data and evaluation quality, then scale.

    For production, consider quantisation, batching, speculative decoding, and retrieval rather than forcing every fact into model weights. Keep sensitive Indian data within the required residency and access controls, and log prompts and outputs only under a documented retention policy.

    If you are building an application rather than a base model, choose frameworks that support rapid evaluation, retrieval, observability, and deployment. This complements the ecosystem covered in AI frameworks for Indian student entrepreneurs.

    9. Governance is part of the training pipeline

    Create dataset documentation before training begins. Record provenance, licence, consent basis, demographic and geographic coverage, annotation instructions, known exclusions, and removal procedures. Establish a process for takedown requests, privacy complaints, and harmful-output reports.

    For healthcare, finance, education, and government use, add domain experts to review prompts and outputs. A language model that sounds fluent can still provide unsafe advice. Retrieval, citations, constrained generation, human escalation, and clear user disclosures are often more valuable than another round of fine-tuning.

    Practical starting plan

    A lean team can begin with this sequence:

    1. Select one high-value use case and two or three target languages.
    2. Audit an open-weight base model for tokenization, fluency, and task performance.
    3. Assemble a licensed pilot corpus with provenance records.
    4. Build language identification, deduplication, PII, and quality filters.
    5. Run a small continued-pretraining or LoRA experiment with held-out data.
    6. Evaluate native-script, Romanised, and code-mixed inputs separately.
    7. Conduct native-speaker safety and factuality reviews.
    8. Compare quality gains against training, inference, and monitoring costs.
    9. Expand languages only after the pipeline is reproducible.

    FAQ

    Do I need to train an LLM from scratch? Usually not. Start with a capable open-weight model and test continued pre-training, SFT, LoRA, or retrieval before considering full pre-training.

    Which Indian languages should I support first? Choose based on users and data quality, not population alone. Hindi has substantial resources, while several other languages may offer stronger product differentiation but require more curation and native-speaker review.

    Is synthetic data enough for low-resource languages? No. It can expand coverage, but synthetic examples need human checks, deduplication, provenance labels, and evaluation against authentic usage.

    What should I measure first? Measure task success by language, script, and input style, alongside token efficiency, latency, factuality, and safety. Aggregate scores alone are misleading.

    AI Grants India supports builders working on language technology and AI products for Indian users. Explore the AI Grants India application if your project needs funding, technical visibility, or ecosystem support.

    Last updated 23 September 2026

AIGI may be inaccurate. Replies seeded from the guide above.