0tokens

Apply for AI Grants India

Financial support for innovators building the future of AI in India.

Apply now

Chat · how to tokenize indian languages for small language models

How to Tokenize Indian Languages for Small Language Models

  1. aigi

    Indian-language tokenization is not a preprocessing detail to postpone. For a small language model, the tokenizer determines how much text fits into context, how efficiently the model learns morphology, and whether rare names, spellings, and mixed-script input remain usable. A poorly designed vocabulary can make Hindi, Tamil, Bengali, Marathi, Telugu, Kannada, Malayalam, Gujarati, Punjabi, Odia, Assamese, and other Indic languages appear far more data-hungry than they are.

    This guide explains how to tokenize Indian languages for small language models with a practical workflow for builders working with limited compute, modest corpora, and real-world Indian text.

    Start with the data, not the algorithm

    Before choosing BPE, Unigram, or WordPiece, profile the corpus you intend to train on. Indian-language data varies sharply by domain: news, government documents, social media, customer support, education, and speech transcripts have different punctuation, spelling, code-mixing, and script usage.

    Create a language and domain inventory covering:

    • Target languages and their scripts
    • Document sources, licences, and deduplication status
    • Native-script text versus Romanised text
    • English and other-language code-mixing
    • Numerals, currency symbols, emojis, URLs, and markup
    • Orthographic variants, spelling errors, and transliteration patterns

    For a low-resource language, a small clean corpus is usually more valuable than a large noisy scrape. The broader principles in this low-resource Indic NLP builder’s guide are especially relevant when data and compute are constrained.

    Normalize Unicode without destroying meaning

    Indic scripts use combining marks and visually similar sequences that may have different underlying Unicode representations. Normalize text consistently before training the tokenizer, preferably with a tested Unicode normalization pipeline. Also inspect script-specific edge cases rather than assuming a generic regex is safe.

    A robust cleaning pass should:

    • Apply consistent Unicode normalization
    • Remove accidental control characters and broken markup
    • Preserve meaningful punctuation and sentence boundaries
    • Standardize whitespace without joining separate words
    • Handle zero-width joiners and non-joiners deliberately
    • Decide how to represent danda characters, quotation marks, and ellipses
    • Preserve meaningful currency symbols and regional numerals

    Do not blindly strip combining marks, joiners, or punctuation. In languages such as Hindi, Bengali, Tamil, Telugu, and Malayalam, these can affect rendering, word identity, or grammatical interpretation. Keep a before-and-after sample set and run character-frequency reports after every cleaning change.

    Choose a tokenization strategy

    Word tokenization: useful for analysis, rarely enough for training

    Whitespace and punctuation splitting is useful for corpus statistics, sentence filtering, and baseline experiments. It is not a reliable final tokenizer for small language models because inflection, compounds, spelling variation, and unseen words create a large unknown-word problem.

    Character tokenization: robust but expensive

    Character-level tokenization handles arbitrary words and spelling errors, but sequences become long. That increases attention cost and reduces the amount of semantic content a compact model can process in one context window. It can be useful for specialist experiments, noisy OCR, or extremely small vocabularies, but it is usually not the best general-purpose choice.

    Subword tokenization: the practical default

    For most Indic small language models, start with a subword tokenizer. BPE merges frequent symbol sequences; Unigram selects a probabilistic vocabulary of useful pieces; WordPiece is another established option. SentencePiece is particularly convenient because it can train directly on raw text without requiring whitespace-based pre-tokenization.

    Subwords should capture recurring linguistic units without fragmenting every common word into too many pieces. Train on representative text, not only on high-resource languages or English-heavy material.

    Design the vocabulary for Indian usage

    Vocabulary size is a trade-off between sequence length, embedding memory, and coverage. A larger vocabulary can reduce token counts but consumes parameters and may memorize noisy or overly specific strings. A smaller vocabulary saves memory but can produce inefficient sequences.

    For a multilingual model, measure each language separately:

    • Average tokens per word
    • Average tokens per character
    • Percentage of <unk> or fallback tokens
    • Share of single-character and very long tokens
    • Token fertility for common words and names
    • Coverage of numbers, punctuation, and code-mixed text

    Do not judge the tokenizer by English compression alone. A tokenizer that produces short English sequences but splits common Tamil or Assamese words into many pieces is poorly balanced. Consider language-aware sampling during tokenizer training so high-resource languages do not dominate the vocabulary.

    A useful approach is to reserve capacity for script and language coverage, then compare several vocabulary sizes on a held-out multilingual set. For a compact model, a smaller balanced vocabulary often beats a larger vocabulary trained on skewed data.

    Treat Romanised and code-mixed text as first-class inputs

    Indian users frequently type Indic languages in Latin script, mix English with an Indic language, or switch scripts within a single message. A native-script-only tokenizer may perform well on formal benchmarks and fail on actual product traffic.

    Decide explicitly whether the model will:

    • Support only native scripts
    • Support native and Romanised forms
    • Normalize Romanised text into a native script
    • Preserve both forms for retrieval or generation

    If Romanised input matters, include representative examples during tokenizer training. Avoid aggressive transliteration that erases distinctions needed by the application. For customer support, search, and voice interfaces, evaluate spelling variation and phonetic transliteration—not just clean literary text. Product teams building multilingual voice experiences can also review these voice agent benefits for Indian businesses to understand why noisy user input matters.

    Train and test the tokenizer systematically

    Keep tokenizer training data separate from evaluation data. Train multiple candidates with controlled vocabulary sizes and compare them on the same held-out set. Useful checks include:

    • Token counts per sentence and per language
    • Unknown-token rate
    • Compression ratio by script
    • Fertility of frequent words, names, and inflected forms
    • Stability under punctuation and whitespace changes
    • Behaviour on code-mixed and Romanised input
    • Preservation of URLs, email addresses, decimals, and dates

    Inspect token sequences manually. Quantitative metrics can hide undesirable merges, such as a token that combines a common stem with an unrelated suffix or a token that memorizes a single website fragment. Examine examples from government forms, school content, chat messages, product reviews, and speech-recognition transcripts.

    Align tokenization with model training

    Tokenizer quality cannot be separated from the model’s training mixture. If the tokenizer sees balanced Indic data but pretraining overwhelmingly uses English, the model will still underperform on Indian languages. Deduplicate documents, track language proportions, and report results by language rather than only as a single average.

    For continued pretraining or fine-tuning, keep the tokenizer fixed unless there is a strong reason to change it. Replacing the vocabulary changes input embeddings and can invalidate checkpoints. If a new language must be added, assess vocabulary expansion or adapter-based approaches before retraining the entire model.

    For builders publishing reusable work, document the corpus sources, normalization rules, training command, vocabulary size, special tokens, and known failure cases. India’s open-source ecosystem is increasingly active; projects listed in this guide to Indian open-source AI developer projects offer useful examples of how reproducibility improves adoption.

    Evaluate downstream usefulness, not just compression

    A tokenizer is successful when it improves the application. Test it on the tasks the model must perform:

    • Next-token prediction and perplexity by language
    • Classification and retrieval across scripts
    • Named-entity recognition for Indian names and places
    • Question answering and instruction following
    • Translation or transliteration where relevant
    • Speech-to-text post-processing and voice-agent queries

    Build a small, reviewed challenge set with spelling variants, compound words, punctuation, numerals, code-mixing, and low-resource languages. Record both token-level metrics and user-facing outcomes such as accuracy, latency, context capacity, and hallucination rate.

    A practical default pipeline

    For many teams in 2026, a sensible starting point is:

    1. Collect licensed, domain-representative text for each target language.
    2. Normalize Unicode and clean markup while preserving linguistic signals.
    3. Deduplicate and split data by document or source.
    4. Train a SentencePiece BPE or Unigram tokenizer on balanced multilingual samples.
    5. Compare vocabulary sizes using per-language fertility and unknown-token metrics.
    6. Add carefully selected Romanised and code-mixed data if the product requires it.
    7. Freeze the best tokenizer and evaluate the model on downstream Indic tasks.
    8. Publish configuration, limitations, and representative tokenization examples.

    The key principle is simple: optimize for useful multilingual coverage, not the smallest token count. A compact model benefits from a tokenizer that respects Indic scripts, handles everyday variation, and gives every target language enough representation to learn efficiently.

    Last updated 23 September 2026

AIGI may be inaccurate. Replies seeded from the guide above.