Indian-language fine-tuning often fails before training begins. A tokenizer that fragments Kannada, Malayalam, or Hindi text excessively can inflate sequence lengths, discard useful distinctions, and make a limited dataset even less effective. The solution is not simply to split text into words: modern Hugging Face models use subword tokenizers whose vocabulary, normalisation rules, and special tokens must match the base checkpoint.
This guide shows how to build a dependable tokenisation pipeline for Indic data in 2026. It covers data cleaning, tokenizer selection, dataset mapping, sequence-length checks, language mixing, and the diagnostics that matter before committing GPU time. For broader modelling decisions, pair this workflow with low-resource Indic NLP guidance and the practical principles in best practices for fine-tuning LLMs.
1. Start with the model, not the language
Load the tokenizer that belongs to the exact model checkpoint you will fine-tune. Do not substitute a generic Hindi or multilingual tokenizer because it appears convenient. The model’s embedding matrix expects the token IDs and special-token conventions produced by its own tokenizer.
from transformers import AutoTokenizer
model_id = "ai4bharat/indicbert-v2-mlm" # replace with your checkpoint
tokenizer = AutoTokenizer.from_pretrained(model_id, use_fast=True)
print(tokenizer.special_tokens_map)For encoder tasks such as classification, token classification, and masked language modelling, an Indic-focused encoder may outperform a general multilingual checkpoint when your language is well represented. For generation, use the tokenizer shipped with the selected causal or sequence-to-sequence model. Multilingual models such as XLM-R can be useful across languages, but performance depends on pre-training coverage and the balance of your fine-tuning set.
Check four properties before training:
- Vocabulary coverage: how many tokens are split into unknown or unusually long pieces?
- Sequence expansion: how many tokens does one sentence require compared with its whitespace word count?
- Script handling: does the tokenizer preserve the intended Unicode script and punctuation?
- Language balance: does one high-resource language dominate mixed batches?
2. Normalise Indic text carefully
Indian-language data commonly combines Unicode variants, copied web text, emoji, English words, numerals, punctuation, and transliterated content. Cleaning should remove accidental noise without erasing information. Never apply ASCII-only lowercasing, strip diacritics blindly, or transliterate everything to Latin script unless that is explicitly the product requirement.
A practical first pass is:
- normalise Unicode consistently, usually with NFC;
- replace null bytes and malformed control characters;
- standardise line breaks and repeated whitespace;
- retain punctuation, numerals, emoji, and meaningful code-switching;
- remove exact duplicates and near-duplicate documents;
- record the original text before any destructive transformation.
import re
import unicodedata
def clean_text(text):
text = "" if text is None else str(text)
text = unicodedata.normalize("NFC", text)
text = text.replace("\x00", " ")
text = re.sub(r"\s+", " ", text).strip()
return textKeep language labels, source identifiers, licence information, and quality flags in separate columns. This makes it possible to audit whether a model is learning a script, a source-specific style, or duplicated content. For high-stakes applications, treat provenance and validation as part of data veracity infrastructure, not as an afterthought.
3. Measure tokenisation before fine-tuning
Tokenise a representative sample for every target language and content type. Include short messages, long documents, code-switched sentences, names, addresses, numbers, and punctuation-heavy text. Inspect both token IDs and decoded tokens; aggregate statistics alone can hide broken examples.
samples = [
"यह एक उदाहरण वाक्य है।",
"தமிழில் ஒரு சோதனை வாக்கியம்.",
"বাংলা এবং English একসঙ্গে ব্যবহার করা হয়েছে।",
]
for text in samples:
encoded = tokenizer(text, add_special_tokens=True)
print(text)
print(tokenizer.convert_ids_to_tokens(encoded["input_ids"]))
print("tokens:", len(encoded["input_ids"]))Track the median and 95th-percentile token count by language. Also calculate the percentage of examples exceeding your planned maximum length and the proportion containing unknown tokens. Excessive fragmentation increases memory use and can truncate important context. If a language consistently needs far more tokens, compare another compatible checkpoint rather than simply increasing the context window.
4. Tokenise with a Hugging Face Dataset
For supervised fine-tuning, store raw text and labels in a Dataset, then use batched mapping. Avoid tokenising the entire CSV into one large in-memory tensor: dynamic padding during collation is usually more efficient.
from datasets import load_dataset
raw = load_dataset("csv", data_files={
"train": "train.csv",
"validation": "validation.csv"
})
max_length = 512
def tokenize_batch(batch):
return tokenizer(
batch["text"],
truncation=True,
max_length=max_length,
padding=False,
)
tokenized = raw.map(
tokenize_batch,
batched=True,
remove_columns=["text"],
desc="Tokenising Indic text",
)Use padding=False during preprocessing and a data collator for batch-level padding:
from transformers import DataCollatorWithPadding
collator = DataCollatorWithPadding(
tokenizer=tokenizer,
pad_to_multiple_of=8,
)For causal language modelling, do not use the classification recipe unchanged. Decide whether documents should be packed into continuous sequences, add an end-of-document token where appropriate, and ensure labels mask padding tokens. For token classification, pass aligned word labels and use the tokenizer’s word_ids() mapping; naïvely copying one label to every subword can corrupt entity boundaries.
5. Handle truncation and long Indic documents
A fixed max_length=512 is only a starting point. Analyse your actual length distribution. For classification, truncation may be acceptable if the decisive evidence appears early; for legal, medical, educational, or support text, it may remove the information the model needs.
Useful strategies include:
- chunk long documents with a documented stride;
- use hierarchical classification when document-level context matters;
- pack short examples for causal language modelling;
- preserve sentence or turn boundaries where possible;
- compare token budgets across languages before setting one global limit.
Do not silently truncate. Log the number of affected examples and sample the truncated regions. A high truncation rate is a data or architecture decision that needs review, not a harmless preprocessing detail.
6. Build language-aware splits and batches
Random row-level splits can leak near-duplicate translations, conversations, or documents across train and validation sets. Split by source, speaker, customer, document, or time period where appropriate. Keep a per-language validation breakdown even when training a multilingual model.
For mixed Indic datasets, inspect:
- macro-F1 or task-appropriate metrics across languages;
- performance by script and code-switching rate;
- token counts and truncation rates per language;
- error patterns involving names, numerals, and punctuation;
- calibration and refusal behaviour for deployment use cases.
Oversampling a smaller language can help, but it changes the effective training distribution. Record the sampling policy and compare it with loss weighting, language tags, or continued pre-training. Builders working on open tooling may also benefit from surveying Indian open-source AI developer projects for compatible checkpoints and evaluation resources.
7. Common mistakes to avoid
- Training a new tokenizer unnecessarily: start with the checkpoint’s tokenizer unless you are also prepared to resize embeddings and retrain or substantially adapt the model.
- Mixing incompatible tokenizers and models: token IDs have meaning only within the model vocabulary they were designed for.
- Removing all punctuation: punctuation carries structure in Indian scripts, chat, speech transcripts, and code-switched text.
- Assuming whitespace equals words: subword boundaries differ across scripts and model families.
- Padding every example to the global maximum: dynamic padding reduces wasted computation.
- Evaluating only aggregate accuracy: a strong overall score can conceal failure on a smaller language.
- Ignoring licensing and consent: web text, speech transcripts, and user-generated content need documented rights and handling controls.
8. A pre-training checklist
Before launching fine-tuning, confirm that you can answer these questions:
1. Does the tokenizer match the exact base model?
2. Have Unicode normalisation and cleaning rules been tested on each script?
3. What are token-length, unknown-token, and truncation rates by language?
4. Are train and validation splits protected against source leakage?
5. Are labels aligned correctly for subword tokens?
6. Is dynamic padding or sequence packing configured for the task?
7. Can you reproduce the dataset version, tokenizer version, and preprocessing code?
8. Will evaluation expose performance gaps across languages and scripts?
Tokenisation is not a cosmetic preprocessing step. It determines how much of an Indian-language example reaches the model, how expensive training becomes, and whether comparisons across languages are meaningful. Measure it on real data, preserve provenance, and make tokenizer decisions together with model, task, and deployment constraints.