Marathi language models often fail before model training begins: the tokenizer fragments common words, mishandles Devanagari combining marks, or wastes vocabulary on formatting noise. A well-designed tokenizer improves sequence efficiency, reduces unknown or overly fragmented tokens, and gives downstream models a stronger representation of Marathi text.
This guide explains how to train a tokenizer for Marathi language models in 2026, with practical choices for corpus construction, normalization, vocabulary size, evaluation, and integration with modern language-model tooling.
1. Define the tokenizer’s job
Start with the model and product requirements, not with a library. A tokenizer for a Marathi-only generative model may use a Devanagari-focused vocabulary. A multilingual model needs to balance Marathi coverage against Hindi, English, other Indic scripts, and code-mixed text.
Decide whether you need to support:
- Marathi in Devanagari only, or Romanised Marathi as well
- News, books, government documents, conversations, or social media
- Generative modelling, classification, speech transcripts, or search
- Long-context efficiency and low token counts
- Compatibility with an existing model and vocabulary
If you are working with limited data, review the principles in Low-Resource Indic Natural Language Processing: A Builder’s Guide. Marathi tokenizer quality depends heavily on data coverage, licensing, and consistent preprocessing.
2. Build a representative Marathi corpus
Collect text from sources that match your intended deployment. Possible sources include licensed news, public-domain literature, educational material, government publications, subtitles, user-provided conversations, and carefully filtered web data. Low-Resource Language Datasets for AI Training in India provides useful context for sourcing and documenting Indian-language datasets.
Keep metadata for each document: source, date, licence, domain, script, and quality label. Split documents into training, validation, and test sets before training the tokenizer. Avoid placing near-duplicate articles or repeated website templates across splits; otherwise, token coverage will look better than it is.
A useful corpus should include:
- Marathi inflections, compounds, names, places, and technical terms
- Formal and informal registers
- Punctuation, numbers, dates, currency, and units used in India
- English and Hindi code-mixing where it appears in real user data
- Spelling variation, typographical errors, and common transliteration patterns
Do not remove every number, symbol, or emoji. A tokenizer should represent the text your application will actually receive.
3. Normalise Devanagari carefully
Unicode normalisation is essential, but aggressive cleaning can destroy useful distinctions. Apply a documented, deterministic pipeline to both tokenizer-training data and model inputs.
Typical steps include:
- Convert text to Unicode NFC where appropriate
- Remove invisible control characters and accidental formatting marks
- Standardise whitespace without deleting meaningful line or paragraph boundaries
- Inspect virama, vowel signs, nukta, zero-width characters, and conjunct formation
- Decide how to handle danda (
।), double danda (॥), quotation marks, and hyphens - Preserve Marathi numerals or convert them only if the model specification requires it
- Keep case handling in mind for Latin-script code-mixed text
Do not strip all punctuation or “stop words” before training a subword tokenizer. Stop-word removal is an NLP-task preprocessing decision, not a tokenizer-training requirement. Maintain a corpus audit with examples before and after every normalisation rule.
4. Choose a subword algorithm
For most Marathi language models, subword tokenization is the practical default. It represents frequent words efficiently while decomposing rare inflected forms and names into reusable pieces.
Common options are:
- Unigram, available through SentencePiece: useful when several segmentations are plausible and often a strong baseline for Indic languages.
- Byte Pair Encoding (BPE): straightforward, widely supported, and effective when trained on a sufficiently varied corpus.
- WordPiece: useful when compatibility with an existing BERT-style ecosystem matters.
- Byte-level BPE: robust to unseen bytes and mixed scripts, but inspect whether it produces efficient Marathi segmentations.
SentencePiece can train directly from raw text and does not require whitespace-based word boundaries. Hugging Face’s tokenizers library offers fast implementations and easy integration with Transformer models. For a new Marathi causal language model, compare Unigram and BPE on the same corpus rather than assuming one is best.
5. Train a baseline tokenizer
Install the tooling in a reproducible environment, pin versions, and save the exact training configuration. A SentencePiece baseline might look like this:
spm_train \
--input=marathi_train.txt \
--model_prefix=marathi_unigram \
--vocab_size=32000 \
--model_type=unigram \
--character_coverage=0.9995 \
--normalization_rule_name=nmt_nfkcThe correct vocabulary size depends on corpus size, model budget, and multilingual scope. Test several values, such as 16k, 32k, and 64k, rather than treating 32k as universal. A vocabulary that is too small causes excessive fragmentation; one that is too large increases embedding and output-layer costs and may memorise rare noise.
For a Marathi-only corpus, inspect whether the vocabulary gives frequent Marathi morphemes and endings their own tokens. For a multilingual model, reserve capacity for all target scripts and measure Marathi performance separately. Add explicit special tokens—such as beginning-of-sequence, end-of-sequence, padding, unknown, and mask tokens—only according to the model architecture.
6. Evaluate token quality, not just vocabulary size
Evaluate the tokenizer on a held-out set containing both ordinary and difficult examples. Track:
- Average tokens per word and tokens per character
- Percentage of text represented as unknown or fallback tokens
- Marathi versus English token efficiency in mixed text
- Fragmentation of frequent words, names, and inflected forms
- Coverage of numbers, punctuation, emojis, and URLs
- Sequence length distribution for real application inputs
A low unknown-token rate is necessary but not sufficient. A tokenizer can avoid unknown tokens while splitting nearly every Marathi word into inefficient fragments. Inspect representative outputs manually, including words with conjuncts, vowel signs, suffixes, proper nouns, and spelling variants.
Compare candidate tokenizers on downstream tasks such as Marathi language modelling, classification, retrieval, or translation. A slightly larger vocabulary may be worthwhile if it materially reduces sequence length and improves validation loss. When fine-tuning an existing model, do not casually replace its tokenizer: changing the vocabulary changes the embedding interface and normally requires continued pretraining or model surgery.
7. Handle code-mixing and Romanised Marathi
Many Indian applications receive Marathi mixed with English, Hindi, Latin-script names, numerals, and chat abbreviations. Decide whether this is part of the supported input distribution. If it is, include representative examples in the corpus and test them separately.
Romanised Marathi deserves special treatment. You can include it in one shared tokenizer, create a transliteration preprocessing layer, or support separate input normalisation paths. Transliteration may improve consistency, but it can also erase distinctions and introduce errors. Measure the full pipeline on real user text rather than evaluating Devanagari alone.
8. Package and integrate the tokenizer
Save the model file, vocabulary, special-token mapping, normalisation rules, pre-tokeniser configuration, and training manifest together. Version them as a single artefact. Record corpus hashes, licences, vocabulary size, library versions, and evaluation results so another team can reproduce the build.
With Hugging Face Transformers, verify that the tokenizer’s special-token IDs match the model configuration and that padding and truncation behave as intended. Test batch encoding, decoding, streaming generation, and long inputs. If you are fine-tuning Llama for Indian regional languages, first confirm whether the base model’s tokenizer already covers Marathi adequately; extending it may be safer than replacing it, but both options require careful retraining and evaluation.
9. Common mistakes to avoid
- Training on scraped text without deduplication or licence review
- Removing punctuation, numbers, or code-mixed text that the product needs
- Applying different Unicode normalisation during inference
- Choosing vocabulary size from convention rather than experiments
- Reporting only vocabulary coverage without downstream metrics
- Replacing a pretrained tokenizer without adapting embeddings
- Ignoring user privacy when building chat or support corpora
- Shipping tokenizer files without versioning or regression tests
For production systems, add tokenizer regression tests to every model release. Include fixed Marathi sentences, mixed-script inputs, malformed Unicode, empty strings, long documents, and special characters.
Final checklist
A production-ready Marathi tokenizer should have a documented corpus, deterministic Unicode handling, tested support for Devanagari and required code-mixed inputs, a justified subword algorithm, and evaluation on held-out and downstream data. Keep the tokenizer and model configuration inseparable in deployment.
Tokenizer work is also part of broader Indic model engineering. If deployment constraints matter, compare your approach with Open-Source Small Language Models for Hindi: A 2026 Guide, especially when deciding between a Marathi-specialised model and a multilingual base model. For local inference and privacy-sensitive applications, How to Deploy Large Language Models Locally covers the next operational layer.
FAQ
Should I use BPE or Unigram for Marathi?
Neither is universally superior. Train both on the same corpus and compare fragmentation, sequence length, unknown-token behaviour, and downstream validation loss.
Should I remove stop words before training?
No. Preserve normal text. Stop-word removal, if needed, belongs to a specific downstream task and should not generally shape the tokenizer vocabulary.
What vocabulary size should I choose?
Benchmark several sizes. Start with 16k–64k depending on corpus scale and whether the tokenizer is Marathi-only or multilingual.
Can I use an existing Hindi tokenizer?
You can use it as a baseline, but Marathi-specific evaluation is essential. Shared Devanagari script does not guarantee efficient Marathi segmentation or adequate coverage.