Marathi is a strong candidate for a focused small language model (SLM): it has substantial written content, clear domain needs, and many applications where a compact model is more practical than a general-purpose system. A Marathi SLM can power search, classification, summarisation, autocomplete, customer support, voice interfaces, and education tools while keeping inference costs and latency under control.
The right objective is not to recreate a frontier model. It is to build the smallest model that performs reliably for a defined Marathi use case. That means selecting high-quality data, preserving linguistic variation, measuring errors on real tasks, and designing deployment constraints from the beginning.
1. Define the task and operating constraints
Start with a narrow product requirement. “Understand Marathi” is too broad to guide data collection or evaluation. Specify:
- Task: generation, classification, retrieval, correction, translation, summarisation, or question answering.
- Users: students, government staff, customer-service agents, journalists, or consumers.
- Language mix: Marathi only, Marathi-English code-switching, or Marathi with Hindi and other Indic languages.
- Deployment: cloud API, private server, Android device, browser, or edge hardware.
- Quality and risk: acceptable hallucination rate, response time, privacy requirements, and moderation needs.
For a narrow classifier, a Marathi encoder model may be enough. For text generation, instruction following, or rewriting, begin with an existing open model and adapt it rather than training from zero. The principles in this guide to low-resource Indic NLP are useful when estimating data requirements and handling limited labelled examples.
2. Build a lawful, representative Marathi corpus
Data quality usually matters more than adding another layer or increasing parameter count. Potential sources include openly licensed Marathi books and articles, government publications, educational material, public-domain text, carefully filtered Common Crawl content, and datasets released for Indic-language research.
Create a data register for every source. Record its licence, collection date, domain, language, and permitted use. Do not scrape private social-media content or assume that publicly visible text is automatically suitable for model training. Remove personal information, account credentials, phone numbers, and other sensitive data before training.
Aim for coverage rather than volume alone. A useful corpus should include:
- Formal Marathi from news, policy, education, and reference material.
- Conversational Marathi, including regional and informal usage where legally available.
- Devanagari punctuation, numerals, dates, names, abbreviations, and formatting.
- Marathi-English code-switching and common technical vocabulary.
- Multiple regions and writing styles, without allowing one publisher or domain to dominate.
Keep a held-out test set from sources that are not present in training. Deduplicate near-identical pages and remove boilerplate, navigation menus, broken encoding, spam, and machine-generated repetitions.
3. Normalise Marathi without erasing useful variation
Marathi uses Devanagari, but Unicode text can contain visually similar characters, different normalisation forms, zero-width characters, inconsistent punctuation, and multiple representations of numerals. Apply Unicode normalisation and write explicit rules for whitespace, punctuation, quotes, danda marks, digits, and malformed sequences.
Avoid blindly lowercasing: Devanagari does not have case, while Latin words in code-switched text may carry useful information. Do not remove stopwords for a generative model, and do not perform aggressive stemming or lemmatisation unless the downstream task specifically benefits from it. Function words and inflections help a language model learn syntax.
Run automated checks and sample the results manually. Inspect character-frequency tables, document lengths, duplicate rates, script proportions, and examples of corrupted text. Marathi quality review should involve fluent speakers, especially for dialects, names, respectful forms, and ambiguous spellings.
4. Select a model strategy and tokenizer
There are three practical routes:
1. Train a compact model from scratch when you have a distinctive domain, enough clean Marathi text, and the infrastructure to run repeated experiments.
2. Continue pretraining an existing multilingual or Indic model on Marathi text to improve language coverage.
3. Parameter-efficient fine-tuning for a specific task using LoRA or QLoRA. This is usually the best starting point for a small team.
For a first prototype, a decoder model in the roughly 100-million-to-1-billion-parameter range can be easier to operate than a larger checkpoint. Choose based on licence, Marathi coverage, tokenizer efficiency, context length, and hardware—not parameter count alone. The open-source small language model guide for Hindi offers a useful comparison point, but Marathi needs its own tests.
Tokenizer quality is especially important. Measure how many tokens are required for representative Marathi sentences, names, inflected words, and code-switched text. A tokenizer that fragments common Marathi words excessively increases memory use and weakens learning. You can reuse a multilingual tokenizer, extend its vocabulary, or train a new SentencePiece or BPE tokenizer; compare these choices on compression, training stability, and downstream accuracy.
5. Train efficiently and reproducibly
Prepare train, validation, and test splits by document or source, not by randomly splitting adjacent sentences. Otherwise, duplicated passages can make results look better than they are. Use packed sequences for causal language modelling, mask padding correctly, and track the number of tokens—not just epochs.
For a practical training run, monitor:
- Training and validation loss or perplexity.
- Learning-rate schedule, gradient norms, and GPU memory.
- Evaluation results by domain, script pattern, and code-switching level.
- Checkpoint recovery and experiment configuration.
Mixed-precision training, gradient accumulation, gradient checkpointing, and distributed training can reduce hardware requirements. Keep configurations, dataset hashes, tokenizer versions, and random seeds under version control. For most Indian builders, rented cloud GPUs may be suitable for experiments, while a smaller quantised model can serve production traffic economically.
Do not treat perplexity as the final product metric. A model can achieve lower loss while still producing poor Marathi, repeating text, or inventing facts. If you are adapting a Llama-family model, compare the trade-offs in this guide to fine-tuning Llama for Indian regional languages.
6. Evaluate Marathi performance with real tests
Build a Marathi evaluation set before tuning the model. Include natural prompts from the intended product and label the expected answer, acceptable alternatives, and failure severity. Test:
- Grammar, spelling, agreement, and fluency.
- Factuality and resistance to unsupported claims.
- Named entities, dates, numbers, and quotations.
- Marathi-English code-switching and transliterated input.
- Regional vocabulary and respectful or formal language.
- Safety, privacy, abusive content, and prompt-injection behaviour.
Use perplexity for language modelling, but use task-specific metrics for the application. Classification may require macro-F1; translation can use chrF alongside human review; summarisation needs factuality and coverage checks. Human evaluation by Marathi speakers remains essential. Ask reviewers to score meaning preservation, naturalness, usefulness, and harmful or misleading output.
Compare against simple baselines: a larger multilingual model, a retrieval system, a rules-based pipeline, and the unadapted base model. This shows whether the SLM creates genuine value rather than merely changing output style.
7. Deploy with privacy, latency, and cost in mind
Quantisation to 8-bit or 4-bit weights can reduce memory and improve serving economics, but test Marathi quality after compression. Batch requests where latency permits, cap output length, cache repeated prompts, and expose confidence or abstention behaviour for high-risk workflows. For mobile or offline use, review AI model optimisation for mobile devices, including memory limits, on-device runtimes, and battery impact.
A production system should log anonymised performance signals, not user secrets. Add rate limits, content filtering, rollback-ready model versions, and a human escalation path. Retrieval-augmented generation can improve factual answers by grounding the model in approved Marathi documents; it is often safer than trying to memorise every policy or knowledge-base article during training.
8. A realistic Marathi SLM project plan
A lean project can proceed in four stages:
- Week 1–2: define use cases, licences, evaluation criteria, and a small representative corpus.
- Week 3–4: clean data, test tokenizers, establish baselines, and create a Marathi reviewer panel.
- Month 2: run continued pretraining or LoRA fine-tuning, then evaluate by domain and failure type.
- Month 3: quantise, pilot with users, monitor errors, and publish a model card with limitations.
Document what the model does not support. Include training sources, known dialect gaps, sensitive-data controls, benchmark results, licence terms, and the date of the last evaluation. That transparency is important for schools, public services, and Indian businesses adopting the model.
FAQ
Can I build a Marathi model with limited data?
Yes. Start with a narrow task, use transfer learning, and invest in high-quality Marathi validation data rather than collecting unfiltered text.
Should I train from scratch?
Usually not for a first release. Continued pretraining or parameter-efficient fine-tuning is cheaper and provides a stronger baseline.
Is Marathi-English code-switching important?
For most real-world products, yes. Measure it explicitly instead of treating it as noise.
What is the best model size?
The best size is the smallest model that meets your accuracy, latency, privacy, and cost targets. Benchmark several sizes on your actual workload.
Where can Indian founders seek support?
Teams building Marathi or other Indic-language products can explore AI Grants India for relevant funding opportunities and ecosystem support.