Marathi small language models should be evaluated as products, not just leaderboard entries. A model can achieve strong next-token scores yet fail on Devanagari spelling, code-mixed Marathi, regional vocabulary, or the practical constraints of an Indian deployment. The right evaluation plan combines automated benchmarks, carefully sampled Marathi test sets, native-speaker review, and production measurements.
This guide explains how to evaluate Marathi small language models for chat, classification, summarisation, translation, search, voice interfaces, and other applications. It is designed for builders working with limited compute, limited labelled data, and real users across Maharashtra and Marathi-speaking communities.
Start with the intended use case
Do not begin by selecting a metric. Begin by defining what the model must do, for whom, and under what constraints. A customer-support model, a speech assistant, and a document classifier need different tests.
Write an evaluation brief covering:
- Tasks: generation, classification, extraction, translation, summarisation, retrieval, or tool use.
- Users: native Marathi speakers, learners, government staff, customers, or internal teams.
- Input formats: clean Devanagari, Roman Marathi, Marathi-English code-mixing, scanned text, or speech transcripts.
- Risk level: low-risk drafting versus health, finance, education, or public-service information.
- Deployment limits: latency, memory, offline operation, per-request cost, and device type.
For background on data scarcity, tokenisation, and transfer learning, see this practical guide to low-resource Indic natural language processing. The same principles apply to Marathi, but Marathi-specific evaluation must account for its grammar, vocabulary, orthography, and regional variation.
Build a representative Marathi test set
A test set should mirror the traffic the model will receive. A single Marathi Wikipedia split is not enough. Create separate, versioned subsets so that a high aggregate score does not hide a serious weakness.
Useful categories include:
- Standard written Marathi: news, educational material, public notices, and formal correspondence.
- Conversational Marathi: short queries, incomplete sentences, local expressions, and informal spelling.
- Code-mixed text: Marathi combined with English, Hindi, numerals, product names, and Latin-script words.
- Roman Marathi: Marathi written in the Latin alphabet, including inconsistent transliteration.
- Regional language: vocabulary and phrasing from Vidarbha, Marathwada, Konkan, western Maharashtra, and urban centres.
- Domain language: agriculture, banking, healthcare, law, education, retail, and government services.
- Adversarial inputs: misspellings, repeated characters, ambiguous words, prompt injection, and misleading context.
Keep training, development, and test examples strictly separated. Near-duplicate news articles and templated government documents can otherwise produce inflated results. Record the source, date, licence, dialect or region where known, annotation method, and personally identifiable information handling for every evaluation item.
For a serious benchmark, use at least 200-500 examples per important task where feasible, with a larger set for simple classification. For generative tasks, maintain a smaller, high-quality challenge set reviewed by Marathi experts. Freeze the test set before model comparison and publish its version with every result.
Measure task performance with the right metrics
Generation and language modelling
Perplexity is useful for comparing models evaluated on the same corpus and tokeniser. It is not a direct measure of usefulness, factuality, or fluency across different tokenisation schemes. Report the corpus composition and tokeniser details alongside the score.
For open-ended generation, measure:
- Instruction-following and task completion rate
- Factual accuracy against a checked answer or source
- Marathi grammaticality and naturalness
- Repetition, truncation, and refusal errors
- Unsupported claims and citation quality
- Robustness to spelling and code-mixing variation
Use exact match only where the answer has a clearly defined form, such as a label or extracted field. For flexible answers, combine rubric-based human scoring with targeted automated checks.
Classification and extraction
Report accuracy only when classes are balanced. In most practical datasets, include precision, recall, macro-F1, and per-class results. Macro-F1 is especially important when minority intents or safety categories matter.
For named-entity recognition and information extraction, report entity-level precision, recall, and F1 rather than token accuracy alone. Check errors involving Marathi inflections, names, locations, abbreviations, and mixed scripts. Calibrate confidence scores before using them to route cases automatically.
Translation and summarisation
BLEU, chrF, and COMET can help compare Marathi translation systems, but no single metric captures meaning, register, or cultural appropriateness. For summarisation, ROUGE is a useful overlap signal, not a quality verdict. Review whether the summary preserves names, numbers, negation, attribution, and the central claim.
A native-speaker rubric should score meaning preservation, fluency, terminology, completeness, and harmful distortion on a consistent scale. Use blind reviews so evaluators do not know which model produced an answer.
Add Marathi-specific linguistic checks
Automated scores often miss the errors users notice first. Build targeted checks for:
- Devanagari rendering, Unicode normalisation, punctuation, and numeral handling
- Matras, conjuncts, spelling variants, and word-boundary errors
- Gender, number, case, tense, and agreement
- Honorifics and formal versus informal register
- Marathi-English and Marathi-Hindi code-switching
- Transliteration between Devanagari and Roman Marathi
- Named entities, place names, agricultural terms, and government terminology
Ask reviewers to label the severity of each error: cosmetic, inconvenient, misleading, or harmful. This produces a more actionable score than a single preference vote. At least two independent native Marathi reviewers should assess high-impact outputs; adjudicate disagreements and report inter-rater agreement where possible.
If your system accepts or produces images, documents, or speech, do not reuse a text-only benchmark. Pair the language evaluation with modality-specific tests. For context on Indian-language multimodal systems, compare the considerations in open-source vision-language models for Indian languages.
Test safety, fairness, and factuality
Safety evaluation must use Marathi prompts, not translations alone. Translation can remove ambiguity, politeness, slang, or culturally specific ways of requesting harmful content. Test direct requests, indirect requests, role-play, code-mixed prompts, transliterated prompts, and misspellings.
Check for:
- Unsafe medical, financial, legal, or self-harm advice
- Stereotypes involving caste, religion, gender, region, occupation, or disability
- Exposure of personal information from prompts or training data
- Overconfident answers when the model lacks evidence
- Unequal performance across dialects, scripts, and demographic contexts
- Prompt-injection resistance when connected to documents or tools
For factual applications, use a retrieval-backed test set with cited source passages. Score both answer correctness and whether the model declines when the source does not support an answer. A smaller model that states uncertainty accurately may be more suitable than a larger model with higher fluency but worse reliability.
Benchmark deployment, not just quality
A small model earns its place through the full quality-cost trade-off. Record results at the exact quantisation, context length, hardware, and serving stack you plan to use.
Track:
- First-token and full-response latency
- Throughput under concurrent load
- Peak RAM or VRAM usage
- Model size and download time
- Energy or inference cost where measurable
- Failure rate, timeout rate, and recovery behaviour
- Quality loss after quantisation or pruning
Evaluate on the target device in India if offline or edge deployment matters. Test low-bandwidth conditions and intermittent connectivity. For voice products, include transcription errors and turn-taking performance; a language model may appear strong on clean text but fail on noisy Marathi speech.
Compare models with a decision matrix
Use a fixed prompt set and identical decoding settings. Compare against at least one simple baseline, such as a rules-based system, a multilingual model, or an existing production model. If you are adapting a general model, document the data and method rather than reporting only the final score; guidance on fine-tuning Llama for Indian regional languages can help structure that process.
A practical weighted matrix might assign:
- 30% task success and factuality
- 20% Marathi fluency and linguistic correctness
- 15% robustness across scripts, dialects, and domains
- 15% safety and fairness
- 10% latency and reliability
- 10% infrastructure cost
Adjust the weights to the product. A public-health assistant should prioritise safety and factuality; an offline keyboard may prioritise latency and correction quality. Keep raw results available so stakeholders can challenge the weighting rather than treating the final score as objective truth.
A repeatable 2026 evaluation workflow
1. Define tasks, users, risk boundaries, and deployment hardware.
2. Assemble and document representative Marathi data.
3. Freeze a clean test set and a separate challenge set.
4. Run automated metrics with confidence intervals where practical.
5. Conduct blinded native-speaker review using a written rubric.
6. Test dialect, script, code-mixing, safety, and factuality failure modes.
7. Benchmark latency, memory, cost, and quantisation effects.
8. Analyse errors by category, not only by average score.
9. Pilot with consenting users and log failures securely.
10. Set release thresholds and schedule regression tests for every model or prompt change.
Evaluation is an ongoing operating process. Marathi usage changes, product domains expand, and fine-tuning can introduce regressions. Maintain a living error set from real failures, remove sensitive data, and rerun it before every release. Builders exploring the broader small-model landscape can also compare open-source small language models for Hindi, while remembering that Hindi results are not a substitute for Marathi evidence.
A Marathi small language model is ready for deployment when it meets clearly defined task, safety, linguistic, and operational thresholds—not merely when it produces plausible text. This disciplined approach helps teams choose smaller, affordable models without compromising the experience or trust of Marathi-speaking users.