Why benchmarking Indic small language models needs a different plan
A compact model can look strong on a single leaderboard and still fail in production. Indian-language systems must handle multiple scripts, code-mixing, spelling variation, transliteration, dialect differences, noisy speech transcripts, and uneven training data. A useful benchmark therefore measures not only task accuracy, but also language coverage, robustness, latency, memory use, safety, and cost.
This guide gives builders a repeatable process for comparing small language models in 2026. It is especially relevant for models intended to run on affordable cloud instances, edge devices, call-centre systems, education products, and public-service interfaces.
For background on data scarcity, tokenisation, and transfer learning, start with this builder’s guide to low-resource Indic NLP. It provides useful context before you design an evaluation set.
1. Define the deployment decision first
Do not begin by collecting every available metric. Begin with the decision the benchmark must support:
- Which model should power a Marathi customer-support assistant?
- Can a Hindi-Tamil system run within a mobile memory budget?
- Does a Bengali classifier work on Romanised input from social media?
- Can a multilingual model maintain quality when one language dominates fine-tuning?
- Is a slightly less accurate model preferable because it is faster and cheaper?
Write down the target languages, scripts, user groups, tasks, hardware, response-time limit, and acceptable failure rate. A model for offline government forms has different requirements from one serving real-time voice queries. If voice is involved, evaluate the full speech-to-text and language-model pipeline rather than testing text alone; product teams may also find related guidance in voice agent software for small businesses.
2. Build a representative evaluation set
A benchmark should contain both public datasets and a carefully governed, private test set. Keep the final test set isolated from training and prompt-tuning workflows.
Cover the following dimensions:
- Languages: Include each target language separately. Do not treat “Indic” as one language category.
- Scripts: Test native scripts as well as Romanised text where users commonly type that way.
- Registers: Include formal writing, conversational language, slang, abbreviations, and professional terminology.
- Geography: Sample regional varieties and urban-rural differences when the product serves multiple states.
- Code-mixing: Include realistic Hindi-English, Tamil-English, and other mixed-language inputs rather than artificial word substitution.
- Input noise: Add spelling errors, missing diacritics, OCR artefacts, speech-recognition errors, and inconsistent punctuation.
- Task types: Cover classification, extraction, retrieval, summarisation, translation, question answering, and generation as required by the product.
Record the source, licence, language, script, domain, annotator instructions, and demographic limitations for every item. Remove duplicates and near-duplicates across training, validation, and test splits. Leakage is particularly damaging when small models are compared on narrow public datasets.
3. Select metrics that reflect real use
No single score is sufficient. Report task metrics alongside operational and quality measures.
Task quality
- Classification: Macro-F1, per-class precision and recall, and confusion matrices. Macro-F1 prevents high-resource or majority classes from hiding poor performance.
- Extraction: Span-level precision, recall, and F1, with strict and relaxed matching where spelling variation is expected.
- Translation: chrF, BLEU, and human adequacy and fluency ratings. chrF is often more informative for morphologically rich languages because it evaluates character n-grams.
- Generation and question answering: Factuality, completeness, citation correctness, refusal quality, and human preference—not just lexical overlap.
- Language modelling: Perplexity can help compare a model under controlled conditions, but it should not be presented as a direct measure of usefulness across different tokenisers or languages.
Operational quality
Measure median and tail latency, tokens per second, peak RAM or VRAM, model size, energy use where relevant, and cost per 1,000 requests. Test batch size one, because interactive Indian-language applications often have low or uneven traffic. Record performance on the actual CPU, GPU, or mobile hardware intended for deployment.
Reliability and safety
Track invalid outputs, repetition, hallucination, prompt sensitivity, toxic completion rates, privacy leakage, and unsafe advice. Evaluate refusals in each target language and in code-mixed forms. A model that refuses correctly in English but complies with the same harmful request in Kannada is not robust.
4. Test tokenisation and script behaviour
Small models are highly sensitive to token efficiency. For each language, calculate average tokens per sentence, characters per token, and the proportion of rare or unknown pieces. Compare native-script and Romanised inputs. A model may appear small by parameter count but become slow and expensive when a language requires many more tokens.
Inspect errors involving:
- conjuncts and combining marks;
- punctuation and numerals;
- names and place names;
- inflections and agglutinative forms;
- transliteration variants;
- mixed scripts within the same sentence.
Use the same normalisation policy across models, and publish both raw-input and normalised-input results. Over-normalising can conceal failures that users will experience in the field.
5. Use a reproducible benchmark harness
Create one evaluation pipeline that fixes prompts, decoding parameters, sampling seeds, maximum output length, context limits, and post-processing rules. Version the code, datasets, model checkpoints, tokenisers, hardware, and dependencies. Store per-example predictions—not only aggregate scores—so errors can be audited.
A practical stack can combine Hugging Face Transformers and Datasets with task-specific scoring libraries, a structured experiment tracker, and a lightweight review interface for human annotators. Run at least three seeds for fine-tuned models where feasible, and report mean scores with variance. For generative systems, test deterministic decoding and a defined sampling configuration separately.
Open-source contributors can use Indian open-source AI developer projects and student developers building open-source AI as useful reference points for reproducible collaboration and local tooling.
6. Add human evaluation with Indian-language reviewers
Automatic metrics miss pragmatics, politeness, cultural references, and subtle meaning changes. Use bilingual or native-language reviewers who understand the target domain. Give them clear rubrics and separate dimensions for correctness, relevance, naturalness, harmfulness, and completeness.
Use blinded comparisons where possible. Measure inter-annotator agreement, adjudicate disagreements, and compensate reviewers fairly. Do not translate every output into English and judge it only there; translation can erase precisely the errors the benchmark should detect.
7. Analyse results by slice, not only by average
Publish an overall score only alongside per-language and per-condition breakdowns. Useful slices include script, domain, input length, code-mixing level, dialect, noise type, and safety category. Identify the worst-performing slices and calculate their user impact.
A weighted score can help with model selection, but publish the weights. For example, a customer-support deployment might assign 35% to task quality, 20% to robustness, 20% to safety, 15% to latency, and 10% to cost. Another team should be able to replace those weights and reach a different, transparent decision.
Common benchmarking mistakes
- Comparing models with different prompts or decoding settings.
- Reporting English results as evidence of Indic capability.
- Using BLEU alone for open-ended generation.
- Mixing training data into public test sets.
- Ignoring Romanised and code-mixed input.
- Averaging all languages so high-resource languages dominate.
- Testing only on powerful GPUs instead of deployment hardware.
- Treating a larger model’s score as proof that it is the better product choice.
- Publishing examples without documenting licences, consent, or sensitive data handling.
A practical benchmark report template
Your final report should include:
1. Model name, parameter count, licence, checkpoint, and quantisation method.
2. Languages, scripts, domains, and dataset sources.
3. Preprocessing, prompts, decoding settings, and hardware.
4. Automatic metrics with confidence intervals or variance.
5. Human-evaluation protocol and reviewer details.
6. Latency, memory, throughput, energy, and estimated serving cost.
7. Per-language error analysis and representative failures.
8. Safety results, known limitations, and data-governance notes.
9. Reproduction instructions and access to code or prediction files where permitted.
Final takeaway
The best way to benchmark small language models for Indian languages is to treat benchmarking as a deployment experiment, not a leaderboard exercise. Build language-specific and script-aware test slices, combine automatic and human assessment, measure resource use on real hardware, and make failures visible. That approach helps Indian builders choose models that are not merely compact, but dependable for the communities and workflows they are meant to serve.