Kannada small language models need more than a strong headline score. A model can perform well on a clean benchmark yet struggle with code-mixed Kannada-English, spelling variation, dialectal vocabulary, names, numerals, or everyday requests from users in Karnataka. A useful evaluation therefore combines reproducible tests with native-speaker judgment and production monitoring.
This guide is designed for researchers, startups, public-interest teams, and developers comparing compact Kannada models in 2026. It applies to models used for classification, search, summarisation, translation, question answering, chat, and speech-adjacent workflows.
Start with a clear evaluation contract
Define what the model must do before selecting metrics. Record:
- Task: generation, classification, extraction, translation, retrieval, or conversational assistance.
- Users: region, age group, literacy level, and expected Kannada-English code-mixing.
- Operating constraints: parameter count, latency, memory, hardware, context length, and inference cost.
- Risk level: whether an error can affect money, health, benefits, employment, or access to services.
- Comparison set: include a baseline, a larger multilingual model, and a simple non-AI alternative where practical.
For teams working with limited data, the principles in this guide to low-resource Indic NLP are especially relevant. Evaluation data should be treated as a controlled asset: version it, document its licence, and prevent test examples from entering training or prompt-tuning data.
Build a representative Kannada test set
A single Kannada validation split is rarely sufficient. Create several slices so that aggregate results do not hide important failures:
- Standard written Kannada from news, education, government, and general web text.
- Informal messages containing abbreviations, typos, emojis, and Kannada-English mixing.
- Regional and social variation, with examples reviewed across relevant user groups.
- Script variation, including Kannada script, transliterated Kannada in Latin script, and inconsistent transliteration.
- Names, places, dates, currency amounts, units, phone numbers, and government terminology.
- Long and short inputs, including incomplete queries and multi-turn context.
- Adversarial prompts, ambiguous wording, negation, sarcasm, and quoted text.
Use stratified sampling rather than choosing only easy or popular examples. Keep a private holdout set for final comparison. If annotated data is scarce, use active learning to select uncertain or high-impact examples for expert labelling instead of repeatedly adding random samples.
Match metrics to the task
Classification and detection: Report accuracy only when classes are balanced. Add macro-F1, per-class precision and recall, confusion matrices, and calibration. Macro-F1 is important when a model performs well on common categories but fails on minority intents or dialectal forms.
Named-entity recognition and extraction: Use entity-level precision, recall, and F1, with exact and partial-match scores separated. Check whether the model preserves Kannada names, locations, dates, and numerals without silently normalising them incorrectly.
Translation and summarisation: BLEU can support comparison, but it should not be the only measure. Add chrF or similar character-aware metrics for spelling and morphology, plus human ratings for adequacy, fluency, factuality, and omission. For Kannada, reviewers should distinguish a grammatically acceptable paraphrase from a literal but unnatural translation.
Question answering and chat: Measure answer correctness, groundedness, refusal quality, instruction following, and citation or evidence accuracy where applicable. Include an abstention test: a reliable small model should say it does not know rather than inventing an answer.
Language modelling: Perplexity is useful only when tokenisation, corpus composition, and evaluation splits are comparable. Report it alongside downstream task results. A lower score on web text does not automatically mean better performance for customer support, education, or public-service queries.
Efficiency: Track time to first token, tokens per second, peak RAM or VRAM, model size on disk, batch behaviour, and cost per request. For an on-device or low-connectivity deployment, these may matter as much as quality scores.
Use human evaluation that Kannada speakers can trust
Native-speaker review is essential for fluency, meaning, cultural fit, and harmful or embarrassing errors. Build a short rubric with anchored examples. Ask reviewers to score:
- Meaning preservation and factual accuracy.
- Naturalness, grammar, spelling, and register.
- Relevance and completeness.
- Respectful handling of caste, religion, gender, disability, and regional identity.
- Whether code-mixed or transliterated input was understood correctly.
- Whether the output is safe to act on.
Use at least two independent reviewers for high-risk items and adjudicate disagreements. Do not ask annotators to judge content outside their linguistic or domain expertise. Record confidence and the reason for each low score; these comments often reveal tokenisation, terminology, or data-quality problems faster than aggregate metrics.
Test robustness, fairness, and safety
Create paired examples that differ in dialect, spelling, script, politeness, or code-mixing while preserving intent. Compare performance gaps across slices, not just the overall average. Examine whether the model over-refuses harmless Kannada requests, follows unsafe instructions more readily in Kannada than English, or produces different answers for equivalent names and social identities.
For sensitive uses, add tests for privacy leakage, prompt injection, fabricated citations, unsafe medical or financial advice, and translation of harmful content. A compact model may be easier to run privately, but it still needs output filtering, access controls, logging, and a human escalation path.
If your project spans text and images, do not assume a Kannada text score predicts multimodal quality. Evaluate the full workflow separately; teams comparing Indian-language multimodal systems may find this overview of open-source vision-language models for Indian languages useful.
Run a reproducible comparison
For every model, preserve the same prompt templates, decoding settings, context limits, retrieval corpus, and hardware class. Report random seeds where applicable. Use bootstrap confidence intervals or significance tests for close comparisons, and publish per-slice results rather than a single leaderboard number.
A practical evaluation table should include:
- Model and checkpoint version.
- Training or fine-tuning data provenance, where known.
- Tokeniser and context window.
- Task and dataset version.
- Prompt, decoding, and retrieval settings.
- Quality, safety, latency, and memory results.
- Known failure modes and reviewer notes.
When adapting an existing model, compare fine-tuning against prompting and retrieval-augmented generation. The guide to fine-tuning Llama for Indian regional languages can help structure that experiment, but the final choice should follow your Kannada test slices and operating constraints.
Monitor after deployment
Offline evaluation cannot capture changing user language, new names, seasonal topics, or shifts in customer queries. Sample production interactions with privacy safeguards, track unanswered and corrected requests, and maintain a labelled regression set from real failures. Monitor drift by input type, script, dialect proxy, task, and safety category.
Set release gates before deployment—for example, no regression on high-risk slices, a maximum latency budget, and a minimum human-rated adequacy score. Re-run the full suite after changes to the model, tokenizer, prompt, retrieval index, quantisation, or safety layer.
A practical decision rule
Choose the smallest Kannada model that meets the required quality and safety thresholds on representative slices, not the model with the best average benchmark score. If two models are close, prefer the one with clearer provenance, better calibration, lower latency, and easier monitoring. For Hindi comparisons and evaluation ideas, see this practical guide to open-source small language models for Hindi.
The strongest evaluation report makes limitations explicit. It states where the model works, where it fails, who reviewed it, how uncertainty was measured, and what safeguards exist. That transparency is more valuable to Kannada users and builders than an impressive but uninterpretable score.