What this benchmark should measure
Kannada instruction-following evaluation is not simply a translation test. A useful benchmark checks whether a model understands the Kannada request, performs the requested task, respects constraints, and produces an answer in the expected format. That distinction matters for assistants used in education, public services, customer support, and India-focused applications.
IndicEval can provide a consistent evaluation structure, but treat repository commands, dataset schemas, and metric names as version-dependent. Verify the official IndicEval documentation and commit version before reporting results. For broader context on dataset selection, compare this workflow with the Indian language LLM benchmark datasets guide and the practical framework for benchmarking multilingual LLMs in India.
A strong Kannada evaluation should separate at least four capabilities:
- Instruction comprehension: Did the model identify the requested action?
- Task completion: Is the answer factually and procedurally adequate?
- Constraint adherence: Did it follow limits on length, structure, language, or style?
- Language quality: Is the Kannada natural, readable, and appropriate for the audience?
Prepare a reproducible environment
Use a clean virtual environment and pin the exact versions used for the run. Do not assume that indic-eval is available as a stable PyPI package or that a placeholder GitHub repository is correct. Clone the maintained IndicEval source, inspect its README, and follow its documented installation and evaluation entry point.
python -m venv .venv
source .venv/bin/activate
python -m pip install --upgrade pip
pip install torch transformers datasets accelerate sentencepiece evaluate
# Install IndicEval using the command documented by its maintained repositoryRecord the following in a README or experiment manifest:
- IndicEval repository URL, commit, and dataset version
- Python, PyTorch, Transformers, and CUDA versions
- Model repository and revision on Hugging Face
- Hardware, quantisation settings, and batch size
- Decoding parameters, prompt template, and random seed
- Any language, safety, or output post-processing
This record is essential because generation settings can change scores as much as model weights do. If you are evaluating a production assistant, also document the exact system prompt and whether retrieval, tools, or conversation history are enabled.
Validate and isolate the Kannada data
Use the benchmark’s official Kannada split where available. Never silently translate an English split and label it as native Kannada instruction following: translation can remove cultural, syntactic, and ambiguity patterns that the benchmark is meant to test.
Before running the model, inspect a sample of records manually and validate:
- Required fields, such as instruction, input, reference answer, category, and ID
- UTF-8 encoding and Kannada Unicode integrity
- Duplicate prompts and near-duplicate templates
- Empty fields, malformed examples, and accidental English leakage
- Train-test contamination and repeated answers
- Category balance across summarisation, classification, extraction, reasoning, and safety tasks
Keep the benchmark split read-only. If you clean records, publish the filtering rules and retain the original IDs so results can be audited. Create a small smoke-test subset of 10–20 examples before a full run; this catches schema and generation errors quickly.
Load a Kannada-capable model with Hugging Face
Select the model class that matches the checkpoint. Causal instruction models generally use AutoModelForCausalLM; encoder-decoder checkpoints use AutoModelForSeq2SeqLM. Do not substitute one class merely because it appears in an older tutorial.
import torch
from transformers import AutoTokenizer, AutoModelForCausalLM
model_id = "org/model-name" # replace with a verified Kannada-capable checkpoint
tokenizer = AutoTokenizer.from_pretrained(model_id, revision="main")
model = AutoModelForCausalLM.from_pretrained(
model_id,
revision="main",
torch_dtype=torch.bfloat16 if torch.cuda.is_available() else torch.float32,
device_map="auto",
)
model.eval()Check the model card for Kannada coverage, licensing, supported context length, chat-template requirements, and known limitations. If the checkpoint provides a chat template, use tokenizer.apply_chat_template() rather than manually inventing role markers. Run a baseline with deterministic decoding first:
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
with torch.inference_mode():
output = model.generate(
**inputs,
max_new_tokens=256,
do_sample=False,
pad_token_id=tokenizer.eos_token_id,
)
answer = tokenizer.decode(output[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True)For fair comparisons, keep max_new_tokens, temperature, top-p, stop rules, and prompt format constant. Report separate results for deterministic and sampled settings only when both are justified by the product use case.
Run IndicEval and preserve raw outputs
Adapt the following conceptual workflow to the current IndicEval CLI or Python API. The exact command may differ by release:
# Example structure; use the maintained IndicEval command for your version
python -m indiceval \
--model your_model \
--language kn \
--task instruction_following \
--data data/kannada_test.jsonl \
--output runs/model-kannada-baselineSave every prompt, raw completion, reference, item ID, category, latency, token count, and error status. Do not save only the final aggregate. Raw outputs enable targeted review when a score changes after a model, tokenizer, or prompt update.
If IndicEval expects a model adapter, implement a thin adapter that exposes one predictable generate(prompt) function. Keep batching, truncation, and device placement outside the scoring logic. This separation makes it easier to compare local Transformers inference with an endpoint or quantised deployment.
Score more than exact match
Instruction following needs layered evaluation. Use exact match only for genuinely exact outputs, such as a fixed label or JSON field. For open-ended Kannada answers, combine task-appropriate automated metrics with human review:
- Format validity: JSON parse rate, required fields, length compliance, and forbidden-content checks
- Task metrics: accuracy, macro-F1, ROUGE, or semantic similarity where the task supports them
- Instruction adherence: rubric-based pass rate for all explicit constraints
- Language review: fluency, grammar, terminology, code-switching, and script correctness
- Safety and reliability: refusal correctness, hallucination rate, and consistency across repeated runs
A Kannada evaluator—human or model-based—must be validated rather than assumed to be neutral. For a credible sample, have Kannada-speaking reviewers score a stratified subset using a written rubric. Measure agreement, adjudicate disagreements, and report confidence intervals where practical. Avoid using the same model family as both generator and judge without a human-checked calibration set.
Analyse failures by category
Aggregate scores conceal the most useful engineering information. Build an error table with columns for category, prompt type, failure mode, severity, and suggested fix. Common Kannada failure modes include:
- Correct reasoning but an answer in English instead of Kannada
- Kannada output with excessive English or transliterated terminology
- Ignoring a requested list, table, word limit, or output schema
- Misreading honorifics, negation, dates, quantities, or named entities
- Producing a fluent answer that does not answer the instruction
- Refusing benign requests because safety rules are over-triggered
Compare results against a strong baseline, not just a single headline score. Test prompt-only changes separately from model changes, and use bootstrap confidence intervals or repeated seeds when differences are small. For downstream chatbot work, connect benchmark failures to product tests; the guide on building a Kannada WhatsApp chatbot with a small language model offers a useful deployment-oriented perspective.
Report results responsibly
Publish a compact evaluation card containing the dataset version, Kannada sample count, categories, model revision, prompt template, decoding configuration, hardware, metrics, human-review method, and known exclusions. Include per-category scores and at least a few anonymised failure examples. State whether results are zero-shot, few-shot, fine-tuned, quantised, or retrieval-augmented.
Do not present IndicEval as proof of real-world Kannada competence. Benchmarks are finite and can contain cultural or domain gaps. Supplement them with fresh, consented product data, red-team cases, and task-specific tests. If fine-tuning is the next step, review guidance on fine-tuning a small language model for Kannada support and fine-tuning with AutoTrain on Indian datasets.
A practical 2026 checklist
- Pin code, model, dataset, and tokenizer revisions.
- Verify native Kannada data and inspect Unicode quality.
- Use the checkpoint’s documented chat template.
- Run a smoke test before the complete benchmark.
- Preserve raw prompts and completions.
- Report format, task, language, safety, and human-reviewed metrics.
- Break down results by task and failure type.
- Re-run the same harness after every model or prompt change.
- Share limitations instead of overstating a single score.
A disciplined IndicEval harness turns Kannada instruction-following evaluation into an engineering process: reproducible inputs, controlled generation, transparent scoring, and actionable error analysis. That is the foundation for improving models that Indian users can rely on.