Kannada model performance should not be judged by a single accuracy number or a few impressive outputs. A useful benchmark shows whether fine-tuning improves the target task, preserves general language ability, and performs reliably across scripts, dialects, domains, and input lengths.
This guide explains how to benchmark a Kannada model before and after fine-tuning on Hugging Face. The workflow applies to text classification, intent detection, named-entity recognition, question answering, translation, and instruction-tuned generation. It also gives you a reproducible comparison that can support a production decision, research report, or grant application.
Define the evaluation contract first
Write down the model’s intended job before choosing metrics. A Kannada customer-support classifier, a translation system, and a conversational model require different tests. Record:
- Task: classification, generation, translation, extraction, or ranking.
- Users and domain: public services, education, healthcare, agriculture, finance, or general chat.
- Input forms: Kannada script, code-mixed Kannada-English, Romanised Kannada, abbreviations, and spelling variation.
- Success criteria: for example, macro-F1 above a threshold, lower translation error, or fewer unsafe responses.
- Constraints: latency, model size, inference cost, licence, and privacy requirements.
For a broader fine-tuning workflow, review these best practices for fine-tuning LLMs on custom data. The benchmark should measure the actual product requirement, not an easier proxy.
Build a clean Kannada evaluation set
Use a fixed, held-out test set that is never used for training, prompt selection, or repeated manual tuning. Split data by document, user, or conversation rather than randomly splitting near-duplicate sentences. This prevents leakage from making the post-training score look better than real-world performance.
Your dataset should include:
- Kannada written in the Kannada script, with UTF-8 normalisation.
- Common spelling variants, punctuation, numerals, and borrowed English terms.
- Code-mixed and Romanised examples if your users produce them.
- Regional, formal, colloquial, and domain-specific language.
- A balanced representation of labels, including rare but important classes.
- A small, manually reviewed challenge set for ambiguity, negation, names, and long inputs.
Use a train, validation, and test split such as 80/10/10 only when the data supports it. Keep the test labels private from the training pipeline. Track dataset version, annotation instructions, annotator agreement, licensing, and personally identifiable information. For translation or generation, ensure references are written by fluent Kannada reviewers rather than produced only by another model.
Benchmark the base model on Hugging Face
Load the exact checkpoint, tokenizer, and revision that you intend to compare. Pin package versions and record hardware, sequence length, batch size, decoding parameters, and random seeds. A benchmark is meaningful only when pre- and post-fine-tuning runs use identical conditions.
For a classification model, use AutoTokenizer and AutoModelForSequenceClassification. Many base multilingual checkpoints have no task-specific classification head, so an untrained head can produce arbitrary predictions. In that case, compare a zero-shot or prompted baseline, a suitable pretrained task checkpoint, or a linear probe—not an uninitialised head presented as a meaningful baseline.
A minimal evaluation pattern is:
from transformers import pipeline
from datasets import load_dataset
model_id = "your-org/your-kannada-checkpoint"
classifier = pipeline(
"text-classification",
model=model_id,
tokenizer=model_id,
device=0,
)
predictions = classifier(dataset["text"], truncation=True)For production benchmarking, prefer Trainer.evaluate() or a batched inference loop and save predictions, probabilities, labels, and example IDs. Use the same test file for both checkpoints. Do not call model.predict() directly on a Transformers model; that method belongs to some high-level libraries, not standard PyTorch Transformers models.
Select metrics that expose Kannada performance
Report aggregate results and class-level results. Accuracy can hide failure on minority labels, so include:
- Macro-F1: gives every class equal weight and is useful for imbalanced Kannada datasets.
- Weighted-F1: reflects the observed class distribution.
- Per-class precision and recall: identifies labels that improved or regressed.
- Confusion matrix: reveals systematic confusion between intents or entities.
- Exact match and token-level F1: useful for extractive question answering.
- SacreBLEU, chrF, and COMET: useful for translation, with human review for Kannada fluency and adequacy.
- ROUGE or BERTScore: useful for summarisation, but never as the only generation metric.
- Validity, citation, refusal, and safety rates: important for instruction-following systems.
For generative models, evaluate multiple outputs with fixed decoding settings and a separate human rubric. Score factuality, relevance, grammaticality, script correctness, harmful content, and code-mixing appropriateness. When possible, use blind pairwise comparisons so reviewers do not know which output is pre- or post-fine-tuning.
Fine-tune with reproducibility controls
Save the baseline before training. Then log the dataset revision, base model revision, training method, learning rate, effective batch size, epochs, warm-up, maximum sequence length, seed, hardware, and checkpoint-selection rule. Parameter-efficient methods such as LoRA can reduce memory requirements, but the adapter configuration and merge procedure must also be recorded.
Use validation performance for early stopping or checkpoint selection. Do not select a checkpoint by repeatedly checking the test set. If you are adapting a regional-language LLM, compare full fine-tuning, LoRA, or instruction tuning against the same Kannada test suite; approaches discussed in fine-tuning Llama for Indian regional languages can help with model and data planning.
Run the post-fine-tuning comparison
Evaluate the original and fine-tuned checkpoints through the same script. Produce a table like this:
| Model | Macro-F1 | Weighted-F1 | Kannada-script accuracy | Code-mixed F1 | Latency |
|---|---:|---:|---:|---:|---:|
| Base checkpoint | | | | | |
| Fine-tuned checkpoint | | | | | |
| Difference | | | | | |
Include confidence intervals or bootstrap estimates where the test set is large enough. A two-point gain may not matter if uncertainty is high. Also measure latency, peak memory, throughput, checkpoint size, and token usage: a modest quality gain may not justify a large deployment cost. If the model will run on constrained hardware, pair quality testing with AI model optimisation for mobile devices.
Inspect errors, not just scores
Export examples where the models disagree and group failures by cause:
- Script or Unicode normalisation problems.
- Romanisation and code-mixing.
- Dialect, spelling, or morphology.
- Long-context truncation.
- Rare labels and class imbalance.
- Names, numbers, dates, and place names.
- Negation, sarcasm, or ambiguous wording.
- Hallucination, refusal, or unsafe completion.
Compare improvements on the challenge set separately from the random test set. A model that gains overall F1 but fails on government-service terminology or healthcare queries may not be ready for deployment. For multilingual systems, broader open-source vision-language models for Indian languages may be relevant only when your application includes images; do not substitute a multimodal benchmark for a text benchmark.
Publish a reproducible benchmark report
Release the evaluation script, dataset card, label definitions, model revisions, preprocessing rules, and limitations. Report whether examples contain personal data, how consent and redaction were handled, and which Kannada varieties are underrepresented. Store predictions securely and avoid publishing sensitive user text.
A strong conclusion states what improved, for whom, under which conditions, and at what cost. If fine-tuning raises macro-F1 while reducing performance on code-mixed inputs, report both facts. This makes the result useful to builders and credible to reviewers.
FAQ
Should I use accuracy for a Kannada classifier?
Use it only alongside macro-F1, per-class recall, and a confusion matrix. Macro-F1 is usually more informative when labels are imbalanced.
Can I use the same test set during training?
No. Keep the test set untouched until the final comparison. Use a validation set for model selection and hyperparameter decisions.
How do I benchmark a Kannada generative model?
Combine automatic metrics with a fixed prompt suite, blind human review, error categories, safety checks, and latency measurements.
What if the base model has no classification head?
Use a task-specific checkpoint, zero-shot prompting, or a trained linear probe. Do not treat random head outputs as a valid baseline.
Where should I share the result?
A Hugging Face model card and dataset card should document the model, data, metrics, limitations, and intended use. Teams building Indian-language systems can also explore support through AI Grants India.