Benchmarking a Telugu model is not simply a matter of running evaluate.load("accuracy") on a held-out split. A useful benchmark must reflect how Telugu is written and spoken across India: formal and conversational registers, code-mixing with English, spelling variation, transliterated Telugu, dialect differences, and domain-specific vocabulary. It should also produce results that another team can reproduce.
Hugging Face provides the building blocks for this process—Transformers, Datasets, Evaluate, the Hub, and model cards. However, Model Card for Performance (MCP) is best treated as a reporting framework, not a benchmark or a single metric. Your evaluation set and methodology determine whether the reported score is meaningful.
1. Define the task and deployment scenario
Start by writing down what the fine-tuned model is expected to do. The correct benchmark depends on the task:
- Text classification: accuracy, macro F1, per-class precision and recall.
- Named-entity recognition: entity-level precision, recall, and F1.
- Translation: chrF, BLEU, COMET where suitable, plus human review.
- Summarisation or generation: ROUGE can be useful, but factuality and human preference matter more than one overlap score.
- Question answering: exact match and token-level F1, with manual checks for acceptable alternative answers.
- Instruction following: rubric-based human evaluation, safety checks, and task success rate.
Also record the production constraints. A model used in a Telugu education application may need stronger performance on children’s language and code-mixed questions. A public-service chatbot may prioritise factuality, refusal behaviour, and robustness to spelling errors. For a mobile deployment, latency and memory should be measured alongside quality; see this guide to AI model optimisation for mobile devices.
2. Build a leakage-resistant Telugu test set
Use a test set that was not used for training, instruction tuning, prompt creation, or repeated manual debugging. Deduplicate at the document, paragraph, and near-duplicate level. Random splitting alone can inflate results when multiple examples come from the same source.
A practical Telugu evaluation set should cover:
- Telugu script, Romanised Telugu, and mixed Telugu-English text.
- Formal news, government language, education, commerce, social media, and conversational prompts.
- Regional vocabulary and common spelling variants.
- Short queries, long passages, noisy text, punctuation variation, and numerals.
- Names, places, dates, currency, measurements, and technical terms.
- Ambiguous prompts and examples that require the model to say it does not know.
Keep a frozen gold test set and create a separate development set for iteration. Store source, licence, annotator instructions, labels, and sensitive-data handling decisions. For Indian-language work, document whether the text was originally written in Telugu or translated from another language; translated test sets can measure translation artefacts rather than natural Telugu ability.
If you are still preparing the training pipeline, review these best practices for fine-tuning LLMs on custom data before fixing your evaluation protocol.
3. Load the model and dataset correctly
For a sequence-classification model, a basic setup looks like this:
from datasets import load_dataset
from transformers import AutoTokenizer, AutoModelForSequenceClassification
model_id = "org/telugu-model"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForSequenceClassification.from_pretrained(model_id)
test = load_dataset("org/telugu-eval", split="test")Check the model card for the expected language, task, label mapping, maximum sequence length, and tokenizer. Confirm that id2label matches your evaluation labels. Silent label mismatches can produce plausible-looking but invalid scores.
Tokenise in batches rather than evaluating one example at a time:
def tokenize(batch):
return tokenizer(
batch["text"],
truncation=True,
max_length=512,
)
test_tokens = test.map(tokenize, batched=True)For generation, set and report decoding parameters such as temperature, top-p, maximum new tokens, repetition penalty, and number of beams. Run deterministic evaluation with a fixed seed when comparing model versions.
4. Select metrics that expose weaknesses
Report aggregate scores, but do not stop there. If classes are imbalanced, macro F1 is usually more informative than accuracy. Include a confusion matrix and per-class results. For NER, calculate scores by entity type—poor performance on locations or personal names may be hidden by a strong overall score.
For generative Telugu tasks, combine automatic and human evaluation:
- Factuality: Does the output preserve the source facts?
- Completeness: Did it answer all required parts?
- Fluency: Is the Telugu grammatical and natural?
- Terminology: Are technical and proper nouns handled consistently?
- Instruction adherence: Did the model follow format and length requirements?
- Safety: Does it avoid harmful, discriminatory, or privacy-invasive responses?
Use a small, carefully designed human-evaluation sample rather than asking reviewers to score everything. Provide Telugu instructions, a clear rubric, anonymised model outputs, and at least two reviewers for a subset. Report agreement and adjudication rules.
5. Run reproducible evaluation
Use the Hugging Face pipeline, Trainer, or a custom PyTorch loop depending on the task. The important requirement is consistent preprocessing and output collection. Save:
- Model repository and exact commit or revision.
- Dataset repository, split, and revision.
- Python, Transformers, Datasets, Evaluate, and hardware versions.
- Random seeds and decoding parameters.
- Prompt templates and system instructions.
- Runtime, peak memory, throughput, and failure counts.
A classification example using evaluate might look like this:
import evaluate
import numpy as np
from transformers import pipeline
accuracy = evaluate.load("accuracy")
classifier = pipeline(
"text-classification",
model=model_id,
tokenizer=tokenizer,
device=0,
)
predictions = classifier(test["text"], batch_size=16, truncation=True)
predicted = [int(item["label"].split("_")[-1]) for item in predictions]
result = accuracy.compute(
predictions=predicted,
references=test["label"],
)
print(result)Adapt label parsing to your model’s actual label names, and add macro F1 with a trusted implementation. For large test sets, write predictions and metadata to a versioned file so that error analysis does not require rerunning inference.
6. Analyse errors, not just scores
Create an error table containing the input, expected output, prediction, category, and reviewer comment. Group failures by linguistic and product dimensions:
- Script versus transliteration.
- Code-mixed versus Telugu-only input.
- Short versus long context.
- Domain and dialect.
- Named entities and numerals.
- Negation, ambiguity, and implicit context.
- Hallucination, refusal, toxicity, or privacy failures.
Compare the base model with the fine-tuned model on the same frozen set. A higher average score can still conceal regressions in safety, long-context handling, or minority varieties. Avoid tuning repeatedly on the test set; use development results for decisions and reserve the test set for final reports.
7. Publish a useful Hugging Face model card
Your MCP-style performance section should make the result auditable. Include the task, intended use, dataset composition, sample counts, language varieties, preprocessing, metrics, confidence intervals where practical, baseline comparisons, known limitations, and ethical considerations. State clearly whether the test data is public, licensed, synthetic, translated, or human annotated.
Include a results table with columns such as:
| Slice | Examples | Metric | Score | Notes |
|---|---:|---|---:|---|
| Telugu script | | Macro F1 | | |
| Romanised Telugu | | Macro F1 | | |
| Telugu-English mixed | | Macro F1 | | |
| Long inputs | | Macro F1 | | |
Document hardware and latency separately from quality. If the model will run on-premises or locally, the deployment approach covered in how to deploy large language models locally can help structure those measurements.
8. A practical release checklist
Before publishing or deploying the model, verify that you have:
- A frozen, representative Telugu test set with documented provenance.
- No train-test contamination or repeated test-set tuning.
- Task-appropriate aggregate and slice metrics.
- Human review for generation, factuality, and safety.
- Baseline and previous-version comparisons.
- Reproducible code, configuration, revisions, and seeds.
- Latency, memory, and failure-rate measurements.
- A model card that states limitations and intended use.
For Telugu and other Indian languages, coverage is part of performance. A benchmark that reports one impressive number but excludes transliteration, code-mixing, regional usage, or real deployment conditions is incomplete. Treat the model card as a living record, rerun the same suite after each major training or data change, and add new slices only with clear versioning.