Telugu instruction-following evaluation needs more than a single accuracy number. A useful benchmark must test whether a model understands Telugu instructions, preserves constraints, responds in the requested format, and avoids inventing information. IndicEval can provide the evaluation structure, while Hugging Face supplies datasets, models, inference utilities, and experiment-sharing tools.
This guide describes a reproducible workflow for running the benchmark in 2026. Because IndicEval package interfaces and task names can change, verify the current repository documentation and dataset card before copying commands. Treat the benchmark definition, prompt template, model revision, and scoring script as part of the experiment—not as implementation details.
Define the Telugu evaluation contract
Start by writing down what counts as a correct response. “Instruction following” is too broad unless the task has explicit acceptance criteria. Your Telugu test set might include:
- Direct execution: answer a question, classify text, extract fields, or perform a transformation.
- Constraint following: obey limits on length, structure, language, tone, or required fields.
- Multi-step instructions: complete several operations in the stated order.
- Grounded response: answer only from supplied context and identify missing information.
- Safety and refusal: decline disallowed requests without becoming irrelevant or switching languages unnecessarily.
Separate language competence from task competence where possible. A response can be grammatically fluent Telugu but still ignore a required constraint. Conversely, a model may complete the task while producing unnatural or mixed-language Telugu. Define these as separate dimensions before evaluation.
For broader test design, compare your plan with this practical framework for benchmarking multilingual LLMs in India. It is particularly useful when Telugu results will be compared with Hindi, Tamil, Kannada, or English.
Obtain and audit the IndicEval data
Use the official IndicEval source, dataset card, or release package rather than an unverified copy. Record the dataset version, commit or revision, license, split names, and any language or task filters. If the current release does not expose a ready-made Telugu instruction-following split, do not silently label a custom collection as IndicEval. Document the adaptation and report it separately.
Before running a model, audit the examples:
- Confirm that instructions are genuinely Telugu, while noting code-mixed and transliterated cases.
- Check that reference answers are valid and not duplicated across train and test splits.
- Identify cultural, regional, and domain coverage gaps.
- Remove prompts that depend on unavailable external facts unless the benchmark explicitly tests retrieval.
- Preserve the original text and store any normalized version in a separate field.
A small manually reviewed sample is worth the time. Telugu spelling variation, punctuation, Unicode normalization, and informal phrasing can create false failures. For additional data sources, consult this guide to open datasets for Telugu language models.
Set up a reproducible Hugging Face run
Create an isolated environment and pin the versions used for the run. A typical starting point is:
pip install -U transformers datasets accelerate evaluate sentencepieceInstall IndicEval according to its current official instructions; the package name and import path should be taken from the release documentation. Log the following for every experiment:
- Model repository, exact revision, tokenizer revision, and license
- Python, PyTorch, Transformers, CUDA, and IndicEval versions
- Hardware, quantization settings, and precision
- Dataset revision and split
- Prompt template, system message, and generation parameters
- Random seed and decoding mode
Load a model using an architecture appropriate to its configuration. Decoder-only instruction models generally use AutoModelForCausalLM; encoder-decoder models use AutoModelForSeq2SeqLM.
from datasets import load_dataset
from transformers import AutoTokenizer, AutoModelForCausalLM
model_id = "your-model-id"
tokenizer = AutoTokenizer.from_pretrained(model_id, revision="your-revision")
model = AutoModelForCausalLM.from_pretrained(
model_id,
revision="your-revision",
device_map="auto",
torch_dtype="auto",
)
data = load_dataset("your-indiceval-dataset", revision="your-dataset-revision")Do not assume that a model marketed as multilingual is strong in Telugu. Include Telugu-capable baselines and, where relevant, a smaller local model. This guide to running a Telugu small language model offline can help when latency, privacy, or GPU availability matters.
Standardise prompting and generation
Prompt formatting can change scores substantially. Use the model’s documented chat template when available, and keep the Telugu instruction unchanged. Avoid translating the prompt into English for one model but not another. If you add an instruction such as “answer only in Telugu,” apply it consistently and disclose it.
Use deterministic decoding for the primary score:
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
outputs = model.generate(
**inputs,
max_new_tokens=256,
do_sample=False,
)
answer = tokenizer.decode(outputs[0], skip_special_tokens=True)For generative tasks, run a separate robustness condition with a fixed seed and controlled temperature. Record input and output token counts, truncation events, stop conditions, and latency. A maximum-token setting that truncates long Telugu answers can look like instruction failure, so report truncation separately.
Score more than exact match
Use the official IndicEval scorer where available, but inspect what it actually measures. Exact match is appropriate for some labels and structured outputs, not for open-ended Telugu responses. A practical scorecard can include:
- Task correctness: matches the reference or satisfies the task rubric.
- Instruction adherence: follows all explicit requirements.
- Constraint accuracy: respects format, length, ordering, and required fields.
- Telugu quality: fluency, spelling, script consistency, and appropriate register.
- Grounding: avoids unsupported claims when context is provided.
- Safety behaviour: refuses or redirects correctly when required.
- Operational metrics: latency, throughput, memory use, and failure rate.
For open-ended outputs, combine automated checks with blinded human review. Use at least two Telugu-proficient reviewers for a representative sample, define a short rubric, and report agreement or adjudication rules. Automated LLM judging can assist triage, but it should not be the only authority—especially when judge models are weaker in Telugu or favour longer answers.
Analyse failures by category
A benchmark becomes useful when it explains what to fix. Save each prompt, raw output, parsed output, score, and error label. Recommended labels include:
- Misread Telugu instruction or ambiguous reference resolution
- Ignored constraint or incomplete multi-step execution
- Hallucination or unsupported detail
- Wrong language, transliteration, or code-switching
- Formatting or parser failure
- Unsafe compliance or inappropriate refusal
- Truncation, timeout, or infrastructure error
Break results down by task type, prompt length, script form, domain, and difficulty. Compare mean scores with worst-case slices; a high aggregate can hide serious failures in government, education, health, or citizen-service prompts. Confidence intervals from bootstrap resampling are useful when the Telugu test set is small.
Publish a defensible benchmark report
A credible report should let another team reproduce the run. Include the dataset and model revisions, prompt template, hardware, decoding settings, scorer version, sample counts, exclusions, and human-evaluation protocol. Publish aggregate and slice-level results, not just a leaderboard number. Redact personal information and review licensing before releasing prompts or outputs.
For model comparisons across Indian languages, use the methods in Benchmarking NLP models for Telugu and Sanskrit as a complementary reference. If your application involves retrieval or legal content, keep those evaluations distinct; this guide to benchmarking LLMs on Indian legal text covers a different risk profile.
Improve the model without contaminating the test set
Use failure examples to refine prompts, retrieval, fine-tuning, or post-processing—but keep the final IndicEval test split untouched. Maintain a development set for iteration and rerun the locked test only at release milestones. Track regressions in Telugu quality when optimising English or other Indian languages.
The strongest workflow is therefore simple: freeze the evaluation contract, version every input, score multiple dimensions, inspect Telugu-specific failures, and report limitations openly. IndicEval and Hugging Face make the mechanics accessible; careful benchmark design determines whether the result is trustworthy.