Telugu and Sanskrit should not be treated as smaller versions of English in an NLP leaderboard. Telugu brings agglutination, rich inflection, dialect variation, informal spelling, and frequent Telugu–English code-switching. Sanskrit adds sandhi, compounds, free word order, dense inflection, and specialised grammatical traditions. A useful benchmark must measure those properties directly—not hide them behind one aggregate score.
This guide presents a practical 2026 workflow for researchers, founders, and public-interest teams building Indic AI. It covers dataset selection, model baselines, evaluation design, error analysis, and reporting standards for both encoder models and generative systems.
Start with the task, not the model
Define the production use case before selecting a leaderboard. A Telugu customer-support assistant, a Sanskrit manuscript search tool, and a speech transcription service need different test sets and success criteria.
Useful benchmark tracks include:
- Text classification: sentiment, topic, intent, toxicity, and language identification.
- Information extraction: named entities, dates, locations, people, organisations, and domain terms.
- Morphosyntax: part-of-speech tagging, morphological features, dependency parsing, and grammatical relation prediction.
- Generation: translation, summarisation, question answering, rewriting, and instruction following.
- Speech: automatic speech recognition, spoken-language understanding, and text-to-speech quality.
- Retrieval: passage ranking, semantic search, and evidence-grounded question answering.
For multimodal or document-heavy projects, pair language evaluation with the guidance in open-source vision-language models for Indian languages. Scanned Sanskrit texts and Telugu forms often fail first at OCR, layout extraction, or script normalisation—not at the language model itself.
Build representative evaluation sets
Public datasets are useful starting points, but they rarely represent deployment conditions. Create a held-out test set that reflects the actual users, domains, and scripts your system will encounter.
For Telugu, include formal news, conversational text, social media, customer-support queries, regional vocabulary, spelling variation, and Telugu–English code-switching. Keep Andhra Pradesh and Telangana examples separate where dialect or terminology affects the task. For Sanskrit, record whether examples come from classical literature, inscriptions, educational material, religious texts, modern prose, or computational grammar resources. A model that performs well on carefully edited prose may still fail on sandhied verses or compound-heavy passages.
Document each example with:
- source, domain, date, and licence;
- script and any transliteration scheme;
- dialect or register, where known;
- annotation guidelines and adjudication process;
- whether the item tests seen or unseen vocabulary;
- personally identifiable or culturally sensitive content controls.
Keep development, test, and challenge sets strictly separated. Near-duplicate contamination is particularly easy when corpora are assembled from Wikipedia, news, religious websites, or parallel translation collections.
Datasets and baseline families
Indic language evaluation can draw on AI4Bharat resources, IndicGLUE-style tasks, LDC-IL collections, Universal Dependencies treebanks, Sanskrit computational grammar resources, and task-specific speech or translation corpora. Check licences and annotation quality before using any dataset commercially; availability in a repository does not automatically mean unrestricted redistribution.
Use at least three baseline families:
- Multilingual encoders: mBERT, XLM-R, and comparable models establish a transferable baseline.
- Indic-focused encoders: IndicBERT and newer Indian-language checkpoints test the value of script-aware and India-focused pretraining.
- Generative models: compare an open multilingual model, an Indic-oriented model where available, and a strong API model under identical prompts and decoding constraints.
For Sanskrit, include grammar-aware or Sanskrit-adapted systems when the task involves parsing, sandhi splitting, or translation. For a deeper adaptation workflow, see fine-tuning large language models for Sanskrit translation. Report the checkpoint, tokenizer, prompt, temperature, context length, quantisation, and decoding settings so results can be reproduced.
Measure tokenisation before measuring accuracy
Tokenisation is a diagnostic, not a footnote. Calculate average tokens per word, characters per token, unknown or fallback rates, sequence truncation, and the proportion of examples exceeding the model context window. Compare these figures by language, domain, and script.
A high token count can increase inference cost and reduce effective context. More importantly, arbitrary fragmentation may make rare inflected forms harder to represent. Test native-script input separately from Latin transliteration and do not assume that transliteration is an improvement: it can discard distinctions, introduce inconsistent spellings, and disadvantage users who write in the original script.
Normalisation must also be explicit. Decide how to handle Unicode variants, punctuation, zero-width characters, numerals, whitespace, diacritics, and Telugu–English mixed text. Apply the same policy to training and evaluation, while retaining a raw-input challenge set to expose brittle preprocessing.
Choose metrics that reflect linguistic behaviour
Use task-appropriate metrics and report confidence intervals where the test set permits them.
- Classification: macro-F1, per-class F1, balanced accuracy, and calibration. Macro-F1 prevents frequent classes from masking minority failures.
- NER and tagging: span-level precision, recall, and F1, with strict and relaxed matching where segmentation is ambiguous.
- Parsing: labelled and unlabelled attachment scores, plus separate results for long-distance dependencies and compound-heavy examples.
- Translation: chrF++ alongside BLEU, COMET or another semantic metric, and human adequacy checks. Character metrics are often informative for rich morphology but are not sufficient alone.
- Summarisation and generation: factuality, coverage, repetition, adequacy, and human preference—not ROUGE alone.
- ASR: word error rate and character error rate, separated by clean speech, accents, noise, and code-switching.
- Retrieval and QA: recall at k, mean reciprocal rank, answer correctness, citation support, and abstention quality.
For Sanskrit, evaluate sandhi splitting and morphological feature accuracy separately from end-task accuracy. A fluent translation can still contain a wrong case, tense, semantic role, or doctrinal interpretation. For Telugu, test spelling variants and colloquial forms rather than evaluating only standard editorial text.
Evaluate generative models safely
Prompt-based scores are sensitive to language, formatting, and examples in the prompt. Create fixed prompt templates in Telugu and Sanskrit, use native-language instructions where relevant, and test zero-shot, few-shot, and retrieval-augmented settings separately. Do not compare a carefully engineered prompt for one model with a minimal prompt for another.
Add adversarial and abstention tests:
- ambiguous words and underspecified questions;
- fabricated citations and invented grammatical explanations;
- unsafe or privacy-sensitive requests;
- code-mixed and transliterated input;
- long passages with distractors;
- culturally sensitive translation and interpretation.
Track unsupported claims, hallucinated sources, script corruption, and inappropriate confidence. When the application handles archival, religious, legal, or public-service material, require evidence-linked answers and human review. Teams deploying voice interfaces can extend this approach with AI call transcript analysis for sales teams, adapting its transcript-quality principles to Telugu speech, code-switching, and domain terminology.
Make the benchmark reproducible
A credible report should publish more than a single score. Include the dataset version, licence, sampling method, preprocessing code, model revision, hardware, latency, memory, and cost per request. Release predictions and error categories where privacy and licensing allow.
Use fixed seeds for conventional experiments, but run multiple trials for generative evaluation. Break results down by script, register, length, morphology, dialect, and domain. Include a small human-evaluated subset with at least two qualified annotators and adjudication for disagreements. For Sanskrit, involve reviewers with relevant grammatical expertise; general fluency is not enough to validate Paninian analysis or classical interpretation.
A practical scorecard can combine:
- task quality;
- robustness on challenge sets;
- latency and memory;
- token and API cost;
- licence and deployment constraints;
- safety, calibration, and abstention;
- performance on underrepresented varieties.
Common mistakes to avoid
- Treating one Telugu news corpus as representative of all Telugu.
- Ranking Sanskrit systems solely by BLEU or surface fluency.
- Reporting accuracy without class balance, confidence intervals, or error examples.
- Mixing transliterated and native-script data without separate results.
- Using test data during prompt development or model selection.
- Claiming that a multilingual model is “Indic” without measuring tokenisation and transfer.
- Ignoring OCR, speech quality, retrieval, and latency in an end-to-end product.
For teams building locally or under data-governance constraints, benchmark deployment options as well as model quality. The trade-offs discussed in how to deploy large language models locally are directly relevant when Telugu or Sanskrit workloads involve sensitive government, education, or archival data.
A practical 2026 benchmark plan
Begin with a frozen, documented test set covering the target users and failure modes. Establish multilingual and Indic-focused baselines, then add a domain-adapted or instruction-tuned model. Measure tokenisation and cost before fine-tuning. Report both aggregate and slice-level metrics, conduct expert review for morphology and meaning, and publish representative failures.
The goal is not to declare one universal winner. It is to identify which system is reliable for a defined Telugu or Sanskrit task, under realistic Indian data, hardware, scripts, and user expectations. That discipline produces benchmarks that builders can trust—and models that are more useful outside the leaderboard.