0tokens

Apply for AI Grants India

Financial support for innovators building the future of AI in India.

Apply now

Chat · how to create a telugu benchmark dataset on hugging face

How to Create a Telugu Benchmark Dataset on Hugging Face

  1. aigi

    Telugu benchmark datasets are useful only when they measure real language capability rather than reward memorisation, script shortcuts, or duplicated web text. This guide explains how to create a Telugu benchmark dataset on Hugging Face that researchers can reproduce, models can be compared against, and Indian AI teams can safely reuse.

    A good benchmark should answer one narrow question clearly: Which Telugu capability are you measuring, under what conditions, and with what evidence? Start with that question before collecting data.

    1. Define the benchmark task

    Choose one primary task for the first release. Common options include:

    • Text classification: topic, intent, toxicity, misinformation, or sentiment.
    • Named entity recognition: people, organisations, locations, dates, and schemes.
    • Question answering: extractive or answer-generation tasks grounded in Telugu passages.
    • Translation: Telugu paired with English or another Indian language.
    • Summarisation: news, government information, or long-form public documents.
    • Instruction following: prompts and expected responses, with careful human evaluation.

    Write a short task specification containing the input, expected output, label definitions, exclusions, and evaluation metric. For example, a sentiment benchmark must state how annotators handle sarcasm, mixed Telugu-English text, emojis, and neutral statements. If your goal is broader multilingual evaluation, compare design choices with this Indian language LLM benchmark guide.

    Avoid combining unrelated tasks in one score. Publish separate configurations or subsets so a model’s strengths and weaknesses remain visible.

    2. Source Telugu data responsibly

    Use sources with clear permission and document provenance for every record. Suitable sources may include:

    • Public-domain or openly licensed government material.
    • News and educational content whose licence permits redistribution.
    • Existing open datasets, subject to their original terms.
    • Consent-based contributions from Telugu speakers.
    • Synthetic examples used only when clearly labelled and separately evaluated.

    Do not scrape private posts, copyrighted pages for redistribution, or personal information without a lawful basis. Remove phone numbers, email addresses, exact addresses, Aadhaar-like identifiers, and other sensitive data. Keep a source field internally, even if it is excluded from the public release.

    Telugu data needs more than a language tag. Record script, region where relevant, domain, date, source type, and whether the text contains code-switching. Include natural variation from Andhra Pradesh and Telangana without treating regional usage as an error. This is especially important when building low-resource language datasets for AI training in India.

    3. Build an annotation guide

    A benchmark becomes dependable when two trained annotators are likely to make the same decision. Create an annotation manual with:

    • Definitions for every label.
    • Positive and negative examples in Telugu.
    • Rules for spelling variation, punctuation, transliteration, and code-mixing.
    • Guidance for ambiguous, offensive, political, or culturally sensitive content.
    • An “uncertain” or “cannot determine” option where appropriate.

    Run a pilot with a small sample before full annotation. Measure agreement using Cohen’s kappa, Krippendorff’s alpha, or task-specific agreement. Review disagreements rather than hiding them; they often reveal unclear labels or genuine linguistic ambiguity.

    For generative tasks, collect reference answers from multiple annotators when feasible. One reference can make a correct but differently worded Telugu answer appear wrong. Human evaluation should assess factuality, relevance, fluency, and harmful content separately.

    4. Clean and structure the records

    Do not apply English-centric preprocessing blindly. Telugu text should generally retain its script, diacritics, punctuation, and meaningful spacing. Normalise only when the rule is documented and reversible.

    A practical record might look like this:

    {
      "id": "telugu_sentiment_000001",
      "text": "ఈ సేవ చాలా వేగంగా ఉంది.",
      "label": "positive",
      "domain": "public_services",
      "source_type": "consented_contribution",
      "language": "te"
    }

    Use stable IDs and keep annotation metadata separate from the text where privacy requires it. Check for duplicate and near-duplicate examples, Unicode inconsistencies, empty fields, invalid labels, and accidental train-test overlap. Deduplicate by normalised text, but inspect collisions manually so legitimate repeated phrases are not removed automatically.

    If the dataset supports model training as well as evaluation, publish a clearly marked training split and protect the test set from casual contamination. For broader dataset planning, see this guide to open-source AI datasets for India.

    5. Create defensible train, validation, and test splits

    Random splitting is often insufficient. Near-identical headlines, translated copies, author-specific patterns, or source templates can leak across splits and inflate scores.

    Use a split strategy suited to the task:

    • Group split: keep documents from the same source, author, or conversation together.
    • Time split: train on earlier material and test on later material to measure robustness over time.
    • Domain split: test whether a model transfers from news to public services or education.
    • Challenge split: isolate code-mixed, colloquial, long-context, or regionally varied examples.

    Publish the split-generation script and fixed random seed. Report the number of examples, label distribution, average length, script or language mix, and source composition for every split. Never tune a model repeatedly on the hidden test set; maintain a development set for iteration.

    6. Load and publish the dataset on Hugging Face

    Install the relevant packages and authenticate with a Hugging Face token that has permission to create or update the repository:

    pip install datasets huggingface_hub evaluate
    huggingface-cli login

    For JSON Lines or CSV files, load the data and create explicit splits:

    from datasets import load_dataset, DatasetDict
    
    raw = load_dataset("json", data_files={
        "train": "data/train.jsonl",
        "validation": "data/validation.jsonl",
        "test": "data/test.jsonl",
    })
    
    raw.push_to_hub("your-org/telugu-benchmark")

    Choose a descriptive repository name and add a dataset card. The card should document the task, intended use, limitations, collection dates, annotation process, licence, known biases, PII checks, split policy, and citation. Include a small loading example and the exact evaluation command. If the data is sensitive or redistribution is restricted, use a gated or private repository rather than claiming it is open.

    Validate the published version by loading it from the Hub in a clean environment. Pin a version or commit hash in benchmark reports so future updates do not silently change results.

    7. Evaluate with Telugu-aware metrics

    Select metrics that match the task and report more than one headline number where necessary:

    • Classification: macro-F1, per-class precision and recall, and accuracy.
    • NER: entity-level precision, recall, and F1 with an explicit matching rule.
    • Translation: chrF alongside BLEU, plus human assessment for adequacy and fluency.
    • Summarisation: factuality and coverage checks, not ROUGE alone.
    • Question answering: exact match or token F1 supplemented by semantic and human review.

    Include confidence intervals or results across multiple seeds for small test sets. Break down performance by domain, text length, code-mixing, regional variation, and difficult examples. A single aggregate score can conceal serious failures for minority categories.

    Establish a simple baseline before comparing large models: a majority classifier, TF-IDF model, or small open model is often enough. For methodology, the article on benchmarking NLP models for Telugu and Sanskrit offers useful comparison principles.

    8. Release responsibly and maintain the benchmark

    Publish a changelog for every revision. Never replace test records without recording what changed, why it changed, and how scores should be compared. Provide a contact route for takedown requests, annotation corrections, and security reports.

    A strong Telugu benchmark is not defined by its size. It is defined by transparent provenance, meaningful linguistic coverage, controlled leakage, reproducible evaluation, and honest limitations. Start with a focused release, involve Telugu-speaking reviewers, and expand only when the evidence shows that the benchmark measures a capability worth improving.

    Last updated 23 September 2026

AIGI may be inaccurate. Replies seeded from the guide above.