0tokens

Apply for AI Grants India

Financial support for innovators building the future of AI in India.

Apply now

Chat · how to create a marathi benchmark dataset on hugging face

How to Create a Marathi Benchmark Dataset on Hugging Face

  1. aigi

    Marathi benchmark datasets are valuable only when they measure real language ability rather than memorisation, translation shortcuts, or quirks in the collection process. A useful release should let researchers compare models fairly, reproduce results, and understand where systems fail across Marathi’s registers, domains, and dialects. This guide explains how to create a Marathi benchmark dataset on Hugging Face, from task definition and sourcing to evaluation, documentation, and maintenance.

    Start with a precise benchmark brief

    Write a short dataset specification before collecting examples. It should answer:

    • What task is being measured? Examples include sentiment classification, intent detection, named-entity recognition, extractive question answering, summarisation, translation, or Marathi language understanding.
    • Who is the intended user? Define whether the benchmark targets foundation models, Indic-language encoders, speech systems, student projects, or production applications.
    • What constitutes success? Choose metrics before annotation. Accuracy may be unsuitable for imbalanced labels; macro-F1, exact match, character error rate, word error rate, or semantic similarity may be more appropriate.
    • Which Marathi varieties are represented? Record region, dialect, formal and informal usage, code-mixing, spelling variation, and scripts where relevant.

    Avoid combining unrelated tasks into one score. A compact, well-controlled benchmark is more useful than a large collection with ambiguous labels. For broader design principles, compare your plan with guidance on low-resource language datasets for AI training in India and Indian-language evaluation practices in benchmarking multilingual LLMs in India.

    Source data lawfully and representatively

    Potential sources include public Marathi news, government documents, educational material, community contributions, licensed corpora, synthetic prompts, and field-collected examples. Do not assume that a publicly accessible webpage is licensed for redistribution. Maintain a source register containing the URL or collection context, access date, licence, consent status, domain, and any removal requirements.

    Representation matters as much as volume. Track coverage across:

    • Urban and rural contexts, age groups, and regions of Maharashtra
    • Formal writing, conversational Marathi, social-media language, and code-mixed Marathi-English
    • Topics such as health, agriculture, education, public services, commerce, and everyday assistance
    • Orthographic variation, transliterated Marathi, spelling errors, and dialectal vocabulary

    Remove personal information unless it is essential, explicitly consented to, and safely handled. For sensitive data, publish transformed or redacted examples and document the process. A benchmark should not expose phone numbers, addresses, account details, private conversations, or identifiable health information.

    Design annotation guidelines before labeling

    Labels need operational definitions, not just names. For every class or span type, include:

    • A definition and at least three positive examples
    • Borderline cases and explicit exclusions
    • Treatment of punctuation, emojis, code-switching, and spelling variation
    • Rules for ambiguous, offensive, or culturally specific language
    • A policy for “uncertain” or “cannot determine” cases

    Use Marathi-first instructions. Translating an English guideline literally can produce labels that do not fit Marathi grammar or social context. Recruit annotators who understand the target variety, pay for expert review where necessary, and separate annotation from adjudication.

    Measure agreement on a pilot batch before labeling the full corpus. Low agreement usually indicates unclear guidelines or overlapping categories, not poor annotators. Resolve disagreements through documented adjudication and version the guidelines whenever a rule changes.

    Build leakage-resistant train, validation, and test splits

    A benchmark’s credibility depends heavily on its splits. Randomly dividing sentences can inflate scores when near-duplicates, articles from the same source, template prompts, or dialogue turns appear in multiple partitions.

    Use a split strategy appropriate to the task:

    • Group related documents, users, speakers, or source outlets before splitting.
    • Keep test examples inaccessible to model developers when a trusted evaluation server is important.
    • Deduplicate exact matches and near-duplicates using normalised text and similarity checks.
    • Preserve meaningful distributions without making the test set predictable.
    • Create challenge subsets for dialects, code-mixing, long contexts, rare entities, or noisy spelling.

    For generative tasks, publish prompts and references carefully. Explain whether multiple valid answers exist and whether evaluation uses human judgment, rubric-based scoring, or automatic metrics. If you plan to compare Marathi systems with other Indic models, the framework in benchmarking NLP models for Telugu and Sanskrit offers a useful cross-language perspective.

    Prepare a clean, machine-readable schema

    Use stable column names and explicit data types. A classification record might contain id, text, label, domain, region, and source_id. A question-answering record may need context, question, answers, answer_start, and is_impossible. Keep metadata separate from the text where possible, and never use labels that reveal the answer through naming conventions.

    A practical repository layout is:

    marathi-benchmark/
    ├── README.md
    ├── LICENSE
    ├── CITATION.cff
    ├── data/
    │   ├── train.jsonl
    │   ├── validation.jsonl
    │   └── test.jsonl
    ├── scripts/
    │   ├── validate.py
    │   └── deduplicate.py
    └── docs/
        ├── annotation-guidelines.md
        └── datasheet.md

    Validate Unicode consistently. Marathi uses Devanagari, so inspect normalisation, combining marks, zero-width characters, punctuation, and numerals. Do not remove diacritics or punctuation automatically unless the benchmark definition requires it. Preserve the original text and, if normalised text is provided, publish the transformation rule.

    Publish the dataset on Hugging Face

    Install the current libraries and authenticate with a write-enabled token:

    pip install -U datasets huggingface_hub
    huggingface-cli login

    For JSONL files, a simple loading path is:

    from datasets import load_dataset
    
    files = {
        "train": "data/train.jsonl",
        "validation": "data/validation.jsonl",
        "test": "data/test.jsonl",
    }
    
    dataset = load_dataset("json", data_files=files)
    dataset.push_to_hub("your-org/marathi-benchmark")

    Before publishing, run schema checks for missing fields, invalid labels, duplicate IDs, unexpected Unicode, and split contamination. Add a dataset card covering the motivation, task, languages and varieties, collection process, annotation workflow, licence, consent, known limitations, intended uses, prohibited uses, and evaluation protocol. Include a reproducible baseline with pinned package versions and the exact commands needed to recreate reported scores.

    If the data is sensitive, restricted, or too large for unrestricted release, use Hugging Face access controls or publish a controlled-access version with a clear request process. Do not upload secrets, raw personal data, or copyrighted material without redistribution rights.

    Evaluate baseline models and publish failure analysis

    A benchmark is incomplete without baselines. Report a majority or heuristic baseline, a strong multilingual baseline, and at least one Marathi-capable model where licensing permits. Record preprocessing, prompt format, decoding settings, hardware, random seeds, and confidence intervals when feasible.

    Do not report only one aggregate number. Break results down by domain, dialect or region, text length, code-mixing, label, and challenge subset. Inspect errors manually and classify them: factual failure, script handling, morphology, named entities, pragmatics, toxic content, or annotation ambiguity. This makes the dataset useful for model improvement and connects naturally to work on fine-tuning AI models for Marathi dialects.

    Version, govern, and maintain the release

    Use semantic or date-based versions and maintain a changelog. Record additions, corrected labels, removed items, changed split assignments, and metric changes. Keep stable IDs so users can trace revisions. Provide a contact or issue template for takedown requests, annotation disputes, and problematic examples.

    Plan a refresh cycle rather than silently editing the test set. Major changes should create a new version and trigger fresh baseline runs. If the benchmark becomes widely used, consider a hidden test server to reduce overfitting and leaderboard gaming. A transparent governance policy is as important as the dataset itself.

    Final checklist

    Before release, confirm that you have:

    • A narrowly defined task and pre-registered evaluation metric
    • Documented licences, consent, provenance, and privacy safeguards
    • Marathi-first annotation guidelines and measured agreement
    • Deduplicated, leakage-resistant splits
    • A validated schema and reproducible loading script
    • Dataset documentation, citation details, and limitations
    • Baselines, subgroup results, and qualitative error analysis
    • Versioning, issue handling, and a process for corrections

    A carefully designed Marathi benchmark can support research, public-interest technology, and production evaluation without treating Marathi as a smaller copy of English. For teams building broader systems, pair the benchmark with guidance on how to train LLMs on Indian datasets and the wider open-source AI datasets for India.

    Last updated 23 September 2026

AIGI may be inaccurate. Replies seeded from the guide above.