0tokens

Apply for AI Grants India

Financial support for innovators building the future of AI in India.

Apply now

Chat · how to benchmark marathi question answering on hugging face datasets

How to Benchmark Marathi Question Answering on Hugging Face

  1. aigi

    Why Marathi QA benchmarking needs care

    A Marathi question-answering score is useful only when the dataset, answer format, and evaluation protocol reflect how people actually ask and answer questions. Marathi introduces challenges that generic English QA recipes often hide: flexible word order, inflection, spelling variation, Devanagari punctuation, code-mixed Marathi-English text, and uneven coverage across domains and dialects.

    A reliable benchmark should therefore measure more than one headline number. It should show whether a model can locate evidence, produce an acceptable answer, abstain when the passage does not support one, and remain consistent across topics and language varieties. This is especially important if the model will power public-information services, education tools, or AI question-answering apps for Indian students.

    Define the task before choosing a dataset

    Start by specifying the QA setting:

    • Extractive QA: the answer is a span copied from a supplied Marathi passage.
    • Generative QA: the model writes an answer, potentially paraphrasing or synthesising evidence.
    • Open-domain QA: the system must retrieve relevant documents before answering.
    • Multiple-choice QA: the model selects one option, useful for educational evaluation.
    • Abstention-aware QA: the model must say that the passage does not contain enough information.

    Do not combine these settings in one score. An extractive model that returns a correct text span is solving a different problem from a retrieval-augmented generator. Write down the intended users, domains, answer length, acceptable sources, and whether spelling variants should count as correct.

    For a broader view of evaluation design, compare this workflow with the Indian language LLM benchmark datasets guide and the practical framework for benchmarking multilingual LLMs in India.

    Audit Hugging Face datasets before using them

    Search the Hugging Face Hub for Marathi QA, reading-comprehension, instruction-tuning, and Indic-language datasets, but verify every candidate rather than relying on its name. Inspect the dataset card, licence, source documents, language labels, annotation process, and train-validation-test schema.

    For each example, confirm that the fields are clear and usable—for example, context, question, answers.text, and answers.answer_start for extractive QA. Check whether answer offsets actually point to the answer in the context; offset errors are common after Unicode normalisation.

    Run a small audit that reports:

    • Number of examples and duplicate questions.
    • Context and answer length distributions.
    • Proportion of Marathi, English, Hindi, and code-mixed text.
    • Devanagari versus Latin-script Marathi.
    • Empty, malformed, or multiple-answer records.
    • Topic and source concentration.
    • Overlap between train and test passages or questions.

    Treat generic Marathi web crawls and Wikipedia-derived material as source text, not automatically as validated QA benchmarks. If you generate questions from these sources, document the generation model, human review rate, and rejection criteria.

    Build leakage-resistant splits

    Randomly splitting rows can produce inflated results when several questions use the same article, paragraph, or template. Prefer document-level or source-level splits: all questions derived from one document should remain in one partition. Keep a final, hidden test set if the benchmark will be used for model selection.

    Deduplicate after normalisation, not only by exact string matching. Compare lowercased text where appropriate, Unicode-normalised text, and near-duplicate passages. Also check for question-answer pairs copied from popular training corpora. A model may appear strong because it memorised an article rather than learned Marathi reading comprehension.

    Stratify the evaluation set by difficulty and use case. Useful slices include short versus long contexts, direct fact retrieval versus multi-sentence reasoning, named entities, numerals and dates, negative questions, code-mixed prompts, and regional or dialectal vocabulary. A benchmark with 500 carefully balanced examples can be more informative than a much larger noisy collection.

    Normalise Marathi without hiding real errors

    For Exact Match and token-level F1, publish the normalisation function. Sensible operations may include Unicode NFC normalisation, trimming repeated whitespace, standardising line breaks, and applying a documented policy for Devanagari danda punctuation. Do not silently remove meaningful words, convert all numerals without reporting it, or erase distinctions between Marathi and Hindi.

    Evaluate at least two ways:

    1. Strict scoring, which preserves spelling, punctuation, and formatting differences.
    2. Indic-aware scoring, which tolerates harmless formatting variation while retaining substantive errors.

    For generative answers, add human or model-assisted checks for factual equivalence, completeness, and unsupported claims. Automatic metrics cannot reliably decide whether two Marathi paraphrases convey the same meaning.

    Establish credible baselines

    Before fine-tuning a large model, create simple reference points:

    • A passage-search baseline that returns the sentence containing a query term.
    • A multilingual extractive model evaluated zero-shot.
    • A Marathi-capable encoder fine-tuned on the training split.
    • A retrieval-plus-reader pipeline for open-domain QA.
    • A generative baseline with a fixed prompt and decoding configuration.

    Record model checkpoint, tokenizer, maximum sequence length, hardware, inference batch size, random seed, training steps, learning rate, and decoding parameters. For fine-tuning Marathi models, compare script and dialect coverage using the guidance in fine-tuning AI models for Marathi dialects.

    Do not compare models using different context windows or retrieval corpora without stating the difference. Report at least three seeds for small datasets, and include mean and standard deviation rather than presenting the best run.

    Use metrics that match the task

    For extractive QA, report:

    • Exact Match (EM): whether the predicted answer matches an accepted reference after the published normalisation.
    • Token F1: overlap between predicted and reference tokens, useful for partially correct spans.
    • Answerability accuracy or F1: performance on answerable versus unanswerable questions.
    • No-answer calibration: whether confidence scores support a sensible abstention threshold.

    For open-domain systems, add retrieval Recall@k, MRR, or nDCG, and report reader performance both with gold passages and retrieved passages. This separates retrieval failures from reading failures. For generative QA, combine lexical scores with human ratings for correctness, relevance, completeness, fluency, and citation or evidence support.

    Always publish slice-level results. A single aggregate F1 can conceal that performance collapses on numerals, long contexts, code-mixed prompts, or questions written in less standard Marathi.

    Run error analysis that leads to fixes

    Sample incorrect and borderline predictions by category. Label whether the failure came from tokenisation, retrieval, entity confusion, numerical reasoning, answer-boundary selection, script variation, unsupported generation, or annotation ambiguity. Keep representative examples in a versioned error set so every future model is tested against known weaknesses.

    Look for systematic issues such as confusing Marathi case markers, truncating long contexts, treating a translated question as equivalent to an original one, or returning a plausible answer not supported by the passage. Report confidence alongside errors; high-confidence wrong answers matter more in production than low-confidence misses.

    Make the benchmark reproducible

    Package the dataset revision, preprocessing code, evaluation script, model identifiers, prompts, and environment details. Pin library versions where possible and publish a machine-readable results file. Respect dataset licences and remove personal or sensitive information before release. If the benchmark contains educational or public-service content, document known demographic, geographic, and topical gaps.

    A strong Marathi QA benchmark is not merely a leaderboard. It is a transparent test that tells builders what works, where it fails, and whether improvements survive outside the training distribution. Revisit the test set periodically, but keep historical versions frozen so progress remains comparable.

    Last updated 23 September 2026

AIGI may be inaccurate. Replies seeded from the guide above.