0tokens

Apply for AI Grants India

Financial support for innovators building the future of AI in India.

Apply now

Chat · how to use hugging face to benchmark marathi on indicgenbench

How to Benchmark Marathi Models on IndicGenBench with Hugging Face

  1. aigi

    What this benchmark is for

    IndicGenBench is useful when you need a repeatable way to compare Marathi-capable language models, not merely a single accuracy number. A sound evaluation should reveal how a model handles Marathi generation, instruction following, factuality, translation, and culturally specific prompts under the same conditions.

    Hugging Face provides the model registry, tokenizers, datasets, inference utilities, and experiment metadata. IndicGenBench supplies the evaluation structure. Used together, they let an Indian-language team test open models before fine-tuning, compare a new checkpoint with a baseline, or document progress for a grant, product review, or research release.

    Before starting, review the wider Indian language LLM benchmark datasets guide so you understand dataset provenance, licensing, language coverage, and contamination risks.

    Prepare a reproducible environment

    Use an isolated environment and record the exact package versions. Benchmarks can change when tokenizers, generation defaults, or evaluation scripts change.

    python -m venv .venv
    source .venv/bin/activate              # Windows: .venv\\Scripts\\activate
    python -m pip install --upgrade pip
    pip install torch transformers datasets accelerate evaluate sentencepiece

    Clone the official IndicGenBench repository and follow its current installation instructions rather than relying on an unverified repository URL. The benchmark may expose a CLI, configuration files, or task-specific runners; use the interface documented by the maintainers. Then capture the environment:

    python --version
    pip freeze > requirements-lock.txt
    git rev-parse HEAD

    For GPU testing, note the GPU model, CUDA version, precision mode, maximum input length, batch size, and whether inference used quantisation. If you are comparing models, keep these settings constant wherever possible. A Marathi model evaluated in 4-bit mode should not be presented as directly equivalent to a full-precision baseline without stating the difference.

    Choose and inspect a Marathi-capable model

    Search the Hugging Face Hub for a model that supports Marathi and fits the task. Check the model card for training languages, licence, intended use, context length, quantisation, and known limitations. The open-source Marathi language models guide is a useful companion when building a shortlist.

    For a causal language model, begin with a minimal loading test:

    from transformers import AutoTokenizer, AutoModelForCausalLM
    
    model_id = "org-or-user/model-name"
    tokenizer = AutoTokenizer.from_pretrained(model_id, use_fast=True)
    model = AutoModelForCausalLM.from_pretrained(
        model_id,
        torch_dtype="auto",
        device_map="auto",
    )
    
    prompt = "मराठी भाषेत भारतातील पावसाळ्याचे थोडक्यात वर्णन करा."
    inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
    outputs = model.generate(**inputs, max_new_tokens=80, do_sample=False)
    print(tokenizer.decode(outputs[0], skip_special_tokens=True))

    Do not assume that a model labelled “multilingual” is equally strong in Marathi. Inspect its tokenizer behaviour, especially how it segments Devanagari text. Compare token counts for equivalent Marathi prompts, and test whether the model preserves punctuation, numerals, names, and code-mixed expressions. Token inefficiency can affect both cost and context capacity.

    Map the model to IndicGenBench tasks

    Read the benchmark task definitions before running anything. Confirm whether the selected model is expected to generate free-form answers, rank candidates, classify labels, or produce structured output. A causal model, encoder model, and instruction-tuned chat model may need different adapters or prompt templates.

    Typical preparation steps are:

    • Select Marathi-only splits where available, and record the split name.
    • Identify whether prompts require a system message, demonstrations, or a fixed template.
    • Match the expected output format exactly, including JSON, labels, or translated text.
    • Disable sampling for deterministic comparison unless the task explicitly evaluates diversity.
    • Set a fixed random seed and generation limits.
    • Keep the original Marathi text intact; do not silently normalise spelling or punctuation.

    If IndicGenBench provides a configuration file, place model ID, language, task names, batch size, and generation parameters there. If it provides a command-line runner, the shape will typically resemble the following, but use the repository’s documented flags rather than copying this literally:

    python run_benchmark.py \
      --model_id org-or-user/model-name \
      --language mr \
      --tasks task_a,task_b \
      --batch_size 4 \
      --seed 42 \
      --output_dir results/marathi-model-name

    The language code may be mr, mar, or a benchmark-specific name. Verify it in the dataset or task configuration.

    Make the evaluation fair

    A credible Marathi benchmark needs more than a leaderboard score. Run at least one strong baseline, preserve raw predictions, and report the full configuration. For each model, store:

    • Hugging Face model ID and revision or commit hash
    • IndicGenBench commit and dataset version
    • Prompt template and chat format
    • Decoding settings, seed, precision, and hardware
    • Runtime, peak memory, and failures
    • Task-level metrics and aggregate scores

    Use the metric appropriate to the task. Accuracy and macro-F1 can be useful for classification; exact match may suit structured answers; BLEU or chrF can support translation analysis; and rubric-based or model-assisted evaluation requires a clearly documented rubric and spot checks by Marathi speakers. Automatic metrics alone can miss grammatical errors, unnatural register, factual mistakes, or regional language variation.

    For stronger conclusions, run repeated seeds or bootstrap confidence intervals. Avoid averaging away a severe weakness in one task: report per-task and per-category results. A model that performs well on short factual prompts may still fail on long Marathi instructions or code-mixed user queries. For broader comparison, see this practical framework for benchmarking multilingual LLMs in India.

    Analyse Marathi-specific errors

    After scoring, sample correct and incorrect outputs. Group errors into categories such as:

    • Devanagari spelling, agreement, gender, or inflection errors
    • Hindi or English substitution where Marathi is expected
    • Hallucinated facts about Maharashtra, India, or local institutions
    • Failure to follow Marathi instructions or preserve named entities
    • Inconsistent handling of numerals, dates, currency, and transliterated names
    • Overly formal, unnatural, or regionally inappropriate register
    • Unsafe or biased responses in Marathi

    Create a small human review set with native or highly proficient Marathi evaluators. Ask reviewers to rate correctness, fluency, instruction following, and cultural appropriateness separately. This separates a genuinely useful model from one that produces fluent but incorrect text. If dialect coverage matters, evaluate it explicitly and consult guidance on fine-tuning AI models for Marathi dialects.

    Troubleshooting and reporting

    Common failures include missing padding tokens, tokenizer-model mismatches, out-of-memory errors, malformed generated output, and inconsistent chat templates. Resolve them by checking the model card, setting an explicit padding strategy, reducing batch size or sequence length, and validating one sample before launching the full run. Never discard failed examples without reporting the failure rate.

    Publish the configuration, evaluation script, dataset access instructions, raw predictions where licensing permits, and a short limitations section. If the benchmark data is private, provide hashes and a reproducible schema. As of 2026, this level of provenance is essential because model revisions and hosted inference behaviour can change quickly.

    A useful benchmark is not just a score. It is an auditable record of which Marathi capabilities were tested, under what conditions, and where the model still needs work.

    Last updated 23 September 2026

AIGI may be inaccurate. Replies seeded from the guide above.