0tokens

Apply for AI Grants India

Financial support for innovators building the future of AI in India.

Apply now

Chat · how to benchmark marathi instruction following on indicifeval using hugging face

How to Benchmark Marathi Instruction Following on IndicEval

  1. aigi

    What this benchmark should answer

    A useful Marathi instruction-following benchmark should do more than produce one score. It should show whether a model can understand Marathi prompts, follow explicit constraints, preserve meaning, avoid unsupported claims, and respond in an appropriate register. It should also make comparisons fair across model versions, decoding settings, and hardware.

    IndicEval can provide the evaluation structure, while Hugging Face supplies the model, tokenizer, datasets, and reproducible inference tooling. Before running anything, define the exact task split, model revision, prompt format, generation parameters, and metrics you will report. This is especially important for Marathi, where script, dialect, code-mixing, transliteration, and tokenisation can materially affect results.

    For broader context on test-set design, use this Indian language LLM benchmark datasets guide. If your comparison includes several languages, the practical framework for benchmarking multilingual LLMs in India covers cross-language controls and reporting.

    1. Confirm the IndicEval task and dataset

    Do not assume that a package name or task identifier in an old tutorial still matches the current release. Check the IndicEval repository or documentation for the supported Marathi instruction-following task, dataset revision, licence, expected input fields, and official scoring script. Record the commit hash or release version in your experiment log.

    Inspect the dataset before evaluation:

    • Confirm that prompts are written in Marathi Devanagari and identify any transliterated or mixed-language examples.
    • Check whether each item has a reference answer, a list of required constraints, or a rubric for human or model-assisted judging.
    • Remove accidental duplicates across development and test splits.
    • Verify that prompts do not expose reference answers or benchmark metadata.
    • Preserve the original test set; create a separate filtered copy if quality control is necessary.

    Instruction following is not equivalent to exact-match classification. A response can be semantically correct while using different wording, so report the benchmark's official metric alongside task-specific checks rather than replacing it with a simple string comparison.

    2. Create a reproducible Hugging Face environment

    Use a fresh virtual environment and pin the main dependencies. The exact versions should match your hardware and the IndicEval release you are using.

    python -m venv .venv
    source .venv/bin/activate                 # Windows: .venv\\Scripts\\activate
    python -m pip install --upgrade pip
    pip install "transformers>=4.40" datasets accelerate evaluate sentencepiece
    # Install IndicEval using its official repository or release instructions.

    If the model is gated, authenticate with Hugging Face through the CLI or an access token stored outside source control. Never commit tokens, private dataset credentials, generated personal data, or raw user prompts to a public repository.

    Capture the environment for repeatability:

    pip freeze > requirements-lock.txt
    nvidia-smi > hardware.txt

    Select a model that actually supports Marathi and instruction tuning. An open Marathi model may be a better baseline than a larger multilingual model for script fidelity or local terminology; compare both where resources allow. The open-source Marathi language models guide is useful when choosing candidate checkpoints.

    3. Load the model and prepare generation

    Use the model architecture advertised on its Hugging Face model card. Causal language models generally require AutoModelForCausalLM; encoder-decoder checkpoints require AutoModelForSeq2SeqLM.

    import torch
    from transformers import AutoTokenizer, AutoModelForCausalLM
    
    model_id = "your-org/your-marathi-instruct-model"
    tokenizer = AutoTokenizer.from_pretrained(model_id, use_fast=True)
    model = AutoModelForCausalLM.from_pretrained(
        model_id,
        torch_dtype=torch.bfloat16 if torch.cuda.is_available() else torch.float32,
        device_map="auto",
    )
    model.eval()
    
    if tokenizer.pad_token_id is None:
        tokenizer.pad_token = tokenizer.eos_token

    Use the checkpoint's chat template when one is provided. Do not silently prepend an English system prompt to Marathi test items: that changes the task. Keep system instructions, user prompts, and output extraction identical across models.

    For a benchmark, deterministic decoding is usually the cleanest starting point:

    def generate_answer(prompt, max_new_tokens=256):
        messages = [{"role": "user", "content": prompt}]
        rendered = tokenizer.apply_chat_template(
            messages, tokenize=False, add_generation_prompt=True
        )
        inputs = tokenizer(rendered, return_tensors="pt").to(model.device)
        with torch.inference_mode():
            output = model.generate(
                **inputs,
                max_new_tokens=max_new_tokens,
                do_sample=False,
                pad_token_id=tokenizer.pad_token_id,
            )
        new_tokens = output[0, inputs["input_ids"].shape[1]:]
        return tokenizer.decode(new_tokens, skip_special_tokens=True).strip()

    Run a small smoke test first. Check that the model answers in Marathi, does not echo the prompt, and stops within the configured limit. Also measure prompt and completion token counts; long Marathi outputs can expose context-window or truncation problems.

    4. Run IndicEval without changing the test conditions

    Adapt the following structure to the current IndicEval API rather than copying an obsolete import or placeholder function:

    from datasets import load_dataset
    import json
    
    # Replace with the dataset identifier and split documented by IndicEval.
    data = load_dataset("your-indiceval-dataset", split="test")
    records = []
    
    for row in data:
        prompt = row["prompt"]
        prediction = generate_answer(prompt)
        records.append({
            "id": row.get("id"),
            "prompt": prompt,
            "prediction": prediction,
        })
    
    with open("marathi_predictions.jsonl", "w", encoding="utf-8") as f:
        for record in records:
            f.write(json.dumps(record, ensure_ascii=False) + "\\n")

    Then invoke IndicEval's documented evaluator on the saved predictions. Saving outputs separately lets you rerun scoring, inspect failures, and compare evaluators without generating responses again. Keep a manifest containing model ID and revision, dataset revision, tokenizer, chat template, seed, decoding settings, maximum input and output lengths, batch size, GPU type, and timestamp.

    5. Score more than one number

    Report the official IndicEval score first, followed by a breakdown that explains model behaviour. Depending on the task, useful measures include:

    • Instruction compliance: whether every requested action or constraint was satisfied.
    • Answer quality: factuality, relevance, completeness, and clarity.
    • Language quality: Marathi fluency, Devanagari consistency, grammar, and appropriate formality.
    • Safety and refusal behaviour: whether harmful or impossible requests were handled correctly.
    • Efficiency: latency, throughput, peak memory, and output-token count.

    If the benchmark uses an LLM judge, fix the judge model, rubric, prompt, and sampling settings, and disclose them. Automated judging can be sensitive to Marathi fluency, verbosity, and formatting. Sample outputs should be reviewed by Marathi speakers, ideally using a blinded rubric. Track common errors such as Hindi substitution, English leakage, literal translation, incorrect honorifics, hallucinated facts, and failure to follow numbered constraints.

    For models that target dialectal or domain-specific Marathi, compare performance by category rather than relying on an aggregate. The guidance on fine-tuning AI models for Marathi dialects can help structure those slices.

    6. Make results comparable and actionable

    Run each model under the same prompt template, test split, context limit, generation policy, and scoring version. Do not tune decoding on the hidden test set. Use a development split for prompt or parameter selection, then run the final test once. If you make multiple model calls or use stochastic decoding, report the number of seeds and confidence intervals or bootstrap intervals where practical.

    A good report includes:

    • model name, revision, parameter count, and licence;
    • Marathi coverage and tokenizer details;
    • IndicEval release, dataset split, and item count;
    • official score plus per-category results;
    • judge or human-review protocol;
    • hardware, runtime, memory, and cost estimates;
    • representative successes and anonymised failure cases;
    • limitations, exclusions, and known data contamination risks.

    The result should guide the next engineering decision. If compliance is weak, improve instruction data and prompt formatting. If language quality is weak, inspect Marathi data balance, tokenisation, and dialect coverage. If quality is strong but latency is too high, test quantisation or a smaller checkpoint while rerunning the same evaluation. For teams building a model rather than only comparing checkpoints, a small language model for Marathi offers a complementary development path.

    FAQ

    Can any Hugging Face model be evaluated?
    Yes, if it can accept the benchmark's input format and produce text, but architecture, chat template, Marathi coverage, and licence must be checked first.

    Should Marathi prompts be translated into English for scoring?
    No. Translate only for analysis if needed. The primary evaluation should preserve the original Marathi prompt and judge the Marathi response against the task requirements.

    Is IndicEval's score enough to select a production model?
    No. Combine it with human review, safety testing, latency, cost, domain-specific evaluations, and monitoring on representative Indian usage data.

    How should I publish the outputs?
    Publish code, configuration, environment details, aggregate scores, and carefully reviewed examples. Remove personal or sensitive content and respect dataset and model licences.

    Apply for AI Grants India

    If you are building Marathi or other Indian-language AI systems, apply for AI Grants India to explore support, funding, and ecosystem resources.

    Last updated 23 September 2026

AIGI may be inaccurate. Replies seeded from the guide above.