0tokens

Apply for AI Grants India

Financial support for innovators building the future of AI in India.

Apply now

Chat · how to benchmark hindi instruction following on indicifeval using hugging face

How to Benchmark Hindi Instruction Following on IndicEval

  1. aigi

    Hindi instruction-following benchmarks are useful only when the evaluation setup is reproducible. A score can change because of the model, prompt template, decoding parameters, tokenizer limits, or evaluator—not just because one model is better. This guide presents a practical workflow for benchmarking Hindi-capable generative models on IndicEval with Hugging Face, while keeping Hindi script, Hinglish, formatting, and safety behaviour visible in the results.

    Before selecting a model, review the wider landscape in this guide to Indian-language LLM benchmark datasets. IndicEval versions, task names, splits, licensing, and scoring conventions may change, so verify the benchmark repository and dataset card rather than assuming that a dataset identifier or configuration is available on the Hugging Face Hub.

    What you are measuring

    Hindi instruction following is broader than text classification. A useful evaluation asks whether a model can:

    • Understand an instruction written in Devanagari, Romanised Hindi, or a mixed Hindi-English form.
    • Produce the requested output type, such as a list, JSON object, summary, translation, or explanation.
    • Respect constraints on length, tone, audience, structure, and prohibited content.
    • Preserve meaning without silently switching to English or another Indic language.
    • Decline unsafe or impossible requests consistently and clearly.

    Report results by task and slice, not only as one aggregate number. At minimum, separate native Hindi script, Hinglish, code-mixed prompts, short instructions, long-context instructions, and examples requiring strict formatting. This makes the benchmark actionable for teams building open-source small language models for Hindi.

    Prepare a reproducible Hugging Face environment

    Use a pinned environment and record the exact model revision, dataset revision, Python version, GPU type, and decoding configuration. A minimal setup might include:

    pip install -U "transformers>=4.45" datasets accelerate evaluate sentencepiece

    For large causal language models, use AutoModelForCausalLM rather than a sequence-classification model. IndicEval instruction-following tasks generally require generated answers; AutoModelForSequenceClassification is appropriate only when the benchmark explicitly defines a classification formulation.

    from datasets import load_dataset
    from transformers import AutoTokenizer, AutoModelForCausalLM
    
    model_id = "your-org/your-hindi-capable-model"
    tokenizer = AutoTokenizer.from_pretrained(model_id, use_fast=True)
    model = AutoModelForCausalLM.from_pretrained(
        model_id,
        device_map="auto",
        torch_dtype="auto"
    )

    If the model uses a chat template, use it exactly as documented. Do not manually concatenate User: and Assistant: labels unless the model documentation specifies that format. A mismatched template can dominate the result.

    Load and inspect IndicEval carefully

    The exact load_dataset() call depends on the published IndicEval package and configuration. Use the official dataset identifier or local files confirmed in the dataset card:

    from datasets import load_dataset
    
    dataset = load_dataset("official-org/indiceval", "hindi")
    print(dataset)
    print(dataset["test"].column_names)
    print(dataset["test"][0])

    If IndicEval is distributed through a repository rather than the Hub, download the pinned release and load its JSON, JSONL, or Parquet files with the matching loader. Before generation, check:

    • Whether the test split contains reference answers or only prompts.
    • The fields used for instructions, context, references, task labels, and metadata.
    • Duplicate prompts and near-duplicates across train, validation, and test sets.
    • Unicode normalisation, Devanagari punctuation, zero-width characters, and accidental HTML.
    • Whether examples contain private, licensed, or sensitive material that should not be sent to a hosted endpoint.

    Keep the original record ID and all metadata in your output file. Never overwrite the source dataset during preprocessing.

    Build the generation loop

    Use deterministic decoding for a primary leaderboard-style run. Fix do_sample=False, set a documented maximum number of new tokens, and define a pad token where the model requires one. Use the benchmark’s prescribed prompt format; if none exists, publish your template with the results.

    import torch
    
    def make_prompt(row):
        instruction = row["instruction"]
        messages = [{"role": "user", "content": instruction}]
        return tokenizer.apply_chat_template(
            messages, tokenize=False, add_generation_prompt=True
        )
    
    def generate_answer(row):
        prompt = make_prompt(row)
        inputs = tokenizer(
            prompt,
            return_tensors="pt",
            truncation=True,
            max_length=tokenizer.model_max_length
        ).to(model.device)
    
        with torch.inference_mode():
            output = model.generate(
                **inputs,
                do_sample=False,
                max_new_tokens=512,
                pad_token_id=tokenizer.eos_token_id
            )
    
        generated = output[0, inputs["input_ids"].shape[1]:]
        return tokenizer.decode(generated, skip_special_tokens=True).strip()

    For batch evaluation, use a data collator and batch tokenisation to improve throughput. Measure truncation rate, input length, output length, latency, and failures. A model that receives truncated Hindi instructions is not being evaluated fairly.

    Score the right properties

    Use the metric specified by IndicEval for official comparison. Depending on the task, this may include exact match, token-level F1, ROUGE-style overlap, classification accuracy, or a structured-output check. Do not replace an official metric with a generic Trainer.evaluate() call: Trainer does not automatically know how to generate answers or score free-form Hindi responses.

    For instruction following, add diagnostic metrics alongside the official score:

    • Format compliance: valid JSON, required headings, list length, or schema validity.
    • Constraint adherence: word limits, requested language, tone, and required facts.
    • Semantic correctness: human or validated model-assisted assessment against references.
    • Hindi fidelity: Devanagari preservation, appropriate vocabulary, and unwanted language switching.
    • Safety and refusal quality: correct handling of disallowed requests without over-refusal.

    If using an LLM judge, treat it as a secondary signal. Keep the judge model, prompt, temperature, rubric, and random seed fixed; sample a human-audited subset; and report disagreement. Hindi evaluation can be distorted by judges that favour English fluency or penalise legitimate variation in Hindi phrasing.

    Create Hindi-specific error slices

    A single average hides the failures that matter in production. Tag each example by script, domain, instruction length, task type, and output format. Then inspect failures in categories such as:

    • Correct meaning but wrong format.
    • Correct format but incomplete or hallucinated content.
    • Devanagari prompt answered mainly in English.
    • Romanised Hindi misunderstood or normalised incorrectly.
    • Numbers, dates, names, and units copied incorrectly.
    • Politeness, gender, register, or regional usage mishandled.
    • Refusal triggered for a benign request, or missing for a harmful one.

    Compare the model against a strong multilingual baseline and a Hindi-focused baseline. For broader methodology, see this practical framework for benchmarking multilingual LLMs in India, and use the same prompt and decoding policy across models.

    Report results so others can reproduce them

    Publish a table containing model revision, quantisation, hardware, dataset revision, number of examples, prompt template, maximum input and output tokens, decoding settings, official metrics, diagnostic metrics, and evaluation failures. Include confidence intervals or bootstrap intervals where the test set is large enough. Report per-slice scores and sample outputs, with sensitive content redacted.

    Do not compare scores produced with different test splits or hidden preprocessing. If you fine-tuned on any IndicEval-derived data, label the result as contaminated or in-domain and exclude it from a clean zero-shot comparison. For teams evaluating Telugu, Sanskrit, or other languages, the same discipline applies; the related Telugu and Sanskrit NLP benchmarking guide offers a useful comparison point.

    A practical evaluation checklist

    Before publishing a Hindi IndicEval result, confirm that you have:

    • Pinned the model, tokenizer, dataset, and code revisions.
    • Verified the official split and task schema.
    • Used the correct causal or classification model architecture.
    • Applied the documented chat template and fixed decoding settings.
    • Measured truncation, failed generations, and malformed outputs.
    • Reported official scores plus Hindi-specific diagnostic slices.
    • Audited automated judging with human review.
    • Saved raw prompts, outputs, IDs, and scoring logs securely.

    Benchmarking is not the final product metric. After IndicEval, test representative user traffic, latency, cost, and safety in your target deployment. A model that performs well on a benchmark may still struggle with noisy WhatsApp-style Hindi, domain terminology, or voice-transcribed text. If speech is part of the product, pair text evaluation with Hindi ASR testing using resources such as Hindi ASR low-WER evaluation guidance.

    Last updated 23 September 2026

AIGI may be inaccurate. Replies seeded from the guide above.