0tokens

Apply for AI Grants India

Financial support for innovators building the future of AI in India.

Apply now

Chat · how to benchmark urdu instruction following on indicifeval using hugging face

How to Benchmark Urdu Instruction Following on IndicEval

  1. aigi

    What this benchmark should measure

    How to benchmark Urdu instruction following on IndicEval using Hugging Face is not simply a matter of loading a model and reporting one score. A useful evaluation tests whether a model understands an Urdu instruction, follows its constraints, produces an appropriate answer, and avoids unsupported or unsafe content.

    Indic language evaluation is especially sensitive to script, dialect, register, and translation artefacts. Urdu examples may use formal literary language, conversational Urdu, Arabic- and Persian-derived vocabulary, English code-switching, or Roman Urdu. Record these characteristics before running the benchmark. For broader dataset selection and comparable reporting, use this Indian language LLM benchmark datasets guide.

    Define the evaluation protocol first

    Before installing packages, write down a protocol that another researcher can reproduce. Include:

    • Model identifier and revision, including whether the checkpoint is base, instruction-tuned, or fine-tuned.
    • Tokenizer and chat template used for every example.
    • IndicEval version, dataset revision, and split.
    • Decoding settings, including temperature, top-p, maximum new tokens, and random seed.
    • Hardware and precision, such as bfloat16, float16, or 4-bit inference.
    • Scoring method, evaluator model, and any human-review procedure.

    Run at least one deterministic configuration with temperature=0 or greedy decoding. If the model supports sampling, add multiple-seed results separately rather than mixing them with the primary score. This distinction matters when comparing models with different generation variance. A wider evaluation plan can follow the principles in benchmarking multilingual LLMs in India.

    Set up Hugging Face

    Create an isolated environment and install the packages required by the current IndicEval repository or release documentation. Package names and APIs can change, so verify the official project instructions instead of assuming that a package called indic-eval exposes a particular Python class.

    python -m venv .venv
    source .venv/bin/activate
    pip install -U transformers datasets accelerate evaluate sentencepiece

    For gated or private Hugging Face models, authenticate through the CLI or an access token stored outside source control:

    huggingface-cli login

    Load a compatible causal or sequence-to-sequence model. Replace the placeholder with an actual checkpoint that supports Urdu and confirm its model card, licence, context length, and chat-template requirements.

    import torch
    from transformers import AutoTokenizer, AutoModelForCausalLM
    
    model_id = "org/model-name"
    tokenizer = AutoTokenizer.from_pretrained(model_id, use_fast=True)
    model = AutoModelForCausalLM.from_pretrained(
        model_id,
        torch_dtype=torch.bfloat16,
        device_map="auto",
    )
    model.eval()

    If the checkpoint is too large for available infrastructure, compare a smaller or quantized model rather than silently changing precision between systems. The discussion of small language models for Urdu can help frame that trade-off.

    Prepare and audit the Urdu dataset

    Use the exact IndicEval task format and field names required by the release you are running. Do not invent a generic input and output schema unless the benchmark documentation specifies it. Preserve the original examples, labels, metadata, and split boundaries in an immutable copy.

    Before evaluation, run a data audit for:

    • Unicode normalisation and invisible characters.
    • Urdu Arabic-script text versus Roman Urdu or mixed-script examples.
    • Duplicate prompts across train, development, and test splits.
    • Empty instructions, broken labels, and truncated responses.
    • Unusual length, excessive English text, or machine-translated phrasing.
    • Personally identifiable, copyrighted, or unsafe content requiring controlled handling.

    Keep a manifest containing dataset hash, row counts, language categories, and exclusions. If you filter examples, publish the reason and the resulting denominator. Never report a score without stating how many examples were actually evaluated.

    Format prompts consistently

    Instruction-following scores can change substantially when the same request is wrapped in a different prompt. Use the model’s documented chat template where available:

    def build_prompt(example):
        messages = [
            {"role": "system", "content": "آپ ایک مددگار اردو معاون ہیں۔"},
            {"role": "user", "content": example["instruction"]},
        ]
        return tokenizer.apply_chat_template(
            messages,
            tokenize=False,
            add_generation_prompt=True,
        )

    Do not translate Urdu instructions into English for the primary result. If you add a translated control condition, label it separately. Test whether system prompts, demonstrations, or explicit language instructions introduce a hidden advantage. For Roman Urdu, run a separate slice rather than combining it with Urdu script.

    Generate model responses reproducibly

    Use batched inference where memory allows, but preserve the original example ID and output order. Set a clear stopping policy and prevent the prompt from being counted as the answer.

    inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
    with torch.inference_mode():
        output = model.generate(
            **inputs,
            max_new_tokens=256,
            do_sample=False,
            pad_token_id=tokenizer.eos_token_id,
        )
    answer = tokenizer.decode(
        output[0][inputs["input_ids"].shape[1]:],
        skip_special_tokens=True,
    )

    Save raw outputs, cleaned outputs, generation settings, latency, input tokens, output tokens, and errors. Raw generations are essential for investigating formatting failures and evaluator disagreements. Also report throughput and peak memory when the benchmark informs deployment decisions.

    Score more than exact match

    Select metrics based on task type. Classification tasks may support accuracy, macro-F1, precision, and recall. Generative instruction-following tasks need a combination of:

    • Task correctness, using exact match, structured validation, or reference-based scoring where appropriate.
    • Instruction adherence, checking required format, language, length, and constraints.
    • Semantic quality, using Urdu-capable human or model-assisted assessment with a documented rubric.
    • Safety and factuality, especially for medical, legal, financial, or public-service prompts.
    • Efficiency, including latency, token usage, and failure rate.

    Do not treat BLEU or ROUGE as a complete measure of instruction following. They can penalise valid Urdu paraphrases and reward surface overlap without checking whether the instruction was followed. For open-ended outputs, sample failures for human review and report agreement between reviewers. Keep evaluator prompts fixed, blind model identity where possible, and validate the evaluator on Urdu examples before using it at scale.

    Analyse Urdu-specific failure modes

    Break down results by script, topic, instruction type, and difficulty. Common failure categories include:

    • Answering in Hindi, English, or Roman Urdu when Urdu script was requested.
    • Misreading Urdu punctuation, numerals, dates, or right-to-left formatting.
    • Dropping negation, politeness, quantities, or ordering constraints.
    • Translating an instruction literally while missing its intended action.
    • Hallucinating culturally specific facts or names.
    • Producing a fluent answer that violates a required JSON, list, or table format.

    Create an error taxonomy before reviewing outputs, then count failures consistently. Compare per-slice scores with confidence intervals or bootstrap estimates rather than relying only on a leaderboard average. If your project also evaluates other Indic languages, the methods in benchmarking NLP models for Telugu and Sanskrit offer a useful cross-language comparison structure.

    Report results others can reproduce

    A strong benchmark report includes the model commit, tokenizer, dataset revision, prompt template, decoding configuration, hardware, exclusions, metrics, confidence intervals, and representative Urdu failures. Publish a machine-readable results file and a small evaluation script where licensing permits. Do not publish sensitive prompts or user data merely to make the run appear transparent.

    For each model, show overall performance alongside script and task slices. Include cost and speed if the intended use is an Indian production application. A model with a slightly lower aggregate score may be the better choice if it is faster, cheaper, more reliable on formal Urdu, or less likely to answer in the wrong language.

    FAQ

    Does IndicEval automatically evaluate every Urdu capability?

    No. It evaluates the tasks and splits included in the version you run. Check the release documentation and add targeted Urdu tests for capabilities not covered by the benchmark.

    Should I fine-tune before benchmarking?

    Benchmark the untouched checkpoint first. Fine-tuning can be evaluated as a separate experimental condition, with training data and hyperparameters recorded to avoid contamination concerns.

    How can I reduce hardware cost?

    Use batching, automatic device placement, shorter generation limits, and validated quantisation. Keep the primary comparison on consistent settings, and report any precision changes clearly.

    Where can Indian builders seek support?

    Teams developing open, responsible language technology can explore the AI Grants India programme and prepare a reproducible technical evaluation alongside their application.

    Last updated 23 September 2026

AIGI may be inaccurate. Replies seeded from the guide above.