0tokens

Apply for AI Grants India

Financial support for innovators building the future of AI in India.

Apply now

Chat · how to benchmark punjabi instruction following on indicifeval using hugging face

How to Benchmark Punjabi Instruction Following on IndicEval

  1. aigi

    Punjabi evaluation needs more than a single aggregate score. A useful benchmark must separate language understanding from instruction adherence, handle Gurmukhi text correctly, preserve a fixed evaluation protocol, and make model outputs easy to audit. This guide presents a reproducible workflow for benchmarking Punjabi instruction following on IndicEval with Hugging Face tools.

    Before starting, review the broader Indian language LLM benchmark datasets guide to confirm that the dataset version, language labels, splits, and licence meet your project’s requirements. IndicEval interfaces and configuration names can change, so verify the current dataset card and task documentation rather than assuming that an illustrative identifier will work unchanged.

    What the benchmark should measure

    Punjabi instruction following is not the same as Punjabi text generation. A strong evaluation asks whether a model:

    • Understands the instruction and any constraints.
    • Produces an answer in the requested language and script.
    • Follows format requirements such as JSON, a numbered list, or a word limit.
    • Completes the task without adding unsupported claims.
    • Handles culturally and linguistically natural Punjabi rather than relying on transliteration or Hindi-like phrasing.

    Use the official IndicEval task definition as the source of truth for the prompt format, expected fields, scoring method, and test split. Do not fine-tune on the test set or use test prompts in prompt development. If you need a development loop, reserve a validation split or create a documented holdout from the permitted training data.

    Prerequisites and environment

    Use a clean Python environment and record package versions. A typical setup is:

    python -m venv .venv
    source .venv/bin/activate
    pip install -U torch transformers datasets accelerate evaluate sentencepiece pandas tqdm

    For gated or private Hugging Face repositories, authenticate separately with hf auth login. Pin versions in requirements.txt or a lockfile, and record the model revision, device type, quantisation settings, decoding parameters, and IndicEval commit or release.

    You will need:

    • A Hugging Face model that supports Punjabi input and output, preferably an instruction-tuned model.
    • The official IndicEval Punjabi instruction-following configuration.
    • A GPU or sufficient CPU capacity for the selected model.
    • A place to save raw prompts, generated outputs, per-example scores, and logs.

    For model selection, compare several candidates rather than assuming that a multilingual encoder such as mBERT is a suitable generation model. Encoder-only models are useful for classification, but instruction-following generation generally requires a decoder or encoder-decoder model. If model size is a constraint, compare compact candidates using the considerations in What is the best small language model for Punjabi?.

    Load and validate the IndicEval data

    First inspect the available configurations instead of hard-coding a guessed dataset name:

    from datasets import get_dataset_config_names, load_dataset
    
    DATASET_ID = "<official-indiceval-dataset-id>"
    print(get_dataset_config_names(DATASET_ID))
    
    dataset = load_dataset(DATASET_ID, "<punjabi-instruction-following-config>")
    print(dataset)
    print(dataset["test"].column_names)
    print(dataset["test"][0])

    Check the following before generating any answers:

    • Punjabi examples are in the intended script, usually Gurmukhi, unless the task explicitly includes transliteration.
    • Test examples contain no accidental answer leakage or duplicate training records.
    • Instructions, inputs, references, and metadata are mapped to the correct fields.
    • Empty, malformed, or unexpectedly non-Punjabi records are flagged rather than silently removed.
    • The split size matches the benchmark documentation.

    For an India-focused evaluation, retain region, dialect, script, and domain metadata where available. These fields enable meaningful slice analysis without changing the official score.

    Build a deterministic generation harness

    Use the model’s chat template when available. Prompt formatting can materially change results, so keep it identical across models unless the benchmark specifies a different format.

    import json
    import torch
    from tqdm import tqdm
    from transformers import AutoModelForCausalLM, AutoTokenizer
    
    MODEL_ID = "<hugging-face-model-id>"
    tokenizer = AutoTokenizer.from_pretrained(MODEL_ID, use_fast=True)
    model = AutoModelForCausalLM.from_pretrained(
        MODEL_ID,
        torch_dtype=torch.bfloat16 if torch.cuda.is_available() else torch.float32,
        device_map="auto",
    )
    model.eval()
    
    
    def make_prompt(example):
        messages = [{"role": "user", "content": example["instruction"]}]
        if getattr(tokenizer, "chat_template", None):
            return tokenizer.apply_chat_template(
                messages, tokenize=False, add_generation_prompt=True
            )
        return example["instruction"]
    
    outputs = []
    for example in tqdm(dataset["test"]):
        prompt = make_prompt(example)
        inputs = tokenizer(prompt, return_tensors="pt", truncation=True).to(model.device)
        with torch.inference_mode():
            generated = model.generate(
                **inputs,
                max_new_tokens=512,
                do_sample=False,
                temperature=1.0,
                top_p=1.0,
            )
        new_tokens = generated[0, inputs["input_ids"].shape[1]:]
        answer = tokenizer.decode(new_tokens, skip_special_tokens=True).strip()
        outputs.append({"prompt": prompt, "prediction": answer})
    
    with open("punjabi_predictions.jsonl", "w", encoding="utf-8") as f:
        for row in outputs:
            f.write(json.dumps(row, ensure_ascii=False) + "\n")

    The code is a template, not a guarantee that every IndicEval release uses an instruction field. Adapt field names to the official schema and preserve the unmodified raw output. Run at least one smoke test first to catch tokenisation, device, context-length, and prompt-template errors.

    Score with the official evaluator

    Use the evaluator supplied by IndicEval or its documented scoring script. Do not replace a task-specific evaluator with generic accuracy, precision, recall, or F1 unless the task is explicitly classification-based. Open-ended instruction following may require exact match, structured validation, reference-based similarity, rubric scoring, or a combination.

    A robust run should produce:

    • The official aggregate score.
    • Per-example scores and failure reasons.
    • Scores by task type, difficulty, topic, and output format.
    • Invalid-output counts, including malformed JSON or missing required fields.
    • Exact evaluator and dataset versions.

    If the benchmark supports an official command-line interface, prefer it. Otherwise, create a small adapter that maps your JSONL predictions to the evaluator’s expected schema. Keep references and predictions separate, and ensure Unicode is read and written with UTF-8.

    For broader comparability, follow the same principles used in a practical framework for benchmarking multilingual LLMs in India: fixed prompts, identical decoding settings, transparent exclusions, and confidence intervals or bootstrap estimates where the test set is small.

    Punjabi-specific quality checks

    Automated scores should be supplemented with targeted review. Sample failures across categories rather than reading only the worst outputs. Check for:

    • Gurmukhi-to-Latin drift or unexpected script switching.
    • Hindi, Urdu, or English substitutions that change meaning.
    • Incorrect handling of Punjabi honorifics, gender, number, and verb agreement.
    • Literal translations of idioms and culturally specific references.
    • Unrequested explanations, refusals, or safety boilerplate.
    • Hallucinated facts in questions that require concise extraction.

    Create a small annotation sheet with fields for instruction adherence, Punjabi naturalness, factuality, completeness, and formatting. Two native or highly proficient Punjabi reviewers can independently label a sample, then resolve disagreements with a written rubric. Report inter-annotator agreement when the sample is large enough to support it.

    Compare models fairly

    Run every model against the same examples and preserve a manifest containing model ID, revision, prompt template, decoding configuration, hardware, and software versions. Do not compare a deterministic run against a high-temperature sampled run.

    Report both the overall score and the distribution of results. A model with a marginally higher mean may be less useful if it fails consistently on structured outputs or produces unsafe, verbose, or non-Punjabi answers. Include latency, peak memory, throughput, and estimated inference cost when the benchmark informs a production decision.

    If you fine-tune before evaluation, keep the benchmark test set isolated. Hugging Face’s AutoTrain workflow can help structure experiments on Indian datasets; see how to fine-tune with AutoTrain on Indian datasets. Fine-tuning should be reported as a separate model condition, with training data, hyperparameters, checkpoints, and contamination checks documented.

    Common mistakes and a reproducibility checklist

    Avoid these frequent errors:

    • Using a guessed dataset or configuration name without checking the current release.
    • Treating generic text-generation quality as instruction-following performance.
    • Translating Punjabi prompts into English before evaluation.
    • Removing difficult examples without publishing the exclusion rule.
    • Reporting only a single score with no raw predictions or error analysis.
    • Using a judge model without disclosing its identity, prompt, language ability, and calibration.

    Before publishing results, confirm that you have:

    • Pinned model and dataset revisions.
    • Saved raw predictions in UTF-8.
    • Used the official split and evaluator.
    • Recorded decoding and hardware settings.
    • Reported Punjabi script and dialect coverage.
    • Published a clear error taxonomy and representative examples.
    • Checked for training-test contamination.

    Conclusion

    A credible Punjabi IndicEval benchmark is a controlled experiment, not just a script that prints a score. Validate the dataset, use a fixed Hugging Face generation harness, apply the official evaluator, inspect Punjabi-specific failures, and publish enough metadata for another team to reproduce the run. This approach gives Indian AI builders evidence they can use to select models, prioritise data work, and improve real Punjabi products.

    Last updated 23 September 2026

AIGI may be inaccurate. Replies seeded from the guide above.