0tokens

Apply for AI Grants India

Financial support for innovators building the future of AI in India.

Apply now

Chat · how to benchmark malayalam instruction following on indicifeval using hugging face

How to Benchmark Malayalam Instruction Following with IndicEval

  1. aigi

    Why this benchmark matters

    Malayalam instruction following needs more than a single overall score. A model may answer factual questions well but fail on formatting, safety, multi-step reasoning, or instructions written in colloquial Malayalam. A reproducible benchmark helps you distinguish genuine capability from prompt sensitivity, translation artefacts, and evaluation noise.

    This guide presents a practical workflow for benchmarking Malayalam instruction following with IndicEval and Hugging Face. Before running experiments, confirm the current IndicEval task name, dataset configuration, split, licence, and evaluator version in the official project documentation or model repository. Tool APIs and benchmark packaging can change; do not assume that an indicieval Python package or a particular class exists simply because a benchmark is associated with Hugging Face.

    For broader dataset-selection principles, see this Indian language LLM benchmark datasets guide. If you are comparing several Indic languages, the same workflow can be extended using a practical framework for benchmarking multilingual LLMs in India.

    Define the evaluation before installing anything

    Write down the evaluation contract first. Record:

    • The exact model revision, tokenizer revision, and quantisation setting
    • The IndicEval task, dataset configuration, split, and evaluator commit
    • Whether the model is evaluated zero-shot, few-shot, or after Malayalam-specific fine-tuning
    • The prompt template, system message, number of demonstrations, and language used for demonstrations
    • Decoding settings, including temperature, top-p, maximum new tokens, and stop sequences
    • Hardware, software versions, random seed, and batch size

    Instruction following is usually a generation task, not ordinary sequence classification. The earlier pattern of loading AutoModelForSequenceClassification and taking logits.argmax() is unsuitable unless the benchmark explicitly defines a classification head and labels. For generative models, use AutoModelForCausalLM or the architecture required by the checkpoint, then score the generated answer against the benchmark’s references or rubric.

    Set up a reproducible Hugging Face environment

    Create an isolated environment and pin dependencies. A typical starting point is:

    python -m venv .venv
    source .venv/bin/activate
    pip install -U "transformers>=4.45" "datasets>=2.20" "accelerate" "evaluate" "sentencepiece" "safetensors"

    Install any IndicEval-specific package only when its documentation requires it. Some benchmarks are distributed as datasets, evaluation scripts, or leaderboard repositories rather than a package named indicieval.

    Authenticate with Hugging Face if the model or dataset is gated:

    huggingface-cli login

    Then load the model with an explicit revision. For a causal language model:

    import torch
    from transformers import AutoModelForCausalLM, AutoTokenizer
    
    model_id = "your-org/your-malayalam-capable-model"
    revision = "main"  # replace with a commit hash for archival runs
    
    tokenizer = AutoTokenizer.from_pretrained(model_id, revision=revision)
    model = AutoModelForCausalLM.from_pretrained(
        model_id,
        revision=revision,
        torch_dtype=torch.bfloat16,
        device_map="auto",
    )
    model.eval()

    Use torch.float16 when your accelerator does not support bfloat16. Check the model card for chat-template requirements; applying the wrong template can change results substantially.

    Load and inspect the Malayalam data

    Use the dataset identifier and configuration specified by IndicEval rather than guessing:

    from datasets import load_dataset
    
    dataset = load_dataset(
        "OWNER/INDIC-EVAL-DATASET",
        "malayalam_instruction_following",
        revision="DATASET_COMMIT_OR_TAG",
    )
    print(dataset)
    print(dataset["test"][0])

    Inspect field names, language labels, reference answers, metadata, and any hidden answer columns. Keep the official test set untouched. If you need local development examples, create a separate validation subset and document how it was sampled.

    Before evaluation, check for:

    • Mixed Malayalam-English prompts and code-switching
    • Unicode normalisation differences, especially combining marks and punctuation
    • Duplicate or near-duplicate prompts across splits
    • Unsafe, ambiguous, or culturally specific requests
    • References that contain multiple acceptable answers
    • Instructions requiring tables, JSON, lists, or exact strings

    Do not silently transliterate Malayalam into Latin script or translate prompts into English. Those can be useful auxiliary experiments, but they are not equivalent to native Malayalam instruction following.

    Build a controlled generation loop

    Use the benchmark’s required prompt format. If the model has a chat template, prefer it over manually concatenating role labels:

    def make_prompt(example):
        messages = [{"role": "user", "content": example["prompt"]}]
        return tokenizer.apply_chat_template(
            messages, tokenize=False, add_generation_prompt=True
        )
    
    def generate_one(example):
        prompt = make_prompt(example)
        inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
        with torch.inference_mode():
            output = model.generate(
                **inputs,
                max_new_tokens=512,
                do_sample=False,
                pad_token_id=tokenizer.eos_token_id,
            )
        new_tokens = output[0, inputs["input_ids"].shape[1]:]
        return tokenizer.decode(new_tokens, skip_special_tokens=True).strip()

    For headline scores, deterministic decoding makes runs easier to compare. For robustness, repeat a smaller sample with controlled temperatures and report score variation. Ensure that the prompt is not included in the decoded answer, and save raw outputs before any cleaning.

    Score the right behaviours

    Follow IndicEval’s official scoring implementation whenever available. Instruction-following evaluation may combine exact match, normalised match, task-specific validators, reference-based similarity, or rubric-based judging. Do not replace these with generic accuracy or F1 without clearly labelling the result.

    Report results by task category, not only as one Malayalam average. Useful slices include:

    • Instruction adherence and required format
    • Factual question answering
    • Summarisation and rewriting
    • Reasoning or multi-step execution
    • Safety and refusal behaviour
    • Code-switching and dialect variation
    • Prompt length and difficulty

    If an LLM judge is used, publish the judge model, prompt, language, sampling settings, and agreement checks. Malayalam answers should be judged by fluent evaluators or validated bilingual protocols; an English-only judge can reward translated-looking output and miss grammatical or cultural errors.

    Analyse failures and publish a credible report

    Save one record per example containing the prompt hash, model revision, generation settings, output, reference, score, and error label. Review a stratified sample of failures manually. Common labels include wrong task, incomplete answer, unsupported claim, formatting violation, language switch, refusal error, and hallucination.

    Report the number of examples, excluded items, confidence intervals or bootstrap ranges where practical, and separate development decisions from final test results. Compare against strong baselines: a Malayalam-capable open model, a multilingual model, and—if relevant—a translated-prompt baseline. Results from other Indic languages can provide useful context; see benchmarking NLP models for Telugu and Sanskrit.

    For production use, add domain-specific tests rather than relying on a public leaderboard. A Malayalam document workflow may need a separate extraction evaluation, such as this guide to Malayalam document extraction, because instruction following and reliable field extraction measure different capabilities.

    Practical checklist

    • Pin model, dataset, evaluator, and dependency revisions.
    • Use the official Malayalam split and prompt template.
    • Evaluate generation models as generators, not classifiers.
    • Keep raw outputs and log every decoding parameter.
    • Score with IndicEval’s prescribed method and report task-level results.
    • Manually inspect Malayalam failures and code-switching cases.
    • Publish exclusions, confidence estimates, and reproducible commands.

    A benchmark becomes useful when another team can reproduce it and understand why a model succeeded or failed. Treat IndicEval as a controlled measurement protocol, not merely a score-producing script, and your Malayalam results will be far more actionable for model selection, fine-tuning, and deployment.

    Last updated 23 September 2026

AIGI may be inaccurate. Replies seeded from the guide above.