Why this benchmark matters
Malayalam instruction following needs more than a single overall score. A model may answer factual questions well but fail on formatting, safety, multi-step reasoning, or instructions written in colloquial Malayalam. A reproducible benchmark helps you distinguish genuine capability from prompt sensitivity, translation artefacts, and evaluation noise.
This guide presents a practical workflow for benchmarking Malayalam instruction following with IndicEval and Hugging Face. Before running experiments, confirm the current IndicEval task name, dataset configuration, split, licence, and evaluator version in the official project documentation or model repository. Tool APIs and benchmark packaging can change; do not assume that an indicieval Python package or a particular class exists simply because a benchmark is associated with Hugging Face.
For broader dataset-selection principles, see this Indian language LLM benchmark datasets guide. If you are comparing several Indic languages, the same workflow can be extended using a practical framework for benchmarking multilingual LLMs in India.
Define the evaluation before installing anything
Write down the evaluation contract first. Record:
- The exact model revision, tokenizer revision, and quantisation setting
- The IndicEval task, dataset configuration, split, and evaluator commit
- Whether the model is evaluated zero-shot, few-shot, or after Malayalam-specific fine-tuning
- The prompt template, system message, number of demonstrations, and language used for demonstrations
- Decoding settings, including temperature, top-p, maximum new tokens, and stop sequences
- Hardware, software versions, random seed, and batch size
Instruction following is usually a generation task, not ordinary sequence classification. The earlier pattern of loading AutoModelForSequenceClassification and taking logits.argmax() is unsuitable unless the benchmark explicitly defines a classification head and labels. For generative models, use AutoModelForCausalLM or the architecture required by the checkpoint, then score the generated answer against the benchmark’s references or rubric.
Set up a reproducible Hugging Face environment
Create an isolated environment and pin dependencies. A typical starting point is:
python -m venv .venv
source .venv/bin/activate
pip install -U "transformers>=4.45" "datasets>=2.20" "accelerate" "evaluate" "sentencepiece" "safetensors"Install any IndicEval-specific package only when its documentation requires it. Some benchmarks are distributed as datasets, evaluation scripts, or leaderboard repositories rather than a package named indicieval.
Authenticate with Hugging Face if the model or dataset is gated:
huggingface-cli loginThen load the model with an explicit revision. For a causal language model:
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "your-org/your-malayalam-capable-model"
revision = "main" # replace with a commit hash for archival runs
tokenizer = AutoTokenizer.from_pretrained(model_id, revision=revision)
model = AutoModelForCausalLM.from_pretrained(
model_id,
revision=revision,
torch_dtype=torch.bfloat16,
device_map="auto",
)
model.eval()Use torch.float16 when your accelerator does not support bfloat16. Check the model card for chat-template requirements; applying the wrong template can change results substantially.
Load and inspect the Malayalam data
Use the dataset identifier and configuration specified by IndicEval rather than guessing:
from datasets import load_dataset
dataset = load_dataset(
"OWNER/INDIC-EVAL-DATASET",
"malayalam_instruction_following",
revision="DATASET_COMMIT_OR_TAG",
)
print(dataset)
print(dataset["test"][0])Inspect field names, language labels, reference answers, metadata, and any hidden answer columns. Keep the official test set untouched. If you need local development examples, create a separate validation subset and document how it was sampled.
Before evaluation, check for:
- Mixed Malayalam-English prompts and code-switching
- Unicode normalisation differences, especially combining marks and punctuation
- Duplicate or near-duplicate prompts across splits
- Unsafe, ambiguous, or culturally specific requests
- References that contain multiple acceptable answers
- Instructions requiring tables, JSON, lists, or exact strings
Do not silently transliterate Malayalam into Latin script or translate prompts into English. Those can be useful auxiliary experiments, but they are not equivalent to native Malayalam instruction following.
Build a controlled generation loop
Use the benchmark’s required prompt format. If the model has a chat template, prefer it over manually concatenating role labels:
def make_prompt(example):
messages = [{"role": "user", "content": example["prompt"]}]
return tokenizer.apply_chat_template(
messages, tokenize=False, add_generation_prompt=True
)
def generate_one(example):
prompt = make_prompt(example)
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
with torch.inference_mode():
output = model.generate(
**inputs,
max_new_tokens=512,
do_sample=False,
pad_token_id=tokenizer.eos_token_id,
)
new_tokens = output[0, inputs["input_ids"].shape[1]:]
return tokenizer.decode(new_tokens, skip_special_tokens=True).strip()For headline scores, deterministic decoding makes runs easier to compare. For robustness, repeat a smaller sample with controlled temperatures and report score variation. Ensure that the prompt is not included in the decoded answer, and save raw outputs before any cleaning.
Score the right behaviours
Follow IndicEval’s official scoring implementation whenever available. Instruction-following evaluation may combine exact match, normalised match, task-specific validators, reference-based similarity, or rubric-based judging. Do not replace these with generic accuracy or F1 without clearly labelling the result.
Report results by task category, not only as one Malayalam average. Useful slices include:
- Instruction adherence and required format
- Factual question answering
- Summarisation and rewriting
- Reasoning or multi-step execution
- Safety and refusal behaviour
- Code-switching and dialect variation
- Prompt length and difficulty
If an LLM judge is used, publish the judge model, prompt, language, sampling settings, and agreement checks. Malayalam answers should be judged by fluent evaluators or validated bilingual protocols; an English-only judge can reward translated-looking output and miss grammatical or cultural errors.
Analyse failures and publish a credible report
Save one record per example containing the prompt hash, model revision, generation settings, output, reference, score, and error label. Review a stratified sample of failures manually. Common labels include wrong task, incomplete answer, unsupported claim, formatting violation, language switch, refusal error, and hallucination.
Report the number of examples, excluded items, confidence intervals or bootstrap ranges where practical, and separate development decisions from final test results. Compare against strong baselines: a Malayalam-capable open model, a multilingual model, and—if relevant—a translated-prompt baseline. Results from other Indic languages can provide useful context; see benchmarking NLP models for Telugu and Sanskrit.
For production use, add domain-specific tests rather than relying on a public leaderboard. A Malayalam document workflow may need a separate extraction evaluation, such as this guide to Malayalam document extraction, because instruction following and reliable field extraction measure different capabilities.
Practical checklist
- Pin model, dataset, evaluator, and dependency revisions.
- Use the official Malayalam split and prompt template.
- Evaluate generation models as generators, not classifiers.
- Keep raw outputs and log every decoding parameter.
- Score with IndicEval’s prescribed method and report task-level results.
- Manually inspect Malayalam failures and code-switching cases.
- Publish exclusions, confidence estimates, and reproducible commands.
A benchmark becomes useful when another team can reproduce it and understand why a model succeeded or failed. Treat IndicEval as a controlled measurement protocol, not merely a score-producing script, and your Malayalam results will be far more actionable for model selection, fine-tuning, and deployment.