What this benchmark should measure
Benchmarking Bengali instruction following is not the same as checking whether a model can generate fluent Bengali. The central question is whether it can understand a Bengali instruction, satisfy every constraint, preserve the requested format, and avoid unsupported claims. A useful evaluation therefore combines task success with language quality and safety checks.
IndicEval can provide a consistent evaluation framework, while Hugging Face supplies the model, tokenizer, dataset tooling, and experiment infrastructure. Before you begin, confirm the exact IndicEval release, task name, dataset split, language code, and scoring protocol. Toolkits and benchmark schemas can change; do not assume that a generic pip install indic-eval command or an invented CLI is valid for your installation.
For wider context on selecting evaluation data, see this Indian language LLM benchmark datasets guide. It helps distinguish instruction-following tests from translation, classification, and general language-understanding benchmarks.
Define the evaluation contract
Write down the benchmark contract before running a model. At minimum, record:
- Language: Bengali (
bn), including whether the test permits code-switching with English or Hindi. - Task types: question answering, summarisation, extraction, rewriting, classification, refusal, and multi-constraint instructions.
- Expected output: free text, JSON, a label, a numbered list, or another strict format.
- Scoring: exact match, reference-based similarity, rubric-based judgement, or a combination.
- Generation settings: model revision, prompt template, maximum tokens, temperature, top-p, and stop conditions.
- Hardware and precision: GPU type, quantisation method, batch size, and inference library.
This contract prevents misleading comparisons. A model evaluated with a permissive chat template and another evaluated with raw user text are not being tested under identical conditions. For a broader methodology, compare your design with this practical framework for benchmarking multilingual LLMs in India.
Prepare IndicEval and Hugging Face
Use an isolated Python environment and pin the versions used for the run. The package name, installation method, and command-line interface should come from the official IndicEval documentation or repository for the release you selected.
python -m venv .venv
source .venv/bin/activate # Windows: .venv\\Scripts\\activate
python -m pip install --upgrade pip
pip install torch transformers datasets accelerate evaluate
# Install IndicEval using its official release instructions.Log the following before inference:
import platform
import torch
import transformers
import datasets
print(platform.platform())
print(torch.__version__)
print(transformers.__version__)
print(datasets.__version__)
print(torch.cuda.get_device_name(0) if torch.cuda.is_available() else "CPU")Load a compatible causal or sequence-to-sequence model with AutoTokenizer and the matching AutoModelForCausalLM or AutoModelForSeq2SeqLM class. Verify the model card for Bengali support, licence terms, context length, recommended chat template, and known limitations. A Bengali-capable small model may be easier to run locally; this guide to Bengali small language models is useful when choosing a baseline.
Validate the Bengali test set
Do not evaluate a dataset merely because its filename says Bengali. Inspect the records manually and programmatically. Each example should make the instruction, any context, and the expected answer unambiguous.
Check for:
- Unicode normalisation, Bengali punctuation, zero-width characters, and accidental Latin-script substitutions.
- Duplicate prompts, near-duplicate references, and train-test contamination.
- Correct handling of Bengali numerals, dates, names, honorifics, and transliterated terms.
- Balanced task categories and difficulty levels.
- Clear separation of public examples from hidden test data.
- References that allow legitimate variations in wording.
Keep a small, locked diagnostic set containing cases such as negation, multiple constraints, long context, tables, code-mixed instructions, and requests requiring refusal. Never tune prompts against the hidden test set. Store the dataset revision or commit hash so another researcher can reproduce the exact split.
Run inference consistently
Use the model’s documented chat template whenever one exists. A simplified Hugging Face pattern is:
from datasets import load_dataset
from transformers import AutoTokenizer, AutoModelForCausalLM
import torch
model_id = "YOUR_MODEL_ID"
tokenizer = AutoTokenizer.from_pretrained(model_id, use_fast=True)
model = AutoModelForCausalLM.from_pretrained(
model_id,
torch_dtype="auto",
device_map="auto",
)
example = {"instruction": "YOUR_BENGALI_INSTRUCTION"}
messages = [{"role": "user", "content": example["instruction"]}]
prompt = tokenizer.apply_chat_template(
messages, tokenize=False, add_generation_prompt=True
)
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
with torch.no_grad():
output = model.generate(
**inputs,
max_new_tokens=256,
do_sample=False,
)
answer = tokenizer.decode(
output[0][inputs["input_ids"].shape[-1]:],
skip_special_tokens=True,
)
print(answer)Use deterministic decoding for the primary score. If you also report sampling results, publish the seed and generation parameters. Save raw outputs, prompts, token counts, latency, and failures—not just aggregate scores. This makes it possible to distinguish model errors from truncation, malformed templates, or out-of-memory retries.
Score instruction following, not just fluency
Run IndicEval’s official evaluator for the selected task and language, then supplement it with targeted checks where appropriate. Do not substitute generic accuracy, precision, recall, and F1 for an instruction-following score unless the task is genuinely classification-based.
Useful dimensions include:
- Task success: Does the answer solve the requested problem?
- Constraint adherence: Did it follow length, ordering, persona, language, and formatting requirements?
- Factuality and groundedness: Does it stay within the supplied context?
- Bengali quality: Is the script readable, grammatical, and natural for the intended audience?
- Safety and refusal quality: Does it refuse disallowed requests without refusing benign ones?
- Format validity: Does JSON parse, do labels match the schema, and are required fields present?
For open-ended outputs, use a Bengali-capable human rubric or a carefully validated judge model. Judge prompts should include the instruction, response, rubric, and explicit scoring anchors. Audit a sample manually because automatic judges can reward verbosity, favour familiar dialects, or miss subtle Bengali meaning errors. Report category-level results, not only one overall number.
Analyse failures and report reproducibly
Create an error taxonomy: misunderstood instruction, missed constraint, hallucination, inadequate Bengali, code-switching, formatting failure, unsafe completion, over-refusal, context loss, and generation truncation. Review examples from every category and separate systematic errors from isolated mistakes.
Report at least:
- Model identifier and immutable revision.
- IndicEval version, dataset revision, split, and language settings.
- Prompt or chat-template format.
- Decoding parameters and hardware.
- Per-task and per-difficulty scores, sample counts, and confidence intervals where feasible.
- Invalid-output and runtime-failure rates.
- Representative Bengali examples with English explanations only when needed.
If your use case is offline or resource-constrained, test the same benchmark with a quantised model and disclose the quantisation method; this offline Bengali model guide covers practical deployment trade-offs. For speech-driven systems, keep transcription quality separate from instruction-following quality—see this speech-to-text benchmarking guide for India.
Turn results into engineering decisions
Use the error breakdown to choose the next experiment. Fine-tune only after confirming that the problem is not prompt formatting or data contamination. For constraint failures, improve instruction templates and add structured decoding. For domain errors, curate representative Bengali data and evaluate on a held-out domain split. For weak fluency, inspect tokenisation, script coverage, and regional variation before increasing model size.
A credible benchmark is a repeatable measurement system, not a one-time leaderboard number. Freeze the protocol, publish the artefacts that can be shared, preserve raw outputs securely, and rerun the diagnostic set after every model, prompt, or data change. This gives Indian-language builders a result they can trust—and a clear path from score to product improvement.