Bengali question answering (QA) benchmarks need more than a model score. Dataset quality, answer-span alignment, Bengali text normalization, script variants, and post-processing can change results substantially. This guide presents a reproducible workflow for benchmarking extractive Bengali QA with Hugging Face Datasets and Transformers, while also showing how to make comparisons useful for Indian-language AI teams.
The focus is answer extraction: given a question and a context, the model identifies the answer span in the context. Generative QA requires additional metrics and a different evaluation pipeline, so do not combine the two task types in one leaderboard without clearly labelling them.
Define the benchmark before writing code
Start with a short benchmark specification. Record:
- Task: extractive, generative, or both.
- Language policy: Bengali script only, or Bengali plus transliterated and code-mixed inputs.
- Evaluation unit: question-level scores, not aggregate token counts.
- Model access: public checkpoints, fine-tuned checkpoints, or zero-shot systems.
- Compute limits: GPU type, batch size, maximum sequence length, and inference precision.
- Reproducibility: dataset revision, model revision, package versions, random seeds, and decoding settings.
A benchmark should also state whether it measures general Bengali comprehension or a specific Indian use case such as education, government information, healthcare, or search. For a broader comparison across Indian languages, use the principles in this Indian-language LLM benchmark guide and keep Bengali-specific results visible rather than hiding them inside one multilingual average.
Choose and audit a Hugging Face dataset
Hugging Face Datasets makes loading and versioning straightforward, but a dataset card is not a substitute for an audit. Inspect the schema, licences, language distribution, question types, and train-validation-test construction before treating a dataset as a benchmark.
A typical extractive QA record contains:
id: a stable question identifier.question: the user query.context: the passage containing the answer.answers: one or more answer strings and their character offsets.
Load a public dataset with an explicit revision where possible:
from datasets import load_dataset
qa = load_dataset(
"dataset-owner/dataset-name",
revision="commit-or-tag"
)
print(qa)
print(qa["train"].features)Do not assume that a dataset described as Bengali is linguistically uniform. Check for duplicated contexts, repeated questions, machine-translated passages, spelling variants, and examples where the annotated answer is not actually present at the supplied offset. Remove or quarantine broken records, but publish the filtering rules and counts.
Prevent leakage by grouping near-duplicate contexts across splits. If the same article appears in both training and test data, a high score may reflect memorisation rather than reading comprehension. Keep a locked test set that is not used for model selection.
Normalise Bengali carefully
Unicode handling is a major source of misleading Bengali QA results. Bengali text can contain different Unicode representations, punctuation styles, zero-width characters, whitespace patterns, and visually similar spellings. Create a documented normalisation function for analysis and metric computation, but preserve the original strings for auditability.
A sensible normalisation policy may include:
- Unicode normalisation such as NFC.
- Consistent whitespace handling.
- Conservative punctuation treatment.
- Removal of accidental zero-width characters where justified.
- Explicit handling of Bengali and Arabic numerals.
- No aggressive stemming or spelling correction unless the benchmark defines it.
Exact Match should be reported under a clearly named policy, such as EM_raw and EM_normalized. This prevents teams from presenting a normalised score as if it were a literal string match. Keep multiple gold answers when annotators use valid variants.
Prepare features and labels correctly
For long contexts, tokenization can create multiple overlapping windows. The answer may appear in one feature but not another, so preserve the mapping between features and original examples using overflow_to_sample_mapping and offset_mapping.
from transformers import AutoTokenizer
tokenizer = AutoTokenizer.from_pretrained("model-name", use_fast=True)
features = tokenizer(
qa["validation"]["question"],
qa["validation"]["context"],
max_length=384,
truncation="only_second",
stride=128,
return_overflowing_tokens=True,
return_offsets_mapping=True,
padding="max_length",
)For each feature, identify the context token range using sequence_ids. Convert the answer’s character start and end into token positions only when the complete span falls inside that window. If it does not, mark the feature appropriately rather than silently assigning an incorrect label.
When benchmarking an existing checkpoint, use the same maximum length, stride, and answer-span post-processing across models. Otherwise, the benchmark measures preprocessing choices as much as model quality.
Run a reproducible baseline
Establish at least two baselines:
- A simple heuristic, such as selecting a sentence containing question keywords.
- A pretrained or fine-tuned Bengali/multilingual extractive QA checkpoint.
Record model name, revision, tokenizer, hardware, inference time, peak memory, and number of parameters. A multilingual model may perform strongly overall but fail on Bengali-specific spelling, named entities, or long contexts. Compare it with a Bengali-focused checkpoint where licensing and data provenance permit.
Use deterministic settings for the first run. Set seeds, disable sampling, fix batch size, and retain raw start and end logits. For production-style benchmarking, run a second configuration with realistic latency constraints and report the trade-off rather than only the highest score.
Score with EM and token-level F1
The standard extractive QA metrics are:
- Exact Match (EM): whether the predicted answer equals at least one gold answer after the declared normalisation.
- Token-level F1: overlap between predicted and gold tokens, combining precision and recall.
Use the current evaluate package rather than older datasets.load_metric examples:
import evaluate
squad_metric = evaluate.load("squad")
results = squad_metric.compute(
predictions=[{"id": "1", "prediction_text": "..."}],
references=[{
"id": "1",
"answers": {"text": ["..."], "answer_start": [0]}
}]
)
print(results)Validate the expected format for your installed metric version. Bengali tokenisation can make token-level F1 sensitive to whitespace and punctuation decisions, so publish the normaliser and, ideally, a small scoring test suite. Report overall EM and F1, then break results down by answer length, context length, question type, and source domain.
For generated answers, do not reuse extractive EM and F1 without qualification. Add semantic or factual evaluation, human review, and citation or entailment checks. A useful QA application should also be assessed for abstention: can it say that the context does not contain enough information?
Add error analysis, not just a leaderboard
Inspect at least 100 incorrect predictions, stratified by score and question type. Label failure modes such as:
- Wrong passage window.
- Correct entity but incorrect boundary.
- Bengali spelling or Unicode mismatch.
- Confusion between similar names or dates.
- Multi-sentence reasoning failure.
- Unanswerable question answered with a guess.
- Context contamination or annotation error.
Create a compact error table with the question, context identifier, gold answer, prediction, normalisation result, and failure label. This turns a benchmark into an engineering backlog. Teams building student-facing systems can also compare the findings with guidance on AI question-answering apps for Indian students, especially for age-appropriate explanations and safe abstention.
Report results responsibly
A credible report should include dataset revision, filtering criteria, split sizes, normalisation rules, model and tokenizer revisions, hardware, runtime, and confidence intervals or bootstrap estimates where feasible. Avoid ranking models on tiny or contaminated test sets. Publish predictions for audit only when dataset and model licences allow it.
For multilingual deployments, pair Bengali results with comparable evaluations for other Indian languages. The practical framework for benchmarking multilingual LLMs in India is useful for selecting common latency, robustness, and fairness measures without erasing language-specific issues.
Recommended benchmark checklist
Before releasing a result, confirm that you have:
- Versioned the dataset and model.
- Locked the test split and checked for leakage.
- Audited answer offsets and duplicate contexts.
- Documented Bengali Unicode normalisation.
- Used consistent tokenisation and post-processing.
- Reported EM, F1, latency, and resource use.
- Included subgroup and error analysis.
- Released scripts, configuration, and environment details.
Benchmarking Bengali QA on Hugging Face is straightforward to start but difficult to do credibly. The strongest result is not merely a larger F1 number; it is a result that another Indian-language AI team can reproduce, inspect, and use to make a better model.