IndicBERT is useful when your NLP system must handle Indian languages, but a credible benchmark requires more than calling trainer.evaluate(). You need a fixed task definition, language-balanced data, reproducible training settings, and metrics that expose both average quality and weak-language behaviour.
This guide shows how to benchmark IndicBERT on Hugging Face for classification and related encoder tasks. The same principles apply when comparing it with multilingual BERT, MuRIL, IndicBERT variants, or another regional-language model. For a broader view of dataset selection, see this Indian language LLM benchmark datasets guide.
What you should benchmark
Start by deciding whether you are measuring fine-tuning quality, zero-shot transfer, or inference efficiency. These are different experiments and should not be mixed.
Useful IndicBERT benchmark tasks include:
- Text classification: sentiment, topic, intent, toxicity, or news classification.
- Named entity recognition: person, organisation, location, and domain-specific entities.
- Natural language inference: entailment and contradiction across Indian languages.
- Retrieval or reranking: only if the model and task formulation support meaningful sentence representations.
- Production efficiency: latency, throughput, peak memory, and model size.
Use a dataset with a documented language, script, label scheme, licence, and train/dev/test split. Avoid reporting results from a test set repeatedly during development. If your project covers Telugu or Sanskrit, compare your protocol with this guide to benchmarking NLP models for Telugu and Sanskrit.
Set up a reproducible Hugging Face environment
Use a current Python environment and pin the major packages. API names change over time, so record the exact versions used for the run.
python -m venv .venv
source .venv/bin/activate
pip install -U torch transformers datasets evaluate accelerate scikit-learn sentencepieceCapture the hardware and software configuration in your experiment log:
- Python, PyTorch, Transformers, Datasets, and CUDA versions
- GPU model, VRAM, CPU, and RAM
- IndicBERT checkpoint and revision
- Dataset version or commit
- Random seed and maximum sequence length
- Batch size, gradient accumulation, learning rate, epochs, and warm-up schedule
IndicBERT checkpoints may use SentencePiece and model-specific configuration. Do not assume that a generic BertTokenizer or BertForSequenceClassification is always correct. Inspect the checkpoint card and load the architecture supported by that checkpoint.
Load and inspect the dataset
For a classification dataset with text and label columns:
from datasets import load_dataset
raw = load_dataset("your-org/your-indic-dataset")
print(raw)
print(raw["train"].features)
print(raw["train"].to_pandas()["label"].value_counts())Before tokenisation, check for duplicated examples, empty strings, inconsistent Unicode, mixed scripts, and accidental overlap between splits. Normalise only when justified: aggressive normalisation can remove distinctions that matter in Indic languages. Report whether punctuation, diacritics, emojis, transliterated text, and code-mixed text were retained.
If the dataset lacks a validation split, create one with stratification where possible. Keep the test split untouched until the final run. For a multilingual benchmark, report results per language, not only a pooled score; a large Hindi subset can hide poor performance on smaller languages.
Tokenise without leaking information
Load the tokenizer from the exact checkpoint you intend to evaluate. Use dynamic padding during training where practical, and cap sequence length based on the observed length distribution rather than selecting an unnecessarily large value.
from transformers import AutoTokenizer
tokenizer = AutoTokenizer.from_pretrained("ai4bharat/indic-bert")
def tokenize(batch):
return tokenizer(
batch["text"],
truncation=True,
max_length=256,
)
tokenized = raw.map(tokenize, batched=True, remove_columns=["text"])Measure truncation rates. A model that appears weak may simply be losing the decisive part of long inputs. Conversely, increasing the limit raises memory use and latency, so treat sequence length as a benchmark variable rather than silently changing it between models.
Fine-tune IndicBERT with fixed settings
The following template is suitable for sequence classification. Confirm the correct model class for your checkpoint before running it.
import numpy as np
import evaluate
from transformers import (
AutoModelForSequenceClassification,
DataCollatorWithPadding,
TrainingArguments,
Trainer,
)
checkpoint = "ai4bharat/indic-bert"
model = AutoModelForSequenceClassification.from_pretrained(
checkpoint,
num_labels=2,
)
collator = DataCollatorWithPadding(tokenizer=tokenizer)
accuracy = evaluate.load("accuracy")
f1 = evaluate.load("f1")
def compute_metrics(eval_pred):
logits, labels = eval_pred
predictions = np.argmax(logits, axis=-1)
return {
"accuracy": accuracy.compute(
predictions=predictions, references=labels)["accuracy"],
"f1_macro": f1.compute(
predictions=predictions, references=labels, average="macro")["f1"],
}
args = TrainingArguments(
output_dir="runs/indicbert",
eval_strategy="epoch",
save_strategy="no",
learning_rate=2e-5,
per_device_train_batch_size=16,
per_device_eval_batch_size=32,
num_train_epochs=3,
weight_decay=0.01,
seed=42,
report_to="none",
)
trainer = Trainer(
model=model,
args=args,
train_dataset=tokenized["train"],
eval_dataset=tokenized["validation"],
tokenizer=tokenizer,
data_collator=collator,
compute_metrics=compute_metrics,
)
trainer.train()
print(trainer.evaluate(tokenized["test"]))For imbalanced labels, macro-F1, per-class recall, and a confusion matrix are more informative than accuracy alone. For NER, use entity-level precision, recall, and F1 with a library such as seqeval; token-level accuracy can be misleading because most tokens are often non-entities.
Run at least three seeds for a publishable comparison. Report mean and standard deviation, not just the best run. Keep the hyperparameter budget equal across competing models. Giving IndicBERT three trials and another model twenty trials is not a fair benchmark.
Compare quality and efficiency separately
A useful report has two tables. The first covers quality:
- Overall score and confidence interval where feasible
- Per-language and per-class metrics
- Results by script, text length, and code-mixing level
- Seed mean and standard deviation
- Number of training examples and labels
The second covers deployment cost:
- Checkpoint size on disk
- Peak GPU or CPU memory
- Median and p95 latency
- Throughput at the intended batch size
- Cold-start and warm inference time
Measure latency with warm-up iterations and identical tokenised inputs. State whether timings include tokenisation. For CPU deployment, test the hardware your Indian users will actually receive rather than quoting only a high-end GPU result. If the benchmark feeds a downstream application, the guidance in building Python-based natural language interfaces can help connect model metrics to system-level behaviour.
Avoid common benchmark errors
- Using the training set for evaluation: this produces inflated scores.
- Changing preprocessing by model: tokenisation and normalisation must be documented and comparable.
- Pooling languages without language-level results: aggregate scores can conceal failures.
- Reporting one lucky seed: small datasets often produce substantial variance.
- Ignoring label imbalance: accuracy may reward a majority-class predictor.
- Comparing different truncation limits: longer inputs can change both quality and cost.
- Treating a checkpoint warning as harmless: newly initialised classification heads must be fine-tuned before evaluation.
- Calling a fine-tuned score zero-shot: clearly label the evaluation setting.
Publish a benchmark others can reproduce
Release the dataset identifier, preprocessing code, checkpoint name, training configuration, random seeds, evaluation script, and raw predictions where licensing permits. Include errors, not only headline scores: examples involving spelling variation, transliteration, dialect, and code-mixing are particularly valuable for Indian-language systems.
For a production decision, choose the model that meets the required quality and latency threshold—not automatically the model with the highest aggregate F1. A careful IndicBERT benchmark gives builders a defensible basis for selecting a checkpoint, identifying language gaps, and deciding whether additional labelled data or task-specific adaptation will deliver the biggest improvement.