What to benchmark—and why
Benchmarking BharatGPT models is not just a race for the highest score. For an Indian-language application, the useful question is whether a model is accurate, culturally appropriate, safe, fast, and affordable for your target users. A model that performs well on Hindi news summarisation may still struggle with code-mixed Marathi, Telugu names, noisy speech transcripts, or low-bandwidth production traffic.
Start by writing a short evaluation brief:
- Languages and varieties: Hindi, Bengali, Tamil, Telugu, Marathi, Sanskrit, Hinglish, or regional dialects.
- Tasks: generation, translation, summarisation, question answering, classification, extraction, or chat.
- Operating limits: maximum latency, memory, GPU type, request volume, and cost per request.
- Risk level: public information, education, financial support, healthcare, or government services.
- Success criteria: minimum quality score plus hard limits for hallucination, refusal errors, latency, and throughput.
If you are comparing BharatGPT with smaller Hindi-focused models, establish the language and data coverage first. The practical considerations in open-source small language models for Hindi are useful when choosing a baseline.
Select models and pin the test setup
On the Hugging Face Hub, record the exact model repository, revision or commit, tokenizer, quantisation setting, and inference library. Do not benchmark a moving main branch and later assume the result is reproducible. Save the model card, license, parameter count, context window, and any stated training-data limitations.
Create an isolated environment and install the core tooling:
python -m venv .venv
source .venv/bin/activate
pip install -U torch transformers datasets evaluate accelerate sacrebleu rouge-score pandasLoad a causal language model with its matching tokenizer. Repository names vary, so replace the placeholder with the approved model ID:
from transformers import AutoTokenizer, AutoModelForCausalLM
model_id = "ORG_OR_USER/MODEL_NAME"
tokenizer = AutoTokenizer.from_pretrained(model_id, revision="COMMIT_OR_TAG")
model = AutoModelForCausalLM.from_pretrained(
model_id,
revision="COMMIT_OR_TAG",
device_map="auto",
torch_dtype="auto"
)Keep generation settings explicit. Greedy decoding, beam search, and sampling answer different questions; mixing them invalidates comparisons. Fix the random seed where sampling is used, and run enough examples to report variation rather than one impressive output.
Build an India-relevant evaluation set
Use a held-out test set that reflects production inputs, not only popular academic datasets. Keep training, development, and test examples separate. A balanced evaluation set might include:
- Native-script prompts and Romanised text.
- Code-mixed queries such as Hinglish and Tamil-English.
- Formal, conversational, rural, urban, and domain-specific language.
- Names, addresses, dates, currency, measurements, and Indian institutions.
- Long context, short queries, misspellings, and noisy user input.
- Adversarial prompts, sensitive topics, and requests requiring refusal.
Check licensing and consent before uploading private or customer data to the Hub. Redact phone numbers, Aadhaar-like identifiers, health information, and other personal data. For Telugu and Sanskrit-specific comparisons, benchmarking NLP models for Telugu and Sanskrit offers a useful starting point for task and language coverage.
Store each example as structured data with fields such as id, language, script, task, prompt, reference, domain, and risk_tag. Stratification lets you report where a model succeeds instead of hiding weak performance inside one average.
Choose metrics that match the task
No single metric captures multilingual generation quality. Combine automatic measures with human review:
- Perplexity: useful for comparing next-token prediction on the same corpus, but not a direct measure of helpfulness or factuality.
- BLEU and chrF: useful for translation, especially when character-level variation matters; report the tokenizer and normalisation process.
- ROUGE: a rough signal for summarisation overlap, not a substitute for checking omissions and factual errors.
- Exact match, accuracy, and F1: appropriate for classification, extraction, and answerable question sets.
- Semantic similarity: helpful for paraphrases, but validate embedding quality across the target Indian languages.
- Human ratings: score correctness, relevance, fluency, script fidelity, cultural fit, and harmful or unsupported claims.
For generative systems, add faithfulness, refusal quality, toxicity, prompt-injection resistance, and language leakage checks. Have bilingual reviewers assess a statistically meaningful sample. Give them blinded model outputs and a clear rubric; otherwise reviewers may favour a familiar model or writing style.
Measure speed, memory, and cost
Quality without production measurements is incomplete. Warm up the model, then measure at least 50-100 requests per scenario where possible. Report median and p95 latency, tokens per second, time to first token, peak memory, throughput, and failure rate. Separate prompt length from generated-token length because both affect cost and latency.
A simple generation harness can capture repeatable outputs:
import time, torch
def generate_one(prompt, max_new_tokens=128):
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
start = time.perf_counter()
with torch.inference_mode():
output = model.generate(
**inputs,
max_new_tokens=max_new_tokens,
do_sample=False,
pad_token_id=tokenizer.eos_token_id,
)
elapsed = time.perf_counter() - start
new_tokens = output.shape[-1] - inputs["input_ids"].shape[-1]
return {
"text": tokenizer.decode(output[0], skip_special_tokens=True),
"seconds": elapsed,
"new_tokens": new_tokens,
"tokens_per_second": new_tokens / max(elapsed, 1e-9),
}Repeat the test on the hardware you expect to deploy. A quantised model running locally may be a better fit than a larger model on an expensive GPU. If deployment is the next step, compare the benchmark assumptions with how to deploy large language models locally or a managed option such as deploying deep learning models on GKE.
Analyse, document, and publish results
Create a results table with one row per model, language, task, decoding configuration, hardware, and dataset slice. Include confidence intervals or bootstrap estimates for quality scores where feasible. Inspect failures manually and group them into actionable categories: untranslated spans, script confusion, hallucinated citations, poor numeracy, unsafe advice, or excessive verbosity.
Publish the evaluation code, dataset card, prompt templates, environment lockfile, and model revisions. Do not publish restricted test data or personal information. A clear model card should state what the benchmark does not prove—for example, a good ROUGE score does not establish factual accuracy or suitability for public deployment.
A practical decision rule
Choose a model only after it passes all non-negotiable gates. For example: minimum bilingual quality, zero critical privacy failures, acceptable refusal behaviour, p95 latency under your product limit, and a cost ceiling per 1,000 requests. Among models that pass, select the one with the best combination of quality, reliability, and operational simplicity—not necessarily the largest checkpoint.
Re-run the benchmark whenever you change the model revision, prompt, tokenizer, quantisation, serving stack, or test distribution. This turns benchmarking into an engineering control rather than a one-time leaderboard exercise.