What this benchmark should answer
A useful benchmark should do more than produce one BLEU number. It should tell you whether a Hugging Face model translates Urdu to another language, another language to Urdu, or both, how it behaves across domains, and whether improvements survive changes in decoding settings.
FLORES-101 and FLORES-200 provide carefully constructed multilingual test sets, making them useful for controlled comparisons. They are not a complete picture of Urdu in India or Pakistan: FLORES sentences are relatively formal, short, and carefully curated. Treat the result as a standardised test, then add an India-relevant evaluation set containing code-mixing, names, administrative language, Roman Urdu, and regional usage.
For broader context, compare your design with this practical framework for benchmarking multilingual LLMs in India and the 2026 guide to Indian-language benchmark datasets.
Choose the exact FLORES task
First define the language direction and dataset version. Urdu is commonly represented by the ISO-style code urd. A benchmark might evaluate eng_Latn -> urd_Arab or urd_Arab -> eng_Latn. Do not combine directions in one score: translation difficulty is asymmetric, and Urdu script generation introduces its own tokenisation and orthographic issues.
Record these details in a configuration file or experiment log:
- FLORES version and dataset split, normally
devfor development anddevtestfor final reporting. - Source and target language codes, including script where FLORES-200 uses them.
- Model repository and exact revision or commit.
- Tokeniser revision, generation parameters, hardware, and software versions.
- Whether the model was prompted, forced to use a language token, or given a task prefix.
Use dev to test your pipeline and reserve devtest for the final comparison. Never tune decoding parameters repeatedly on devtest; that turns the test set into a hidden training signal.
Install a reproducible Hugging Face stack
Create an isolated environment and install the libraries needed for loading data, running generation, and scoring:
python -m venv .venv
source .venv/bin/activate
pip install -U torch transformers datasets evaluate sacrebleu sentencepiece unbabel-cometThe exact model class varies. AutoModelForSeq2SeqLM works for many encoder-decoder systems, while some multilingual checkpoints need a language-code setting or a model-specific processor. Check the model card before assuming that a checkpoint supports Urdu or the requested direction.
Load FLORES through datasets where possible rather than manually copying files:
from datasets import load_dataset
flores = load_dataset("facebook/flores", "eng_Latn-urd_Arab")
test = flores["devtest"]
print(test.column_names)Dataset configuration names can differ between releases. If this identifier is unavailable, inspect the configurations with get_dataset_config_names and select the pair matching your installed version. Keep the raw dataset unchanged and create a separate processed file with stable row IDs.
Run translation without accidental inconsistency
A benchmark is only comparable when every model receives the same inputs and uses documented generation settings. Set deterministic decoding for the main run, then perform a separate sensitivity test with beam search or sampling if your application requires it.
import torch
from transformers import AutoTokenizer, AutoModelForSeq2SeqLM
model_id = "your-supported-multilingual-checkpoint"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForSeq2SeqLM.from_pretrained(model_id).eval()
def translate(texts):
batch = tokenizer(texts, return_tensors="pt", padding=True,
truncation=True, max_length=512)
with torch.no_grad():
output = model.generate(
**batch, max_new_tokens=256, num_beams=5,
do_sample=False
)
return tokenizer.batch_decode(output, skip_special_tokens=True)For models such as mBART or NLLB, set the source language and target language exactly as the model documentation specifies. For example, a target language token may need to be passed through forced_bos_token_id. Omitting it can produce fluent text in the wrong language, invalidating the benchmark.
Save predictions, references, source text, model metadata, and failures as JSONL or Parquet. Do not retain only aggregate scores: sentence-level outputs are essential for debugging.
Score with complementary metrics
Install and use SacreBLEU so that metric tokenisation and signatures are recorded consistently. BLEU remains useful for comparison with earlier literature, but it is weak at recognising valid paraphrases and can be unstable for small test sets.
import sacrebleu
bleu = sacrebleu.corpus_bleu(predictions, [references])
chrf = sacrebleu.corpus_chrf(predictions, [references])
print(bleu.score, bleu.signature)
print(chrf.score, chrf.signature)Report at least:
- SacreBLEU, with its signature and number of references.
- chrF or chrF++, often informative for morphologically rich languages and spelling variation.
- COMET, when the model and runtime are available; it uses learned evaluation and can capture adequacy better than surface overlap.
- Exact or normalised accuracy for specialised fields such as names, dates, numbers, and entity strings.
Do not present scores from different tokenisation schemes as if they were interchangeable. Urdu normalisation deserves explicit documentation: decide how to handle zero-width non-joiners, Arabic versus Persian characters, punctuation, whitespace, and Unicode normalisation. Keep raw text for human review and apply any normalisation only in a clearly labelled scoring variant.
Add human and targeted error analysis
Automatic metrics cannot reliably detect whether an Urdu sentence has the right politeness, named entity, tense, or negation. Sample errors from the full test set rather than reviewing only the lowest-scoring sentences. Have bilingual reviewers rate adequacy, fluency, and critical errors on a defined scale.
Build an error taxonomy relevant to Indian deployments:
- Wrong named entities, transliteration, dates, numbers, or currency.
- Negation, tense, honorifics, gender, and agreement errors.
- Omitted content or hallucinated details.
- Urdu-script spelling and diacritic inconsistencies.
- Code-mixed English, Hindi, Punjabi, or regional terms.
- Unsafe or misleading translations in public-service, health, legal, or financial text.
If the model will power a product, supplement FLORES with a small, licensed, domain-specific set. This is particularly important for teams comparing AI translation platforms for Indian regional languages or building production systems for linguistics professionals, where terminology and consistency matter more than a single general-domain score.
Make results reproducible and decision-ready
Publish a compact benchmark table containing model, direction, FLORES split, BLEU, chrF, COMET, latency, throughput, and hardware. Include confidence intervals or bootstrap significance tests when comparing close results. A difference of 0.5 BLEU is not automatically meaningful.
Track operational metrics alongside quality:
- Tokens per second and latency at the batch size used in production.
- Peak GPU memory and approximate cost per million characters.
- Failure rate, truncation rate, and unsupported-language behaviour.
- Performance on long sentences and domain-specific subsets.
Use a fixed seed where applicable, pin package versions, store the model revision, and commit the evaluation script. A small repository containing config.yaml, run_translate.py, score.py, predictions, and a README is more valuable than an unreproducible leaderboard claim.
Common mistakes to avoid
- Benchmarking a model in the wrong direction or with the wrong language token.
- Tuning on
devtestand reporting the result as an untouched test score. - Comparing case-sensitive and normalised outputs without saying so.
- Reporting BLEU without its SacreBLEU signature.
- Using sampling for one model and deterministic decoding for another.
- Treating FLORES as representative of conversational Urdu or every Indian deployment.
- Ignoring latency, memory, and entity preservation when selecting a production model.
A disciplined FLORES benchmark gives you a defensible baseline. Pair it with targeted Urdu evaluation, bilingual review, and deployment measurements before choosing a model for users. The same experiment structure can also support benchmarking NLP models for Telugu and Sanskrit and other Indian-language systems.