0tokens

Apply for AI Grants India

Financial support for innovators building the future of AI in India.

Apply now

Chat · how to benchmark marathi translation on flores using hugging face

How to Benchmark Marathi Translation on FLORES with Hugging Face

  1. aigi

    FLORES is useful for comparing machine translation systems, but a credible Marathi benchmark requires more than loading a dataset and printing a BLEU score. You need the correct FLORES release, direction, split, tokenizer, generation settings, and metric implementation. This guide presents a reproducible workflow for Marathi-to-English (mr-Deva → eng-Latn) evaluation with Hugging Face, while showing how to adapt it for English-to-Marathi and other Indian-language pairs.

    For broader model comparisons, pair this workflow with the Indian language LLM benchmark datasets guide and the practical framework for benchmarking multilingual LLMs in India.

    What FLORES measures

    FLORES-200 is a multilingual evaluation dataset built from professionally translated sentences. It is designed for controlled comparison, not for estimating performance across every real-world Marathi use case. Its sentences are relatively clean and general-domain, so results can overstate performance on government forms, conversational Marathi, code-mixed text, dialects, or specialised content.

    Use the benchmark to answer a narrow question: how well does a system translate the same fixed Marathi test sentences under clearly documented conditions? Report the following alongside every score:

    • FLORES version and configuration
    • Translation direction, such as mar_Deva-eng_Latn
    • Dataset split, normally devtest for final evaluation
    • Model revision and decoding parameters
    • Metric implementation and tokenisation settings
    • Number of evaluated examples and any filtering applied

    Do not tune prompts, checkpoints, or decoding settings on devtest. Use dev for development and reserve devtest for the final result.

    Set up a reproducible environment

    Install the libraries needed for dataset access, generation, and evaluation:

    pip install -U datasets transformers accelerate sentencepiece sacrebleu evaluate

    Pin versions in a requirements.txt or lock file. Reproducibility matters because dataset scripts, model defaults, and metric packages can change. Set a seed for sampling and record whether inference used CPU, CUDA, or another accelerator.

    FLORES configurations and language-code conventions can vary across releases. Inspect the available configurations rather than assuming that an older example will work:

    from datasets import get_dataset_config_names, load_dataset
    
    configs = get_dataset_config_names("facebook/flores")
    print([c for c in configs if "mar" in c.lower()][:20])

    If your installed version exposes a different dataset identifier, use the current Hugging Face dataset page and confirm the field names before writing the evaluation script. A common configuration is a language-pair name such as mar_Deva-eng_Latn.

    Load Marathi FLORES data correctly

    from datasets import load_dataset
    
    config = "mar_Deva-eng_Latn"
    dev = load_dataset("facebook/flores", config, split="dev")
    devtest = load_dataset("facebook/flores", config, split="devtest")
    
    print(devtest.column_names)
    print(devtest[0])

    The exact columns may include language-specific text fields and metadata. Inspect one record and identify the source and reference fields explicitly. Avoid blindly using dataset['train']: FLORES evaluation sets commonly use dev and devtest, and treating development data as a test set produces misleading reporting.

    For Marathi, preserve Devanagari exactly during input loading. Do not lowercase, transliterate, strip punctuation, or normalise Unicode unless the same policy is applied consistently and documented. Unicode normalisation can be appropriate, but it should be deliberate rather than an accidental preprocessing side effect.

    Select and configure a translation model

    Choose a model that supports the required direction and language codes. Examples include NLLB checkpoints and other multilingual encoder-decoder models available on Hugging Face. Verify that the model card explicitly lists Marathi; a model that accepts arbitrary language tags is not necessarily reliable for Marathi.

    The following example uses NLLB-style language tags. Adjust the checkpoint and fields to match your model:

    import torch
    from transformers import AutoTokenizer, AutoModelForSeq2SeqLM
    
    model_id = "facebook/nllb-200-distilled-600M"
    tokenizer = AutoTokenizer.from_pretrained(model_id, src_lang="mar_Deva")
    model = AutoModelForSeq2SeqLM.from_pretrained(model_id).eval()
    
    device = "cuda" if torch.cuda.is_available() else "cpu"
    model.to(device)

    Set the target language during generation. For NLLB, the target token is obtained from the tokenizer:

    target_id = tokenizer.convert_tokens_to_ids("eng_Latn")

    For Marian or another architecture, use the model-specific source and target conventions. Never compare a model with the wrong language tag and interpret the result as a Marathi capability score.

    Generate translations in batches

    Batching improves throughput, while a fixed maximum length and deterministic decoding make runs comparable. Start with beam search and record the settings; greedy decoding and sampling are different experiments.

    from tqdm.auto import tqdm
    
    source_col = "sentence_mar"
    reference_col = "sentence_eng"
    source_texts = devtest[source_col]
    references = devtest[reference_col]
    
    predictions = []
    for start in tqdm(range(0, len(source_texts), 16)):
        batch = source_texts[start:start + 16]
        inputs = tokenizer(
            batch,
            return_tensors="pt",
            padding=True,
            truncation=True,
            max_length=512,
        ).to(device)
    
        with torch.inference_mode():
            output = model.generate(
                **inputs,
                forced_bos_token_id=target_id,
                num_beams=5,
                max_new_tokens=256,
                do_sample=False,
            )
    
        predictions.extend(tokenizer.batch_decode(output, skip_special_tokens=True))

    Check for empty outputs, repeated text, unexpected scripts, and severe truncation before calculating metrics. Save predictions, references, model metadata, and generation parameters as JSON or JSONL so another researcher can audit the run.

    Calculate BLEU, chrF, and COMET

    Use sacreBLEU for standardised BLEU rather than a hand-written NLTK calculation. BLEU is useful for comparison but can undervalue valid Marathi or English paraphrases and is sensitive to tokenisation.

    import sacrebleu
    
    bleu = sacrebleu.corpus_bleu(predictions, [references])
    chrf = sacrebleu.corpus_chrf(predictions, [references])
    
    print("BLEU:", bleu.score)
    print("chrF:", chrf.score)
    print("signature:", bleu.signature)

    chrF is especially useful for morphologically rich and lower-resource settings because it evaluates character n-gram overlap. Add COMET when its supported model and compute budget are available; it often correlates better with human judgements, but it introduces another model, version, and possible language-coverage limitation. Report the COMET checkpoint, batch size, and whether the model supports the Marathi direction.

    Do not collapse these metrics into a single “accuracy” number. Present them separately, and include confidence intervals or bootstrap estimates when comparing small score differences. A one-point change may not be meaningful without significance testing.

    Analyse Marathi-specific errors

    Scores tell you that systems differ; inspection explains why. Sample errors by sentence length and manually review at least 50–100 examples. Track categories such as:

    • Named entities, place names, and transliterated terms
    • Honorifics, gender, tense, and agreement
    • Ambiguous Marathi words and compound expressions
    • Devanagari punctuation and number formatting
    • English insertions, code-mixing, and untranslated spans
    • Hallucinated content, omissions, and duplicated phrases

    Compare output against the source as well as the reference. A fluent English sentence can still omit a crucial Marathi qualifier. For deployment decisions, supplement FLORES with representative Marathi data from your target domain and, where possible, human evaluation by Marathi speakers.

    If your system must support dialect variation, a FLORES score is only a baseline. Consider the separate workflow for fine-tuning AI models for Marathi dialects. For products such as live interpretation, evaluate latency, streaming stability, terminology control, and failure recovery alongside translation quality; the AI translation platforms for Indian regional languages topic provides useful product-level context.

    What to publish with the benchmark

    A useful benchmark report should include:

    • Exact model ID, revision, tokenizer, and language tags
    • FLORES dataset ID, configuration, split, and row count
    • Hardware, software versions, batch size, and decoding parameters
    • BLEU, chrF, and any COMET configuration
    • Examples of errors and the review methodology
    • Predictions and evaluation code, subject to model and dataset terms

    For comparisons with Telugu, Sanskrit, or other Indian languages, keep the same pipeline but do not assume that scores are directly comparable across language directions. The NLP benchmarking guide for Telugu and Sanskrit is a useful companion for designing a consistent cross-language evaluation.

    Final checklist

    Before publishing a Marathi FLORES result, verify that you used the correct language pair, evaluated devtest only once, applied the target language token correctly, and recorded the metric signature. Re-run a small sample manually, inspect Unicode and empty outputs, and retain the generated translations. This turns a one-off score into a defensible, reproducible benchmark that can guide model selection and future fine-tuning.

    Last updated 23 September 2026

AIGI may be inaccurate. Replies seeded from the guide above.