0tokens

Apply for AI Grants India

Financial support for innovators building the future of AI in India.

Apply now

Chat · how to benchmark marathi model on indicglue using hugging face

How to Benchmark a Marathi Model on IndicGLUE with Hugging Face

  1. aigi

    IndicGLUE can give Marathi NLP systems a consistent evaluation target, but only if you treat benchmarking as an experiment rather than a single evaluate() call. Dataset configurations, label mappings, tokenisation, split names, and task-specific metrics all affect the result. This guide presents a reproducible workflow for benchmarking a Marathi model with Hugging Face in 2026, while flagging the implementation details that commonly produce misleading scores.

    What IndicGLUE measures

    IndicGLUE is a benchmark collection for Indian-language natural language understanding. Depending on the release and task configuration, it may include classification, inference, sentiment, question answering, named-entity recognition, or related datasets. Marathi is generally represented using the mr language code, but the exact configuration names and available splits should be verified rather than assumed.

    Before writing evaluation code, inspect the dataset card or repository and record:

    • The exact task and Marathi configuration.
    • Available train, validation, and test splits.
    • Input columns, label names, and label encoding.
    • Whether the official score uses accuracy, macro-F1, exact match, span F1, or another metric.
    • Any licensing, attribution, or test-submission requirements.

    This matters because a generic accuracy score is not comparable across tasks. If your goal is broader Indian-language evaluation, compare the setup with benchmarking NLP models for Telugu and Sanskrit, but keep Marathi results separated from results for other languages.

    Set up a reproducible environment

    Create a clean virtual environment and pin the principal libraries. Versions change quickly, so save the output of pip freeze with every benchmark run.

    python -m venv .venv
    source .venv/bin/activate
    pip install -U torch transformers datasets evaluate accelerate sentencepiece
    pip freeze > benchmark-requirements.txt

    Use a fixed random seed, identify the GPU or CPU used, and record the model revision. For public reporting, include the dataset revision and preprocessing script—not just the final score. Hugging Face model cards and dataset cards are useful places to preserve this information.

    Load and inspect the Marathi data

    Do not assume that load_dataset("indic_glue", "mr") is valid for every IndicGLUE release. Configuration names are often task-specific. First inspect the available configurations and follow the current dataset card.

    from datasets import get_dataset_config_names, load_dataset
    
    repo_id = "<indicglue-dataset-id>"
    configs = get_dataset_config_names(repo_id)
    print(configs)
    
    # Replace with the verified Marathi task configuration.
    dataset = load_dataset(repo_id, "<marathi-task-config>")
    print(dataset)
    print(dataset["train"].features)
    print(dataset["train"][0])

    Check for null text, duplicated examples, unexpected Unicode, and class imbalance. Marathi uses Devanagari, so normalisation must be conservative: do not remove characters merely because they are non-ASCII. Preserve punctuation and combining marks unless the benchmark instructions explicitly require normalisation.

    A simple audit can expose problems early:

    for split, data in dataset.items():
        print(split, len(data), data.column_names)
        print(data.filter(lambda row: row.get("text") in (None, "" )).num_rows)

    For sensitive or licensed data, avoid uploading raw examples to external experiment trackers.

    Select a suitable Marathi-capable model

    Start with a model whose tokenizer and pre-training data support Devanagari and Marathi. Multilingual encoders such as Indic-focused BERT variants can be strong baselines, while larger multilingual encoders may improve results at higher compute cost. Verify the model card for supported languages, intended tasks, vocabulary coverage, and license.

    For sequence classification, load the model with the correct number of labels:

    from transformers import AutoTokenizer, AutoModelForSequenceClassification
    
    model_id = "ai4bharat/indic-bert"  # verify the current model card
    num_labels = dataset["train"].features["label"].num_classes
    
     tokenizer = AutoTokenizer.from_pretrained(model_id, revision="main")
    model = AutoModelForSequenceClassification.from_pretrained(
        model_id,
        num_labels=num_labels,
        id2label={i: str(i) for i in range(num_labels)},
        label2id={str(i): i for i in range(num_labels)},
    )

    The leading space before tokenizer in the example above should be removed in a real script; keep benchmark code linted and version-controlled. If the model is not already fine-tuned for the task, its randomly initialised classification head means that the first evaluation is not a meaningful benchmark. Fine-tune on the official training split, then evaluate once on validation and, where permitted, once on test.

    For a domain-specific Marathi system, you may also test dialect coverage separately. The workflow in fine-tuning AI models for Marathi dialect is relevant when benchmark examples do not reflect the language variety used by your application.

    Tokenise without leaking information

    Map the actual text columns, not a guessed text field. Many benchmark tasks use paired inputs such as a premise and hypothesis.

    def tokenize_batch(batch):
        return tokenizer(
            batch["text"],
            truncation=True,
            max_length=256,
        )
    
    encoded = dataset.map(tokenize_batch, batched=True)

    For sentence pairs:

    def tokenize_pairs(batch):
        return tokenizer(
            batch["sentence1"],
            batch["sentence2"],
            truncation=True,
            max_length=256,
        )

    Prefer dynamic padding with DataCollatorWithPadding to reduce wasted memory. Choose max_length after inspecting Marathi sequence lengths; truncation can disproportionately remove the second part of long examples and change the task.

    Fine-tune and evaluate with the right metric

    Use TrainingArguments and a task-specific compute_metrics function. In current Transformers releases, use eval_strategy where supported; older environments may require evaluation_strategy.

    import evaluate
    import numpy as np
    from transformers import Trainer, TrainingArguments, DataCollatorWithPadding
    
    accuracy = evaluate.load("accuracy")
    f1 = evaluate.load("f1")
    collator = DataCollatorWithPadding(tokenizer=tokenizer)
    
    def compute_metrics(pred):
        predictions = np.argmax(pred.predictions, axis=-1)
        return {
            "accuracy": accuracy.compute(
                predictions=predictions, references=pred.label_ids
            )["accuracy"],
            "macro_f1": f1.compute(
                predictions=predictions,
                references=pred.label_ids,
                average="macro",
            )["f1"],
        }
    
    args = TrainingArguments(
        output_dir="runs/marathi-indicglue",
        eval_strategy="epoch",
        save_strategy="epoch",
        load_best_model_at_end=True,
        metric_for_best_model="macro_f1",
        greater_is_better=True,
        learning_rate=2e-5,
        per_device_train_batch_size=16,
        per_device_eval_batch_size=32,
        num_train_epochs=3,
        weight_decay=0.01,
        seed=42,
        report_to="none",
    )
    
    trainer = Trainer(
        model=model,
        args=args,
        train_dataset=encoded["train"],
        eval_dataset=encoded["validation"],
        tokenizer=tokenizer,
        data_collator=collator,
        compute_metrics=compute_metrics,
    )
    
    trainer.train()
    print(trainer.evaluate())

    For imbalanced classification, macro-F1 usually reveals weak minority-class performance better than accuracy. For NER and extractive question answering, use the benchmark’s official span-level or entity-level metric instead of reusing classification code.

    Make the result comparable

    Run at least three seeds when compute allows, and report mean and standard deviation. A single run can overstate progress, particularly on small Marathi datasets. Keep constant across model comparisons:

    • Dataset version and split.
    • Tokenisation and maximum length.
    • Training budget and early-stopping rule.
    • Random seeds and hardware details.
    • Metric implementation and label mapping.

    Save predictions, confidence scores, configuration files, and checkpoints. Never tune repeatedly on the test set. Use validation for model selection and reserve test results for the final report.

    Analyse Marathi-specific errors

    A score tells you how much the model fails, not why. Build a confusion matrix and manually inspect errors by category:

    • Devanagari spelling variation and tokenisation splits.
    • Code-mixed Marathi-English text.
    • Dialectal vocabulary and transliterated Marathi.
    • Named entities, abbreviations, and numerals.
    • Long inputs affected by truncation.
    • Ambiguous labels or annotation disagreement.

    Compare errors across length, script usage, and domain. If the production target is a mobile application, measure latency and memory as well as accuracy; AI model optimisation for mobile devices covers the deployment trade-offs that a leaderboard score misses. Teams building wider Indic-language systems can also review open-source small language models for Hindi when selecting a compact baseline, while remembering that Hindi performance does not guarantee Marathi performance.

    Publish a useful benchmark report

    A credible report should include the exact model revision, dataset configuration, preprocessing decisions, hyperparameters, seeds, hardware, evaluation script, and confidence intervals. State whether the model was fine-tuned, whether external Marathi data was used, and whether any examples were removed. Release code and predictions where licensing permits.

    IndicGLUE is a valuable reference point, not a complete Marathi quality assessment. Pair it with a small, carefully documented application set—such as customer support, education, agriculture, or public-service queries—and label that set independently. This gives Indian builders both a standardised score and evidence that the model works for the people and contexts it is meant to serve.

    Last updated 23 September 2026

AIGI may be inaccurate. Replies seeded from the guide above.