0tokens

Apply for AI Grants India

Financial support for innovators building the future of AI in India.

Apply now

Chat · how to benchmark marathi model before and after fine tuning on hugging face

How to Benchmark a Marathi Model Before and After Fine-Tuning

  1. aigi

    Why a before-and-after benchmark matters

    Fine-tuning can improve a Marathi model on a narrow task while reducing general language quality, increasing hallucinations, or damaging performance on code-switched Marathi-English text. A credible benchmark therefore compares the same model family, data conditions, prompts, decoding settings, and held-out examples before and after training.

    This guide focuses on Hugging Face workflows for classification, generation, instruction following, and translation. It also applies to models adapted for customer support, education, public services, and other Indian-language use cases. If you are still choosing a base model or training strategy, review best practices for fine-tuning LLMs on custom data before starting.

    1. Define the task and success criteria

    Do not begin with a generic claim that the model should “understand Marathi better”. Write down the production task and the measurable outcome.

    Examples include:

    • Classification: intent, sentiment, toxicity, topic, or eligibility labels.
    • Generation: helpful answers, summaries, rewriting, or structured extraction.
    • Translation: Marathi-English or English-Marathi translation.
    • Speech-adjacent NLP: normalising transcripts, spelling correction, or transliteration.
    • Retrieval and question answering: answering only from an approved knowledge base.

    Set a primary metric and guardrail metrics. For example, macro-F1 may be primary for an imbalanced classifier, while accuracy, calibration, and subgroup recall are guardrails. For a generative assistant, combine human ratings with automatic checks for factuality, instruction following, unsafe content, and Marathi fluency.

    2. Build a leakage-safe Marathi evaluation set

    Your test set is more important than the evaluation library. Keep it separate from training and validation data, and do not tune prompts or hyperparameters against it. Store the test set privately or restrict repository access so future experiments cannot accidentally train on it.

    A useful Marathi test set should include:

    • Devanagari text with varied sentence lengths and vocabulary.
    • Formal Marathi, conversational Marathi, dialectal variation, and code-switching.
    • Numbers, dates, names, abbreviations, punctuation, and spelling variation.
    • Hard examples identified by native speakers, not only randomly sampled text.
    • Relevant safety cases, including harassment, scams, self-harm, and sensitive personal data.
    • Duplicate and near-duplicate checks across train, validation, and test splits.

    Record the dataset version, licence, source, annotator guidance, label definitions, and disagreement rate. For generative tasks, use multiple reference answers where more than one Marathi response is valid. A single reference can unfairly penalise correct wording.

    3. Choose metrics that fit Marathi NLP

    Classification: report accuracy, macro-F1, per-class precision and recall, and a confusion matrix. Macro-F1 prevents large classes from hiding weak performance on minority intents.

    Language modelling: report perplexity only when the tokenizer, corpus, and tokenisation scheme are comparable. Perplexity is useful for measuring next-token prediction, but it does not prove that answers are helpful or factually correct.

    Generation: use exact match or token-level F1 for structured outputs; ROUGE or chrF for selected summarisation and translation settings; and human evaluation for fluency, adequacy, relevance, and factuality. Marathi morphology and word-order variation make character-level and semantic review particularly valuable.

    Translation: use chrF, COMET where supported, and bilingual human assessment. Ask reviewers to separately score meaning preservation, naturalness, omissions, and additions.

    Safety and reliability: measure refusal correctness, prompt-injection resistance, hallucination rate, and performance on adversarial or out-of-domain examples. Report confidence intervals or bootstrap intervals rather than presenting small score changes as meaningful improvements.

    4. Run a reproducible baseline on Hugging Face

    Pin the model revision, tokenizer, dataset revision, Python environment, random seed, device type, and evaluation batch size. Use identical preprocessing and generation parameters for both checkpoints. For generation, fix max_new_tokens, temperature, top-p, repetition penalty, and stopping rules; otherwise the comparison is not fair.

    For a supervised task, the datasets and evaluate libraries can provide a repeatable pipeline:

    from datasets import load_dataset
    from transformers import AutoTokenizer, AutoModelForSequenceClassification
    from transformers import Trainer
    
    model_id = "your-org/marathi-base"
    tokenizer = AutoTokenizer.from_pretrained(model_id, revision="main")
    model = AutoModelForSequenceClassification.from_pretrained(model_id, revision="main")
    
    def tokenize(batch):
        return tokenizer(batch["text"], truncation=True, max_length=512)
    
    test = load_dataset("your-org/marathi-eval", split="test")
    test = test.map(tokenize, batched=True)
    trainer = Trainer(model=model, tokenizer=tokenizer)
    raw = trainer.predict(test)

    Add a compute_metrics function for the task, save predictions and probabilities, and log the exact command used. A benchmark should be rerunnable by another engineer without reconstructing undocumented decisions.

    5. Fine-tune without contaminating the comparison

    Fine-tune from a fresh copy of the baseline checkpoint, not from a model that has already seen the test set. Keep an untouched baseline artefact so later experiments cannot overwrite it. For a small Marathi dataset, parameter-efficient methods such as LoRA or QLoRA can reduce memory use and make ablations easier; they do not remove the need for careful evaluation.

    Track:

    • Base model and immutable revision.
    • Training examples and their provenance.
    • Learning rate, epochs, effective batch size, and sequence length.
    • LoRA rank, target modules, quantisation settings, and adapter version where relevant.
    • Random seeds and checkpoint-selection rule.
    • Training, validation, and final test results.

    Use early stopping or checkpoint selection based on validation data only. Further guidance is available in this workflow on fine-tuning Llama for Indian regional languages, especially when adapting a multilingual model rather than a Marathi-specific checkpoint.

    6. Evaluate the fine-tuned checkpoint identically

    Load the final selected checkpoint and rerun the exact baseline command. Do not change the test prompts, examples, decoding settings, or post-processing. Produce a comparison table with absolute scores and deltas:

    | Measure | Baseline | Fine-tuned | Change |
    |---|---:|---:|---:|
    | Macro-F1 | 0.00 | 0.00 | +0.00 |
    | Marathi fluency rating | 0.00 | 0.00 | +0.00 |
    | Hallucination rate | 0.00 | 0.00 | -0.00 |

    Include confidence intervals where possible. Inspect bootstrap samples or repeated runs to determine whether the apparent gain exceeds evaluation noise. For generation, save the baseline and fine-tuned outputs side by side, then have bilingual reviewers blind-score them so knowledge of the checkpoint does not bias judgments.

    7. Analyse regressions, not just average gains

    A fine-tuned model is ready only when it improves the target task without unacceptable regressions. Slice results by text style, length, topic, dialect, code-switching, named entities, and safety category. Review examples where the baseline was correct but the fine-tuned model failed, and vice versa.

    Common Marathi-specific failure modes include:

    • Incorrect handling of inflections, postpositions, or honorifics.
    • Devanagari spelling drift and unnecessary transliteration into Latin script.
    • Marathi-English code-switching being copied or “corrected” incorrectly.
    • Loss of numbers, dates, names, and domain terminology.
    • Verbose or confident answers that invent details.
    • Overfitting to annotation phrases instead of learning the intended task.

    Use these failures to improve labels, add hard negatives, rebalance the data, or adjust the training recipe. Do not simply train for more epochs.

    8. Publish a benchmark card and deployment decision

    Document the model revision, dataset composition, metrics, limitations, known failure cases, compute, licence, and intended use. State whether the model is suitable for production, human-in-the-loop use, or research only. If latency and memory matter, measure them with the same quantisation and serving stack planned for deployment; optimisation can change output quality, so benchmark the deployed artefact too. See this guide to optimising AI models for mobile devices if the model will run on phones or edge hardware.

    For larger models, compare throughput, latency, and cost alongside quality. A modest Marathi quality gain may not justify a large increase in inference cost, while a small regression in a safety-critical workflow may be unacceptable. Teams deploying privately can also review how to deploy large language models locally before selecting an evaluation and serving setup.

    Practical checklist

    • Freeze a representative, leakage-safe Marathi test set.
    • Pin model, tokenizer, dataset, code, and environment revisions.
    • Select task-specific metrics plus safety and regression measures.
    • Run baseline and fine-tuned evaluations with identical settings.
    • Save predictions, not only aggregate scores.
    • Use native-speaker review for fluency, meaning, and factuality.
    • Report slices, confidence intervals, limitations, and deployment cost.
    • Re-run the benchmark after every data, prompt, tokenizer, or serving change.

    The strongest answer to how to benchmark a Marathi model before and after fine-tuning on Hugging Face is not a single score. It is a controlled, reproducible evidence trail showing where the model improved, where it regressed, and whether the change is valuable for the people who will use it.

    Last updated 23 September 2026

AIGI may be inaccurate. Replies seeded from the guide above.