0tokens

Apply for AI Grants India

Financial support for innovators building the future of AI in India.

Apply now

Chat · how to benchmark malayalam model before and after fine tuning on hugging face

How to Benchmark a Malayalam Model Before and After Fine-Tuning

  1. aigi

    Fine-tuning can improve a Malayalam model—or simply make it memorise a narrow dataset. A credible comparison requires more than running trainer.evaluate() twice: you need a fixed test set, task-appropriate metrics, reproducible settings, and error analysis that reflects Malayalam usage in India.

    This guide shows how to benchmark a Malayalam text-classification model before and after fine-tuning with Hugging Face Transformers and Datasets. The same workflow can be adapted for named-entity recognition, sentiment analysis, topic classification, and other supervised tasks. If you are still choosing a training strategy, start with these best practices for fine-tuning LLMs on custom data.

    Define the comparison before training

    Write down the experiment before opening a GPU session. Your baseline and fine-tuned model should be evaluated on:

    • The same held-out test examples
    • The same label mapping and preprocessing
    • The same sequence length and truncation policy
    • The same evaluation metrics and averaging method
    • The same hardware or clearly documented inference settings

    The baseline is not necessarily a Malayalam-specific model. It may be a multilingual encoder such as IndicBERT, MuRIL, XLM-R, or another model available on the Hugging Face Hub. Check the model card for language coverage, tokenizer behaviour, licence, and intended use. Do not assume that a model with Malayalam data in pre-training will perform well on your domain.

    For a fair study, define an experiment identifier containing the model revision, dataset version, random seed, code commit, and fine-tuning configuration. This matters particularly for small Malayalam datasets, where results can shift substantially between random splits.

    Prepare a Malayalam evaluation dataset

    Create separate train, validation, and test splits. Keep the test split untouched until the final comparison. Avoid near-duplicates across splits—for example, the same news story copied by several Malayalam websites—because this can inflate the apparent score.

    Inspect the data for issues that commonly affect Indian-language benchmarks:

    • Unicode normalisation differences and invisible characters
    • Malayalam punctuation, numerals, emojis, and mixed Malayalam-English text
    • Duplicate or contradictory labels
    • Highly imbalanced classes
    • Transliteration written in Latin script
    • Domain gaps between formal news, social media, government text, and conversational Malayalam

    Use stratified splits for classification where possible. If your application receives time-ordered content, use a temporal test set as an additional measure rather than relying only on a random split. Preserve the original text in a separate column so that later error analysis is possible.

    from datasets import load_dataset
    
    raw = load_dataset("your-org/your-malayalam-dataset")
    print(raw)
    print(raw["train"].features)

    Map labels consistently and verify that every split contains the expected classes. A benchmark with accidental label remapping is worse than no benchmark because it produces confident but meaningless comparisons.

    Install and load the baseline model

    Use pinned versions when reporting results. A practical starting point is:

    pip install -U "transformers>=4.45" datasets evaluate accelerate

    Load the tokenizer and sequence-classification model using the same checkpoint that will later be fine-tuned:

    from transformers import AutoTokenizer, AutoModelForSequenceClassification
    
    checkpoint = "your-org/your-malayalam-checkpoint"
    label_names = ["negative", "neutral", "positive"]
    
     tokenizer = AutoTokenizer.from_pretrained(checkpoint, use_fast=True)
    model = AutoModelForSequenceClassification.from_pretrained(
        checkpoint,
        num_labels=len(label_names),
        id2label=dict(enumerate(label_names)),
        label2id={name: i for i, name in enumerate(label_names)},
    )

    The leading space before tokenizer in the example should be removed in a real script. More importantly, confirm whether the checkpoint already has a task-specific classification head. If the head is newly initialised, the “before fine-tuning” score may be close to random; that is still a valid baseline, but label it clearly. A zero-shot or prompted generative baseline can be reported separately, not mixed with supervised classification results.

    Tokenise once using a version-controlled function:

    def tokenize(batch):
        return tokenizer(
            batch["text"],
            truncation=True,
            max_length=256,
        )
    
    encoded = raw.map(tokenize, batched=True)
    encoded = encoded.rename_column("label", "labels")
    encoded = encoded.remove_columns(["text"])

    Choose max_length after inspecting Malayalam text lengths. Excessive truncation can hide the information needed for classification, while an unnecessarily large value increases cost and memory use.

    Benchmark the pre-fine-tuning model

    Accuracy alone is inadequate when one class dominates. Report macro-F1, weighted-F1, precision, recall, and a confusion matrix. Macro-F1 gives each class equal weight, which is useful when minority categories matter.

    import evaluate
    import numpy as np
    from transformers import TrainingArguments, Trainer
    
    accuracy = evaluate.load("accuracy")
    f1 = evaluate.load("f1")
    
    
    def compute_metrics(eval_pred):
        logits, labels = eval_pred
        predictions = np.argmax(logits, axis=-1)
        return {
            "accuracy": accuracy.compute(
                predictions=predictions, references=labels)["accuracy"],
            "macro_f1": f1.compute(
                predictions=predictions, references=labels,
                average="macro")["f1"],
            "weighted_f1": f1.compute(
                predictions=predictions, references=labels,
                average="weighted")["f1"],
        }
    
    args = TrainingArguments(
        output_dir="runs/malayalam-baseline",
        per_device_eval_batch_size=32,
        report_to="none",
        seed=42,
    )
    
    trainer = Trainer(
        model=model,
        args=args,
        eval_dataset=encoded["test"],
        compute_metrics=compute_metrics,
    )
    
    before = trainer.evaluate()
    print(before)

    For a randomly initialised classification head, this result is best interpreted as a pre-training task baseline rather than the quality of the underlying language encoder. If you want to measure transfer before supervised fine-tuning, compare against an already trained task checkpoint, a linear probe with the encoder frozen, or a carefully designed zero-shot baseline.

    Save the exact output as JSON, along with the test-set hash and checkpoint revision. Also record inference latency, peak memory, and model size if deployment is part of the decision. These measurements are useful when comparing accuracy against production constraints, including AI model optimisation for mobile devices.

    Fine-tune with controlled settings

    Fine-tune only on train; use validation for model selection and reserve test for the final report:

    from transformers import DataCollatorWithPadding
    
    collator = DataCollatorWithPadding(tokenizer=tokenizer)
    
    train_args = TrainingArguments(
        output_dir="runs/malayalam-finetuned",
        learning_rate=2e-5,
        num_train_epochs=3,
        per_device_train_batch_size=16,
        per_device_eval_batch_size=32,
        weight_decay=0.01,
        evaluation_strategy="epoch",
        save_strategy="epoch",
        load_best_model_at_end=True,
        metric_for_best_model="macro_f1",
        greater_is_better=True,
        seed=42,
        report_to="none",
    )
    
    trainer = Trainer(
        model=model,
        args=train_args,
        train_dataset=encoded["train"],
        eval_dataset=encoded["validation"],
        tokenizer=tokenizer,
        data_collator=collator,
        compute_metrics=compute_metrics,
    )
    
    trainer.train()

    Set metric_for_best_model to the metric that reflects product risk. For a safety-sensitive minority class, macro-F1 or class-specific recall may be more meaningful than accuracy. Use class weighting or resampling only after documenting the choice; these can improve minority recall while reducing calibration or overall precision.

    For larger language models, parameter-efficient approaches such as LoRA can reduce GPU requirements. Keep the adapter configuration, base-model revision, and quantisation settings in the experiment record. For regional-language model work, the workflow in fine-tuning Llama for Indian regional languages provides useful context on data and evaluation decisions.

    Benchmark after fine-tuning

    Evaluate the selected checkpoint on the untouched test set using the same trainer configuration:

    after = trainer.evaluate(eval_dataset=encoded["test"])
    print(after)

    Create a comparison table rather than reporting only the best number:

    | Metric | Before | After | Change |
    |---|---:|---:|---:|
    | Accuracy | | | |
    | Macro-F1 | | | |
    | Weighted-F1 | | | |
    | Minority-class recall | | | |
    | Inference latency | | | |

    Calculate absolute and relative changes, but do not overstate small gains. Bootstrap confidence intervals or run several seeds when the test set is small. If a 0.5-point improvement falls within expected variation, describe the result as inconclusive rather than as a meaningful upgrade.

    Go beyond aggregate scores

    Inspect predictions that changed after fine-tuning. Export the text, gold label, baseline prediction, fine-tuned prediction, and confidence. Group errors by:

    • Script and transliteration style
    • Topic or source domain
    • Text length
    • Named entities and code-switching
    • Ambiguous or sarcastic language
    • Rare labels and spelling variants

    A confusion matrix can show that the overall F1 gain comes entirely from one easy class. Calibration checks can reveal whether the fine-tuned model is overconfident. Manually review at least a stratified sample with Malayalam-fluent evaluators; automatic metrics cannot reliably detect culturally specific ambiguity, offensive content, or annotation problems.

    If your task is generation rather than classification, replace accuracy and F1 with task-specific measures such as chrF, BLEU, ROUGE, COMET where appropriate, and human evaluation. For translation work, pair automatic scores with adequacy, fluency, terminology, and named-entity checks. Do not compare generative and classification metrics as if they measured the same capability.

    Common mistakes to avoid

    • Changing the test set after fine-tuning
    • Tuning on test results, which turns the test set into validation data
    • Reporting accuracy only on imbalanced Malayalam classes
    • Ignoring Unicode and transliteration variation during preprocessing
    • Comparing different tokenizers or maximum lengths without disclosure
    • Using a newly initialised head as proof of poor base-model quality
    • Claiming general Malayalam performance from one narrow domain
    • Publishing data or model artefacts that contain private personal information

    What a credible report includes

    Publish the dataset source and licence, split methodology, label definitions, model and tokenizer revisions, preprocessing code, hyperparameters, random seeds, hardware, evaluation scripts, confidence intervals, and representative errors. State whether the data includes Kerala news, government documents, social media, education, healthcare, or other domains. This makes the result reproducible and helps other builders judge whether it transfers to their application.

    The strongest conclusion is not always “fine-tuning improved the model.” A useful benchmark may show that gains are limited to one domain, that minority recall worsened, or that a smaller model delivers comparable quality at lower cost. For teams deploying locally, compare quality with latency and memory using the same test suite; for teams extending multilingual systems, related open-source small language models for Hindi can provide a useful cross-language comparison without substituting for Malayalam evaluation.

    Last updated 23 September 2026

AIGI may be inaccurate. Replies seeded from the guide above.