0tokens

Apply for AI Grants India

Financial support for innovators building the future of AI in India.

Apply now

Chat · how to benchmark telugu model before and after fine tuning on hugging face

How to Benchmark a Telugu Model Before and After Fine-Tuning

  1. aigi

    Fine-tuning a Telugu model can improve classification, translation, summarisation, question answering, or instruction following—but only if the improvement is measured against a reliable baseline. A single accuracy score is not enough. Telugu data varies by script quality, dialect, domain, spelling conventions, transliteration, and code-mixing with English. A model can appear stronger while becoming less reliable on clean, unseen Telugu text.

    This guide presents a reproducible Hugging Face workflow for benchmarking the same model before and after fine-tuning. It covers dataset design, baseline evaluation, task-specific metrics, statistical confidence, Telugu-focused error analysis, and practical reporting.

    Define the benchmark before training

    Start by writing down the task, model checkpoint, data version, evaluation split, metrics, and acceptance threshold. Do this before fine-tuning so that the benchmark cannot be adjusted to favour the final model.

    Your benchmark should specify:

    • Task: classification, token labelling, generation, translation, retrieval, or conversational response quality.
    • Input format: Telugu script, transliterated Telugu, Telugu-English code-mixed text, or a combination.
    • Target domain: news, education, government services, healthcare, customer support, or general web text.
    • Primary metric: the metric used for the go/no-go decision.
    • Secondary metrics: safety, latency, calibration, robustness, and subgroup performance.
    • Model and tokenizer versions: record the exact Hugging Face repository revision or commit.

    If your project includes a custom instruction dataset, follow established best practices for fine-tuning LLMs on custom data, especially around deduplication, prompt formatting, and held-out evaluation data.

    Build a trustworthy Telugu evaluation set

    Use separate train, validation, and test splits. The test set must remain untouched until the final comparison. Never select checkpoints or tune prompts repeatedly against the test set.

    A useful Telugu test set should represent the conditions in which the model will operate. Include variation in:

    • Formal and conversational Telugu
    • Regional vocabulary and dialectal phrasing
    • Spelling variation and punctuation differences
    • Telugu script and common Roman transliteration
    • Telugu-English code-mixing
    • Short queries, long documents, and noisy user input
    • Names, places, dates, currency, and government terminology

    Check for near-duplicates across splits. Duplicate news articles, paraphrased prompts, or records from the same conversation can make a model look better than it is. For sensitive applications, remove personally identifiable information and document consent and licensing.

    Keep a small challenge set separate from the main test set. It can contain difficult grammar, rare words, ambiguous queries, and examples from underrepresented domains. This set is not a replacement for the main benchmark; it exposes failure modes that average scores conceal.

    Install the evaluation stack

    A typical environment for sequence classification or token classification is:

    pip install -U transformers datasets evaluate accelerate scikit-learn pandas sentencepiece

    For generative tasks, add a task-appropriate evaluation library and define text normalisation rules in advance. Do not silently strip punctuation, change Unicode characters, or remove Telugu stopwords unless that policy applies equally to both model versions.

    Load data with an explicit revision where possible:

    from datasets import load_dataset
    
    raw = load_dataset(
        "your-org/your-telugu-dataset",
        revision="main"
    )
    
    print(raw)

    For production experiments, save a frozen export or dataset commit hash. Benchmark results are not reproducible if the underlying records change.

    Evaluate the pre-fine-tuning baseline

    Use the exact base checkpoint that will be fine-tuned. For a classification task, initialise the model with the correct label count and evaluate it on the same test set later used for the fine-tuned model.

    from transformers import AutoTokenizer, AutoModelForSequenceClassification
    
    checkpoint = "your-org/telugu-compatible-checkpoint"
    tokenizer = AutoTokenizer.from_pretrained(checkpoint)
    model = AutoModelForSequenceClassification.from_pretrained(
        checkpoint,
        num_labels=2
    )

    If the checkpoint was not originally trained for your task, its classification head may be randomly initialised. That is still a valid baseline, but label it clearly as a zero-shot or task-head baseline. For generative models, record the decoding settings—temperature, top-p, maximum tokens, and stop conditions—because changing them can alter results as much as fine-tuning.

    Use the same preprocessing and evaluation code for both checkpoints. A simple classification metric function might include macro-F1, weighted-F1, accuracy, and per-class recall:

    import evaluate
    
    accuracy = evaluate.load("accuracy")
    f1 = evaluate.load("f1")
    
    
    def compute_metrics(eval_pred):
        logits, labels = eval_pred
        predictions = logits.argmax(axis=-1)
        return {
            "accuracy": accuracy.compute(
                predictions=predictions, references=labels
            )["accuracy"],
            "macro_f1": f1.compute(
                predictions=predictions,
                references=labels,
                average="macro"
            )["f1"],
        }

    Macro-F1 is particularly useful when Telugu labels are imbalanced. Accuracy can hide poor performance on minority classes.

    Fine-tune without changing the experiment

    Keep the test set fixed and save the training configuration:

    from transformers import TrainingArguments, Trainer
    
    args = TrainingArguments(
        output_dir="telugu-run",
        eval_strategy="epoch",
        save_strategy="epoch",
        learning_rate=2e-5,
        per_device_train_batch_size=16,
        per_device_eval_batch_size=32,
        num_train_epochs=3,
        weight_decay=0.01,
        load_best_model_at_end=True,
        metric_for_best_model="macro_f1",
        report_to="none",
    )
    
    trainer = Trainer(
        model=model,
        args=args,
        train_dataset=tokenized["train"],
        eval_dataset=tokenized["validation"],
        tokenizer=tokenizer,
        compute_metrics=compute_metrics,
    )
    
    trainer.train()

    The current Transformers API may use eval_strategy; check the installed version if older examples use evaluation_strategy. Save the tokenizer, data collator, random seed, hardware details, and package versions. If you are adapting a large model with LoRA or another parameter-efficient method, report adapter rank, target modules, and merge settings. This matters when comparing results across experiments, including work involving fine-tuning Llama for Indian regional languages.

    Evaluate the fine-tuned checkpoint fairly

    Run the identical evaluation command on the selected fine-tuned checkpoint. Compare absolute scores and deltas:

    Metric          Before       After        Change
    Macro-F1        0.61         0.74         +0.13
    Accuracy        0.68         0.77         +0.09
    Minority recall 0.42         0.58         +0.16

    A meaningful report should include:

    • Overall and per-class metrics
    • Results on the challenge set
    • Performance by input type and domain
    • Telugu script versus transliterated text
    • Error counts and representative examples
    • Inference latency, memory use, and model size
    • Any regressions in general-purpose capability

    For generative tasks, combine automatic metrics with human evaluation. BLEU or chrF can help with translation; ROUGE can help with summarisation; exact match may suit structured question answering. For open-ended Telugu generation, use a rubric covering factuality, fluency, relevance, instruction following, and harmful or culturally inappropriate output. Human raters should see anonymised, randomly ordered outputs without knowing which checkpoint produced them.

    Test whether the gain is real

    Do not treat a one-point increase as meaningful without uncertainty. Use bootstrap confidence intervals for aggregate metrics, paired bootstrap tests for generation, or McNemar’s test for paired classification predictions. Report the number of test examples and the confidence interval where practical.

    Inspect the confusion matrix and compare examples where the models disagree. A fine-tuned model may improve the target domain while losing performance on general Telugu, English code-mixed input, or rare labels. If deployment is on constrained hardware, include efficiency measurements and review AI model optimisation for mobile devices before selecting the larger checkpoint.

    Common benchmarking mistakes

    • Evaluating the baseline and final model on different records
    • Letting test examples leak into prompts, training files, or model selection
    • Reporting only accuracy on an imbalanced dataset
    • Changing tokenisation or Unicode normalisation between runs
    • Comparing different decoding parameters
    • Removing difficult or noisy Telugu examples after inspecting results
    • Treating automatic generation metrics as a complete quality assessment
    • Ignoring latency, memory, licensing, and inference cost

    A practical release checklist

    Before publishing a model or deploying it, confirm that you have:

    • Frozen dataset and model revisions
    • Documented preprocessing and Telugu normalisation
    • Baseline and fine-tuned results from the same test set
    • Per-class, domain, and script-level analysis
    • Confidence intervals or paired significance testing
    • Qualitative review of model disagreements
    • Safety, privacy, and licensing checks
    • A model card describing limitations and intended use

    The benchmark should answer more than “did the score increase?” It should show where the Telugu model improved, where it regressed, and whether the gain matters for the intended users. For teams building an India-focused language stack, this evidence is more valuable than a headline metric and makes future experiments easier to reproduce.

    Last updated 23 September 2026

AIGI may be inaccurate. Replies seeded from the guide above.