0tokens

Apply for AI Grants India

Financial support for innovators building the future of AI in India.

Apply now

Chat · how to fine tune a model using indian language news summaries on hugging face

How to Fine-Tune a Model on Indian-Language News Summaries

  1. aigi

    Fine-tuning a summarisation model for Indian news is not simply a matter of uploading articles and running Trainer. The quality of the result depends on language coverage, summary style, script handling, licensing, and evaluation that reflects how Indian readers actually consume news.

    This guide shows how to fine-tune a model using Indian language news summaries on Hugging Face. It focuses on an extractive or abstractive summarisation dataset containing an article and its human-written summary, and uses Hugging Face datasets, transformers, and evaluation tools. The same workflow can be adapted for Hindi, Bengali, Tamil, Telugu, Marathi, Gujarati, Kannada, Malayalam, Punjabi, Odia, Urdu, and multilingual datasets.

    If your dataset is small or a target language is under-represented, first review this builder’s guide to low-resource Indic natural language processing. It covers the data and modelling constraints that matter before training begins.

    1. Define the task and success criteria

    Decide what your model should produce before choosing an architecture:

    • Language: one language, several Indic languages, or code-mixed text.
    • Summary type: headline, short bullet, news brief, or long-form digest.
    • Length: define a target range in words or tokens.
    • Audience: local readers, journalists, internal researchers, or an end-user product.
    • Latency: a larger encoder-decoder model may score better but cost more to serve.

    For most article-to-summary projects, use a sequence-to-sequence model such as mT5, Indic-T5, or another multilingual encoder-decoder checkpoint. BERT and XLM-R are useful for classification and tagging, but they are not the natural starting point for generation. Compare available checkpoints on the Hugging Face Model Hub, checking their language coverage, licence, tokenizer, context length, and published benchmarks.

    2. Build a legally usable, representative dataset

    Collect only data you are permitted to use. News articles may be copyrighted, and a public URL does not automatically grant permission to download, redistribute, or train on the content. Keep source URLs, publication dates, licence notes, and consent or permission records alongside every example.

    A practical record contains:

    {"article": "पूर्ण समाचार लेख...", "summary": "मानवीय सारांश...", "language": "hi", "source": "publisher", "date": "2026-02-10"}

    Prioritise human-written summaries over summaries generated by another model. Remove duplicate wire stories, boilerplate navigation, advertisements, and malformed text. Preserve meaningful named entities, numbers, dates, quotes, and locations; aggressive cleaning can damage factual accuracy.

    Create train, validation, and test splits by story or source, not by randomly splitting near-identical syndicated articles. A typical starting point is 80/10/10, but a time-based test set is often more realistic for news. Include every target language in each split where data permits, and report results per language rather than hiding weaker languages behind a single average.

    3. Normalise Indic text carefully

    Indic text needs language-aware preprocessing. Do not remove punctuation or diacritics indiscriminately: punctuation can signal sentence boundaries, and Unicode differences can inflate vocabulary size. Useful checks include:

    • Unicode normalisation and consistent script encoding.
    • Removal of HTML, tracking parameters, and repeated whitespace.
    • Detection of empty, extremely short, or unusually long records.
    • Validation that the declared language matches the text.
    • Identification of transliterated or code-mixed content.
    • Checks for copied article text appearing in the summary.

    Keep a raw immutable copy and store preprocessing versions. This makes it possible to reproduce a run and investigate data leakage. For broader dataset engineering patterns, see best practices for fine-tuning LLMs on custom data.

    4. Load and tokenise the dataset

    Install the core packages in a clean environment:

    pip install -U transformers datasets evaluate accelerate sentencepiece sacrebleu rouge_score

    Load a local JSONL file and tokenize the article-summary pairs:

    from datasets import load_dataset
    from transformers import AutoTokenizer
    
    model_id = "google/mt5-small"
    tokenizer = AutoTokenizer.from_pretrained(model_id)
    dataset = load_dataset("json", data_files={
        "train": "train.jsonl",
        "validation": "validation.jsonl",
        "test": "test.jsonl",
    })
    
    max_source_length = 768
    max_target_length = 128
    
    def preprocess(batch):
        inputs = ["summarize: " + text for text in batch["article"]]
        model_inputs = tokenizer(
            inputs, max_length=max_source_length, truncation=True
        )
        labels = tokenizer(
            text_target=batch["summary"],
            max_length=max_target_length,
            truncation=True,
        )
        model_inputs["labels"] = labels["input_ids"]
        return model_inputs
    
    tokenized = dataset.map(preprocess, batched=True, remove_columns=dataset["train"].column_names)

    Measure how often articles exceed the source limit. Truncating the beginning may remove the story’s key facts, while truncating the end can remove the conclusion. For long articles, consider section-aware extraction, chunking followed by second-stage summarisation, or a long-context checkpoint.

    5. Fine-tune with Seq2SeqTrainer

    Use a model compatible with your tokenizer and language goals. Begin with a small checkpoint to validate the pipeline, then scale after the data and evaluation are reliable.

    import numpy as np
    import evaluate
    from transformers import (
        AutoModelForSeq2SeqLM, DataCollatorForSeq2Seq,
        Seq2SeqTrainer, Seq2SeqTrainingArguments
    )
    
    model = AutoModelForSeq2SeqLM.from_pretrained(model_id)
    collator = DataCollatorForSeq2Seq(tokenizer=tokenizer, model=model)
    rouge = evaluate.load("rouge")
    
    def compute_metrics(eval_pred):
        predictions, labels = eval_pred
        decoded_preds = tokenizer.batch_decode(predictions, skip_special_tokens=True)
        labels = np.where(labels != -100, labels, tokenizer.pad_token_id)
        decoded_labels = tokenizer.batch_decode(labels, skip_special_tokens=True)
        scores = rouge.compute(
            predictions=decoded_preds, references=decoded_labels, use_stemmer=False
        )
        return {key: round(value, 4) for key, value in scores.items()}
    
    args = Seq2SeqTrainingArguments(
        output_dir="./indic-news-summariser",
        eval_strategy="steps",
        eval_steps=500,
        save_steps=500,
        learning_rate=3e-5,
        per_device_train_batch_size=4,
        per_device_eval_batch_size=4,
        gradient_accumulation_steps=4,
        num_train_epochs=3,
        predict_with_generate=True,
        generation_max_length=128,
        logging_steps=50,
        load_best_model_at_end=True,
        metric_for_best_model="rougeLsum",
        greater_is_better=True,
        report_to="none",
    )
    
    trainer = Seq2SeqTrainer(
        model=model,
        args=args,
        train_dataset=tokenized["train"],
        eval_dataset=tokenized["validation"],
        tokenizer=tokenizer,
        data_collator=collator,
        compute_metrics=compute_metrics,
    )
    trainer.train()

    The exact argument names can vary across Transformers releases, so pin and record package versions for reproducibility. Use mixed precision when your GPU supports it, and reduce batch size if you encounter out-of-memory errors. Parameter-efficient methods such as LoRA can reduce memory use for larger models; they are especially useful when experimenting on rented or shared GPUs.

    6. Evaluate factuality, not only ROUGE

    ROUGE and BLEU are useful for regression testing, but lexical overlap does not prove that a summary is accurate. Evaluate:

    • Factual consistency: names, figures, dates, places, and claims match the article.
    • Coverage: the summary captures the main event and avoids irrelevant detail.
    • Fluency and script quality: grammar, punctuation, spelling, and natural phrasing.
    • Compression: the output is materially shorter without losing essential context.
    • Safety: allegations, communal references, medical claims, and breaking-news uncertainty are handled responsibly.

    Use a human review set sampled by language, publisher, topic, and article length. Ask bilingual reviewers to score faithfulness and readability separately. Also test adversarial cases: conflicting numbers, repeated names, code-mixed sentences, and articles whose key fact appears near the end.

    7. Publish and monitor the model

    Save the model and tokenizer, then publish a model card that states training data provenance, supported languages, limitations, evaluation results, known failure modes, and intended use:

    trainer.save_model("./indic-news-summariser")
    tokenizer.save_pretrained("./indic-news-summariser")

    If you upload to the Hub, do not include restricted news text in the repository unless redistribution is allowed. Consider publishing dataset metadata, hashes, preprocessing code, and evaluation scripts instead. For an India-focused open-source workflow, related Indian open-source AI developer projects can provide useful examples of documentation and collaboration.

    In production, monitor language-wise quality, unsupported claims, latency, token usage, and user reports. Keep a rollback version and re-evaluate after adding new publishers or changing prompts. A summariser should support editorial review—not silently replace it—especially for live news.

    Practical checklist

    Before calling the model ready, confirm that you have:

    • Permission to use and process the articles and summaries.
    • Deduplicated, language-labelled, versioned data.
    • Story-level or time-based validation and test splits.
    • A tokenizer and model that support the target scripts.
    • Per-language ROUGE or BLEU plus human factuality review.
    • Tests for long articles, numbers, names, code-mixing, and sensitive claims.
    • A model card, reproducible training configuration, and monitoring plan.

    Fine-tuning is the easy part once these foundations are sound. For Indian-language news, data provenance, script-aware processing, and factual evaluation usually matter more than adding another training epoch.

    Last updated 23 September 2026

AIGI may be inaccurate. Replies seeded from the guide above.