0tokens

Apply for AI Grants India

Financial support for innovators building the future of AI in India.

Apply now

Chat · how to fine tune a model using indian public domain text on hugging face

How to Fine-Tune a Model with Indian Public-Domain Text

  1. aigi

    Fine-tuning a language model on Indian public-domain text can improve its handling of local vocabulary, literary styles, names, places, scripts, and cultural references. But the hard part is not calling trainer.train(). It is building a lawful, representative dataset and choosing a training method that matches your goal.

    This guide shows how to fine-tune a Hugging Face model using Indian public-domain text. It covers licence verification, corpus preparation, continued pretraining versus supervised fine-tuning, a practical Python workflow, evaluation, and publishing safeguards. If you need a broader training strategy, pair this guide with best practices for fine-tuning LLMs on custom data.

    Decide what “fine-tuning” means

    There are two common approaches:

    • Continued pretraining: Train a causal or masked language model on raw Indian-language text so it learns vocabulary, spelling, style, and domain patterns.
    • Supervised fine-tuning: Train on instruction-response pairs, classifications, summaries, or other labelled examples.

    A collection of books, newspapers, poems, or government documents is usually suitable for continued pretraining. It is not automatically an instruction-tuning dataset. If your target is a question-answering assistant, you must create high-quality examples or use retrieval-augmented generation alongside the model.

    For most small teams, parameter-efficient fine-tuning (PEFT), especially LoRA or QLoRA, is the practical starting point. It reduces GPU memory requirements, preserves the base model, and makes experiments easier to compare. Review the current ecosystem of Indian open-source AI developer projects before selecting a base model or tokenizer.

    Verify public-domain status before collecting text

    “Available online” does not mean “public domain”. Check the rights status of every source and retain evidence in a dataset card or internal registry.

    Before downloading material, record:

    • Author, title, publication date, language, and source URL.
    • The applicable copyright term and whether copyright has expired in India.
    • Any scan, transcription, translation, annotation, or website-specific licence.
    • Restrictions on redistribution, commercial use, attribution, or database reuse.
    • Whether personal data, sensitive content, or third-party excerpts appear in the text.

    A public-domain original can coexist with a copyrighted modern translation or transcription. Prefer primary scans and clearly licensed transcriptions. Keep provenance fields such as source_id, source_url, author, year, language, and licence with every document. Do not remove these fields during cleaning.

    Build a representative Indian-language corpus

    Indian text collections often contain mixed scripts, OCR errors, archaic spelling, and uneven coverage across languages. Start with a narrow, measurable scope rather than combining every available source.

    Useful corpus design decisions include:

    • Language and script: Distinguish Hindi in Devanagari from Romanised Hindi, Bengali from Assamese, and Urdu in Perso-Arabic script. Use language identification, but review samples manually.
    • Genre balance: Track prose, poetry, plays, essays, religious works, legal documents, and technical material separately.
    • Document boundaries: Preserve title, author, chapter, and paragraph boundaries where possible.
    • Deduplication: Remove repeated scans, boilerplate, mirrored repositories, and near-duplicate chapters.
    • OCR quality: Flag pages with excessive symbols, broken words, missing matras, or abnormal character distributions.
    • Contamination control: Keep a held-out set from distinct works, not random paragraphs from the same book.

    For multilingual training, log token counts by language and script. A large Hindi corpus can otherwise overwhelm smaller languages. Consider temperature-based sampling or explicit per-language caps, then report the final mixture.

    Prepare the dataset with Hugging Face Datasets

    Install a reproducible environment and pin versions in requirements.txt or a lockfile:

    pip install -U transformers datasets accelerate peft evaluate sentencepiece

    A simple JSONL record might look like this:

    {"text":"...paragraph or document text...", "language":"mar", "script":"Devanagari", "source_id":"book_001", "licence":"Public domain"}

    Load and split the data by document, not by individual lines:

    from datasets import load_dataset
    
    raw = load_dataset("json", data_files="data/indian_public_domain.jsonl", split="train")
    sources = sorted(set(raw["source_id"]))
    cut = int(len(sources) * 0.9)
    train_ids, eval_ids = set(sources[:cut]), set(sources[cut:])
    
    train = raw.filter(lambda row: row["source_id"] in train_ids)
    eval_ds = raw.filter(lambda row: row["source_id"] in eval_ids)

    In production, use a deterministic shuffle with a documented seed and stratify by language where the dataset is large enough. Remove empty records, but do not silently normalise away meaningful punctuation or script features. Store both the original and processed versions when feasible.

    Choose a compatible base model and tokenizer

    Select a model that already supports your target language and script. Check its model card, tokenizer vocabulary, licence, context length, and commercial-use terms. A multilingual model with poor tokenisation for an Indian script may be less efficient than a smaller language-specific model.

    Inspect tokenisation before training:

    from transformers import AutoTokenizer
    
    model_id = "your-compatible-causal-model"
    tokenizer = AutoTokenizer.from_pretrained(model_id)
    
    sample = "यह एक परीक्षण वाक्य है।"
    print(tokenizer.tokenize(sample))
    print(tokenizer(sample, return_tensors="pt")["input_ids"].shape)

    If text is split into an excessive number of tokens, training and inference costs rise sharply. Do not casually replace a pretrained tokenizer: changing its vocabulary requires carefully resizing embeddings and usually more training.

    Continued pretraining with Trainer

    For a causal language model, tokenise the text and group tokens into fixed-length blocks. This avoids wasting capacity on very short examples:

    from transformers import (AutoModelForCausalLM, DataCollatorForLanguageModeling,
                              Trainer, TrainingArguments)
    
    def tokenize(batch):
        return tokenizer(batch["text"], truncation=False)
    
    tok_train = train.map(tokenize, batched=True, remove_columns=train.column_names)
    tok_eval = eval_ds.map(tokenize, batched=True, remove_columns=eval_ds.column_names)
    
    block_size = 1024
    def group_texts(batch):
        joined = {k: sum(batch[k], []) for k in batch}
        length = (len(joined["input_ids"]) // block_size) * block_size
        return {k: [values[i:i+block_size] for i in range(0, length, block_size)]
                for k, values in joined.items()}
    
    tok_train = tok_train.map(group_texts, batched=True)
    tok_eval = tok_eval.map(group_texts, batched=True)
    
    model = AutoModelForCausalLM.from_pretrained(model_id)
    collator = DataCollatorForLanguageModeling(tokenizer=tokenizer, mlm=False)
    args = TrainingArguments(
        output_dir="outputs/indian-corpus",
        num_train_epochs=1,
        learning_rate=2e-5,
        per_device_train_batch_size=2,
        gradient_accumulation_steps=8,
        eval_strategy="steps",
        eval_steps=500,
        save_steps=500,
        logging_steps=25,
        bf16=True,
        report_to="none",
    )
    
    trainer = Trainer(
        model=model,
        args=args,
        train_dataset=tok_train,
        eval_dataset=tok_eval,
        data_collator=collator,
    )
    trainer.train()
    trainer.evaluate()

    The exact arguments depend on your installed Transformers version and GPU. On hardware without bfloat16 support, use fp16 where stable. Reduce block_size for limited memory, or use gradient checkpointing and PEFT. For larger experiments, launch with accelerate and track GPU memory, tokens processed, throughput, and checkpoint size.

    Evaluate language quality and usefulness

    Perplexity is useful for comparing runs on the same held-out corpus, but it is not a complete quality measure. Evaluate separately by language, script, genre, and document source.

    Include:

    • Held-out perplexity or loss by language.
    • Human review for spelling, fluency, cultural appropriateness, and historical accuracy.
    • Prompt tests for summarisation, extraction, translation, and continuation.
    • Repetition, memorisation, toxicity, and unwanted quotation checks.
    • Tokenisation efficiency and inference latency.

    Use distinct works in evaluation to detect memorisation. Never report only an aggregate score if a smaller language performs poorly. For task-specific models, publish accuracy, macro-F1, recall, or exact match with a clearly defined test protocol.

    Publish responsibly on Hugging Face

    A model repository should include the base model, training method, dataset sources, licences, languages, scripts, limitations, known failure modes, evaluation results, and intended use. Do not upload copyrighted material merely because it appeared in a scraped corpus. If provenance is uncertain, exclude it.

    Add a dataset card and model card, use a suitable licence, and document whether the adapter can be used commercially. Consider access controls for sensitive or low-confidence material. Builders developing multilingual products may also benefit from open-source vision-language models for Indian languages when their application combines text with scanned documents or images.

    A practical 2026 checklist

    • Define the target language, script, task, and success metric.
    • Verify rights and preserve source-level provenance.
    • Deduplicate and inspect OCR before tokenisation.
    • Split by document or author to prevent leakage.
    • Benchmark tokenisation with representative samples.
    • Start with a small PEFT or continued-pretraining run.
    • Compare against the untouched base model.
    • Evaluate every language separately and review samples manually.
    • Publish cards, licences, limitations, and reproducible training details.

    A well-documented, modest corpus is more valuable than a large, legally ambiguous scrape. Treat data quality, language coverage, and evaluation as first-class engineering work, and your Hugging Face model will be easier to trust, improve, and deploy in Indian products.

    Last updated 23 September 2026

AIGI may be inaccurate. Replies seeded from the guide above.