0tokens

Apply for AI Grants India

Financial support for innovators building the future of AI in India.

Apply now

Chat · how to fine tune with lora on hugging face using indian datasets

How to Fine-Tune with LoRA on Hugging Face Using Indian Datasets

  1. aigi

    LoRA is one of the most practical ways for Indian AI teams to adapt an existing model without updating every parameter. It is especially useful when your dataset is modest, GPU access is limited, or you need separate adapters for Hindi, Tamil, Bengali, Hinglish, or a domain such as fintech, education, healthcare, or customer support.

    This guide explains how to fine-tune with LoRA on Hugging Face using Indian datasets. The examples use a text-classification task, but the same workflow applies to instruction tuning with a causal language model. For broader data and training decisions, pair this guide with best practices for fine-tuning LLMs on custom data.

    What LoRA changes

    Full fine-tuning updates all trainable weights in a model. LoRA freezes the base model and inserts small trainable matrices into selected layers, usually the attention projections. The adapter learns task-specific behaviour while the original model remains unchanged.

    The main benefits are:

    • Lower VRAM use: only adapter parameters and selected training states need updating.
    • Faster experiments: teams can test several datasets, languages, or domains without copying full model checkpoints.
    • Reusable adapters: one base model can support separate Hindi, Marathi, or customer-support adapters.
    • Simpler deployment: adapters can remain separate or be merged into the base model when required.

    LoRA is not a substitute for good data. It cannot reliably fix mislabeled examples, script inconsistencies, leakage, or a base model that does not support your target language.

    Choose the base model and task first

    For classification, use a multilingual encoder such as a suitable multilingual BERT-family checkpoint. For chat or generation, select a causal language model that covers the target Indian languages and fits your inference budget. Check the model card for licence terms, language coverage, tokenizer behaviour, and known limitations before training.

    Define the task precisely:

    • Classification: sentiment, intent, toxicity, eligibility, or complaint routing.
    • Generation: responses, summaries, translations, or structured extraction.
    • Instruction tuning: prompt-response examples with a clearly defined output format.

    A narrow task with consistent labels usually benefits more from LoRA than a vague attempt to make a general model “understand India”. If your product serves multilingual users, evaluate each major language separately rather than reporting one aggregate score.

    Prepare an Indian-language dataset

    Load data from a trusted Hugging Face dataset repository or create a versioned dataset from your own consented records. Public availability does not automatically make a dataset suitable for commercial training. Confirm its licence, collection method, personally identifiable information, and permitted uses.

    Before tokenisation, check:

    • Language and script labels, including Devanagari, Bengali, Tamil, Telugu, Kannada, Malayalam, Gujarati, Gurmukhi, and Romanised text.
    • Duplicate rows and near-duplicates across train and test splits.
    • Code-switching such as Hinglish and mixed-script messages.
    • Label balance and ambiguous examples.
    • PII such as phone numbers, Aadhaar-like identifiers, email addresses, account numbers, and exact addresses.
    • Transliteration variants, spelling variation, emojis, punctuation, and Unicode normalisation.

    Keep a genuine held-out test set. If examples from the same user, ticket, document, or conversation appear in multiple splits, your score may be inflated. For production systems, maintain a second evaluation set reflecting real traffic, including regional spelling, noisy mobile input, and code-switching.

    Install the Hugging Face stack

    Use a clean environment and pin versions for reproducibility. The following packages cover a standard classification workflow:

    pip install -U transformers datasets peft accelerate evaluate bitsandbytes

    bitsandbytes is optional and is mainly useful for quantised training on compatible NVIDIA hardware. Use a GPU where possible, but start with a small sample to validate the pipeline before paying for a long run.

    Load, tokenise, and attach LoRA

    Assume your dataset has text and label columns and train and test splits. Replace the dataset identifier with a verified Hugging Face repository or local dataset.

    from datasets import load_dataset
    from transformers import (
        AutoTokenizer, AutoModelForSequenceClassification,
        TrainingArguments, Trainer, DataCollatorWithPadding
    )
    from peft import LoraConfig, TaskType, get_peft_model
    
    model_id = "bert-base-multilingual-cased"
    dataset = load_dataset("your-org/indian-language-intents")
    
    tokenizer = AutoTokenizer.from_pretrained(model_id)
    
    def tokenize(batch):
        return tokenizer(batch["text"], truncation=True, max_length=256)
    
    tokenized = dataset.map(tokenize, batched=True)
    tokenized = tokenized.remove_columns(["text"])
    tokenized = tokenized.rename_column("label", "labels")
    
    model = AutoModelForSequenceClassification.from_pretrained(
        model_id, num_labels=2
    )
    
    lora_config = LoraConfig(
        task_type=TaskType.SEQ_CLS,
        r=8,
        lora_alpha=16,
        lora_dropout=0.1,
        target_modules=["query", "value"],
        modules_to_save=["classifier"]
    )
    model = get_peft_model(model, lora_config)
    model.print_trainable_parameters()

    The exact target_modules depend on the architecture. Inspect the model before training if the names differ. For decoder models, common targets include attention projections such as q_proj and v_proj, but blindly copying settings between architectures can produce an adapter that trains poorly or attaches to nothing.

    Train with an evaluation metric

    Accuracy alone can hide failures on minority languages or rare intents. Add macro-F1, per-language scores, and a confusion matrix where class imbalance matters.

    import numpy as np
    import evaluate
    
    f1 = evaluate.load("f1")
    
    def compute_metrics(eval_pred):
        logits, labels = eval_pred
        predictions = np.argmax(logits, axis=-1)
        return f1.compute(
            predictions=predictions,
            references=labels,
            average="macro"
        )
    
    args = TrainingArguments(
        output_dir="./lora-indian-intents",
        learning_rate=2e-4,
        per_device_train_batch_size=16,
        per_device_eval_batch_size=32,
        num_train_epochs=3,
        evaluation_strategy="epoch",
        save_strategy="epoch",
        load_best_model_at_end=True,
        metric_for_best_model="f1",
        greater_is_better=True,
        logging_steps=50,
        report_to="none",
    )
    
    trainer = Trainer(
        model=model,
        args=args,
        train_dataset=tokenized["train"],
        eval_dataset=tokenized["test"],
        tokenizer=tokenizer,
        data_collator=DataCollatorWithPadding(tokenizer),
        compute_metrics=compute_metrics,
    )
    
    trainer.train()
    trainer.evaluate()
    model.save_pretrained("./lora-indian-intents")
    tokenizer.save_pretrained("./lora-indian-intents")

    For larger models, combine LoRA with 4-bit quantisation through QLoRA, gradient accumulation, gradient checkpointing, and a lower sequence length. Monitor loss, GPU memory, throughput, and validation performance—not only the final checkpoint.

    Evaluate for Indian use cases

    Report results by language, script, and input type. A model that performs well on clean Hindi may fail on Romanised Hindi, regional names, or code-switched customer messages. Include:

    • Macro-F1 and per-class recall.
    • Per-language and per-script performance.
    • Robustness to spelling variation and noisy text.
    • False-positive and false-negative examples.
    • Abstention or escalation behaviour for uncertain inputs.
    • Human review by speakers familiar with the target communities.

    For applications such as education, lending, healthcare, or government services, establish a human escalation path. If the model will power a voice workflow, test transcripts and accents separately; related product design considerations appear in guides to AI voice solutions for Indian real estate developers and top-rated voice agent services for Indian businesses.

    Publish and deploy the adapter responsibly

    Upload the adapter, tokenizer, training configuration, dataset reference, evaluation results, and licence information to a private or public Hugging Face repository as appropriate. Do not publish raw sensitive records or memorised personal data. Document the languages, exclusions, known failure modes, and intended use.

    Keep the adapter separate when you need multiple domain or language variants. Merge it into the base model only after checking inference compatibility and latency. In production, log model and adapter versions, sample inputs safely, monitor drift, and schedule re-evaluation as user language changes.

    Practical checklist

    • Validate licences, consent, PII handling, and annotation quality.
    • Create leakage-resistant train, validation, and test splits.
    • Confirm tokenizer coverage for every target script.
    • Start with a small run and verify trainable parameter counts.
    • Use macro-F1 and per-language evaluation, not accuracy alone.
    • Compare LoRA against the unfine-tuned baseline.
    • Record hyperparameters, library versions, hardware, and dataset revision.
    • Document limitations before exposing the model to users.

    LoRA makes experimentation affordable, but responsible dataset preparation and evaluation determine whether the resulting model is genuinely useful. For Indian builders, the strongest workflow is usually a small, well-governed dataset, language-specific testing, and an adapter that can be improved without repeatedly retraining the entire base model.

    Last updated 23 September 2026

AIGI may be inaccurate. Replies seeded from the guide above.