0tokens

Apply for AI Grants India

Financial support for innovators building the future of AI in India.

Apply now

Chat · how to fine tune a model using indian msme data on hugging face

How to Fine-Tune a Model on Indian MSME Data

  1. aigi

    What you will build

    Fine-tuning is useful when a general-purpose model does not understand the language, workflows, or terminology of a specific business domain. For Indian MSMEs, a carefully prepared dataset can improve tasks such as classifying support tickets, extracting invoice fields, routing procurement queries, analysing customer feedback, or answering questions about internal policies.

    This guide uses a text-classification example because it is economical to train and straightforward to evaluate. The same workflow applies to other supervised tasks, with changes to the model head, labels, and evaluation metrics. Before choosing a base model, review best practices for fine-tuning LLMs on custom data and decide whether fine-tuning is actually required. Prompting, retrieval-augmented generation, or a rules-based system may be better for small datasets or frequently changing information.

    Start with a defensible MSME dataset

    The quality and provenance of the data matter more than the number of training rows. Possible sources include anonymised customer-support conversations, product descriptions, invoice narratives, loan-service queries, public government documents, and consented business records.

    Before uploading anything to Hugging Face:

    • Define the task and label set. For example: billing, delivery, technical, and other.
    • Remove personal and confidential information. Mask names, phone numbers, email addresses, GSTINs, bank details, addresses, order IDs, and proprietary financial figures unless they are essential and legally authorised for the task.
    • Record provenance. Keep a data sheet covering source, collection date, language, consent or licence, transformations, and known limitations.
    • Preserve Indian language variation. Decide whether the model must handle English, Hindi, Hinglish, or other Indian languages. Do not silently translate everything if language itself carries useful signal.
    • Check representation. Include different regions, sectors, business sizes, writing styles, and levels of digital maturity where the model will be used.

    For high-stakes systems, treat data quality as an engineering control. A useful reference point is the approach described in data veracity infrastructure for high-stakes AI. Keep a private source dataset and upload only the minimum transformed dataset needed for training.

    Choose a model and training approach

    For classification, start with a compact encoder such as multilingual BERT, IndicBERT, or another actively maintained multilingual checkpoint available on the Hugging Face Hub. Select a model based on language coverage, licence, tokenisation quality, inference cost, and benchmark performance—not popularity alone.

    For generation or instruction following, parameter-efficient fine-tuning methods such as LoRA or QLoRA can reduce GPU memory requirements. They are often more practical for Indian startups than updating every model parameter. However, adapters do not solve poor labels, duplicated examples, leakage, or a weak evaluation set.

    Use a held-out test set that is never used for model selection. A stratified split is generally preferable for classification, but split by customer, company, or time period when near-duplicate records could otherwise appear in both training and test data.

    Prepare the dataset in Hugging Face format

    A simple CSV can contain two columns:

    text,label
    "Payment failed for our monthly subscription",billing
    "Shipment has not arrived at the warehouse",delivery

    For a production workflow, use integer labels and preserve the original text only in a controlled environment. Install the core packages:

    pip install -U transformers datasets evaluate accelerate scikit-learn

    Load, inspect, and split the data:

    from datasets import load_dataset
    
    raw = load_dataset("csv", data_files={"all": "msme_support.csv"})["all"]
    raw = raw.shuffle(seed=42)
    
    split = raw.train_test_split(test_size=0.2, seed=42)
    train_test = split["train"].train_test_split(test_size=0.125, seed=42)
    dataset = {
        "train": train_test["train"],
        "validation": train_test["test"],
        "test": split["test"],
    }

    In a real project, validate required columns, remove empty rows, deduplicate near-identical messages, and inspect class counts before training. If one label dominates, accuracy alone will be misleading. Use macro F1 and per-class recall as primary measures.

    Tokenise and fine-tune

    Replace the checkpoint below with a model whose licence and language support fit your project. The label mapping must be saved with the model so that predictions remain interpretable.

    from transformers import (
        AutoTokenizer, AutoModelForSequenceClassification,
        TrainingArguments, Trainer
    )
    
    checkpoint = "ai4bharat/indic-bert"
    labels = sorted(set(dataset["train"]["label"]))
    label2id = {name: i for i, name in enumerate(labels)}
    id2label = {i: name for name, i in label2id.items()}
    
    tokenizer = AutoTokenizer.from_pretrained(checkpoint)
    
    def encode(batch):
        encoded = tokenizer(batch["text"], truncation=True, max_length=256)
        encoded["labels"] = [label2id[x] for x in batch["label"]]
        return encoded
    
    tokenised = {
        key: value.map(encode, batched=True, remove_columns=value.column_names)
        for key, value in dataset.items()
    }
    
    model = AutoModelForSequenceClassification.from_pretrained(
        checkpoint,
        num_labels=len(labels),
        label2id=label2id,
        id2label=id2label,
    )
    
    args = TrainingArguments(
        output_dir="msme-classifier",
        learning_rate=2e-5,
        num_train_epochs=3,
        per_device_train_batch_size=16,
        per_device_eval_batch_size=32,
        eval_strategy="epoch",
        save_strategy="epoch",
        load_best_model_at_end=True,
        metric_for_best_model="f1",
        report_to="none",
    )

    Add metrics and train:

    import evaluate
    
    f1 = evaluate.load("f1")
    accuracy = evaluate.load("accuracy")
    
    def metrics(prediction):
        predictions, references = prediction
        predicted = predictions.argmax(axis=-1)
        return {
            "f1": f1.compute(
                predictions=predicted, references=references, average="macro"
            )["f1"],
            "accuracy": accuracy.compute(
                predictions=predicted, references=references
            )["accuracy"],
        }
    
    trainer = Trainer(
        model=model,
        args=args,
        train_dataset=tokenised["train"],
        eval_dataset=tokenised["validation"],
        processing_class=tokenizer,
        compute_metrics=metrics,
    )
    trainer.train()

    Package versions change, so check the current Transformers API if an argument is deprecated. For limited compute, reduce batch size, use gradient accumulation, enable mixed precision on compatible GPUs, or use a LoRA-based training library.

    Evaluate for Indian MSME deployment

    Run evaluation only after finalising hyperparameters:

    results = trainer.evaluate(tokenised["test"])
    print(results)

    Do not stop at one score. Create an error report covering:

    • confusion between similar categories;
    • English, Hindi, Hinglish, and code-mixed inputs;
    • short, misspelled, and voice-transcribed messages;
    • sector-specific terms and regional references;
    • sensitive or adversarial prompts;
    • performance by company, language, geography, and time period.

    Have domain users review false positives and false negatives. If the model will influence credit, employment, collections, benefits, or compliance decisions, add human review and document an escalation path. A dashboard built with a no-code data analytics platform in India can help operations teams monitor drift, but it should not replace reproducible evaluation.

    Publish safely on Hugging Face

    Save the artefacts locally first:

    trainer.save_model("msme-classifier")
    tokenizer.save_pretrained("msme-classifier")

    When publishing to the Hub, create a private repository unless the dataset and model are clearly safe to share. Include a model card with intended use, prohibited use, training-data summary, licence, evaluation results, language coverage, known biases, and hardware used. Never publish raw customer records, memorised secrets, credentials, or an unreviewed dataset.

    For inference, pin dependency versions, add input validation, log confidence and model version, and provide a fallback for low-confidence predictions. Re-test after changes to product taxonomy, language mix, or upstream data. If your team is building a broader open-source stack, compare this workflow with Indian open-source AI developer projects for reusable patterns.

    A practical launch checklist

    • Task, labels, and success metrics are documented.
    • Data is consented or appropriately licensed and anonymised.
    • Train, validation, and test records are separated without leakage.
    • Evaluation includes macro F1, per-class recall, and slice analysis.
    • Human review exists for uncertain or consequential outputs.
    • Model, dataset, code, and dependency versions are reproducible.
    • Hugging Face visibility, access controls, and model-card disclosures are correct.
    • Monitoring covers drift, latency, cost, errors, and user feedback.

    Fine-tuning Indian MSME data can produce a focused, affordable model, but the durable advantage comes from disciplined data governance and continuous evaluation. Start with a narrow workflow, prove value against a strong baseline, and expand only when real operating data justifies it.

    FAQ

    How much data is needed?

    There is no universal threshold. A few hundred high-quality examples can establish a baseline for a narrow classification task, while broader multilingual or generative use cases typically need substantially more coverage. Measure performance by label and language rather than relying on row count.

    Should I upload MSME data publicly?

    Only data that is legally shareable, genuinely anonymised, and safe to disclose should be public. Use private Hugging Face repositories, access controls, and internal storage for confidential material.

    Is fine-tuning better than retrieval?

    Fine-tuning changes model behaviour and is useful for stable patterns, tone, or classification. Retrieval is usually better for private facts that change frequently, such as product catalogues, policies, prices, and scheme rules. Many production systems use both.

    Last updated 23 September 2026

AIGI may be inaccurate. Replies seeded from the guide above.