0tokens

Apply for AI Grants India

Financial support for innovators building the future of AI in India.

Apply now

Chat · how to fine tune a model using indian retail product data on hugging face

How to Fine-Tune a Model on Indian Retail Data with Hugging Face

  1. aigi

    Indian retail data is rarely clean, uniform, or monolingual. A product catalogue may combine English, Hindi, transliterated text, regional brand names, pack sizes, GST fields, discounts, and marketplace-specific attributes. Fine-tuning can make a general model much more useful for tasks such as category prediction, attribute extraction, duplicate detection, search relevance, and product-title normalisation—but only if the dataset and evaluation design reflect those realities.

    This guide explains how to fine tune a model using Indian retail product data on Hugging Face, with a focus on text classification and product-catalogue workflows. The same process can be adapted to token classification, semantic matching, and instruction-tuning projects.

    Define the task before choosing a model

    Start with one measurable task. “Understand retail data” is too broad to train or evaluate reliably. Common starting points include:

    • Category classification: map titles or descriptions to a controlled taxonomy.
    • Attribute extraction: identify brand, colour, size, quantity, material, flavour, or model number.
    • Duplicate detection: decide whether two listings describe the same product.
    • Title normalisation: convert inconsistent seller titles into a standard format.
    • Search relevance: rank products against a customer query.

    Decide whether the task is single-label, multi-label, regression, ranking, or generation. For example, a product can belong to one primary category but have several attributes. For multi-label classification, your label format and model head must support multiple active labels.

    If the project will eventually power a customer-facing workflow, define the acceptable failure modes early. A wrong category may be inconvenient; extracting the wrong medicine strength or food allergen can be materially more serious. This risk-based approach should shape your review process and confidence thresholds.

    For broader training decisions, compare your plan with these best practices for fine-tuning LLMs on custom data.

    Build a representative Indian retail dataset

    The dataset should reflect the catalogue your model will see in production—not an idealised sample. Include variation across:

    • Grocery, fashion, electronics, beauty, home, pharmacy, and other target categories.
    • Marketplaces, distributors, brands, and seller quality levels.
    • English, Hindi, Hinglish, transliterated Indian-language text, abbreviations, and spelling variants.
    • Pack sizes such as 500 g, 0.5kg, 6 x 100 ml, and regional formatting conventions.
    • Seasonal products, promotional language, misspellings, and incomplete descriptions.

    Create a data dictionary before labelling. Record each field’s meaning, allowed values, source, and known limitations. Keep raw source data separate from the cleaned training table so that transformations remain auditable.

    Remove or mask phone numbers, email addresses, customer names, addresses, order IDs, and other personal data. Check licensing and marketplace terms before downloading or redistributing listings. A model repository should not expose private records or copyrighted data without permission.

    Clean, label, and split without leakage

    Basic cleaning includes Unicode normalisation, whitespace cleanup, duplicate removal, and consistent handling of missing values. Do not erase every punctuation mark or number: pack size, dosage, storage capacity, model number, and variant codes often carry the strongest signal.

    Create clear labelling rules with positive and negative examples. For category classification, specify how to handle bundles, accessories, multipacks, private labels, and ambiguous products. Measure agreement between annotators and send disputed records for adjudication. A small, consistently labelled dataset usually beats a larger dataset with contradictory labels.

    Avoid random row-level splitting when the same product appears across sellers or marketplaces. Instead, split by product family, brand, seller, or time period where appropriate. Otherwise, near-identical titles can appear in both training and test sets, producing an inflated score. Keep a locked test set that is not used for repeated model decisions.

    A practical split might be 80% training, 10% validation, and 10% testing, but the correct design depends on catalogue size and business risk. Preserve minority categories in each split and report performance by category, language, source, and text quality—not only an overall average.

    Choose a suitable Hugging Face model

    For short product titles and descriptions, encoder models are often more efficient than large generative models. Candidate families include multilingual BERT-style models, XLM-R, and Indian-language-focused encoders available on the Hugging Face Hub. Check the model card for training languages, licence, maximum sequence length, intended use, and known limitations.

    Choose a model based on the input and task:

    • Use sequence classification for one label per product.
    • Use token classification for extracting spans such as BRAND or SIZE.
    • Use a sentence-transformer-style model for similarity and deduplication.
    • Use a causal or encoder-decoder model only when the output is genuinely generative, such as structured title rewriting.

    Benchmark a simple baseline first. A keyword or TF-IDF model can reveal whether fine-tuning is adding value. Also compare full fine-tuning with parameter-efficient approaches such as LoRA when GPU memory or budget is limited. The Indian open-source AI developer projects guide is useful context for selecting community-supported tooling and models.

    Load and tokenise the dataset

    Install the core libraries:

    pip install -U transformers datasets evaluate accelerate scikit-learn

    Assume a CSV contains text and an integer label column. Load it, create a validation split if needed, and tokenise with truncation. Avoid padding every example to the maximum length during preprocessing; dynamic padding generally saves memory.

    from datasets import load_dataset
    from transformers import AutoTokenizer
    
    model_name = "xlm-roberta-base"
    dataset = load_dataset(
        "csv",
        data_files={"train": "train.csv", "test": "test.csv"}
    )
    dataset = dataset["train"].train_test_split(test_size=0.1, seed=42)
    
    tokenizer = AutoTokenizer.from_pretrained(model_name)
    
    def tokenize(batch):
        return tokenizer(
            batch["text"],
            truncation=True,
            max_length=192,
        )
    
    tokenised = dataset.map(tokenize, batched=True)

    Inspect the truncation rate. If important attributes routinely fall after the limit, increase max_length, restructure the input, or extract fields separately. Do not assume longer sequences automatically improve results.

    Fine-tune with the Trainer API

    For a single-label classification task, initialise the model and define metrics that reflect class imbalance. Macro-F1 is often more informative than accuracy when a few categories dominate the catalogue.

    import evaluate
    import numpy as np
    from transformers import (
        AutoModelForSequenceClassification,
        DataCollatorWithPadding,
        TrainingArguments,
        Trainer,
    )
    
    num_labels = 12
    model = AutoModelForSequenceClassification.from_pretrained(
        model_name,
        num_labels=num_labels
    )
    
    accuracy = evaluate.load("accuracy")
    f1 = evaluate.load("f1")
    
    def compute_metrics(eval_pred):
        logits, labels = eval_pred
        predictions = np.argmax(logits, axis=-1)
        return {
            "accuracy": accuracy.compute(
                predictions=predictions, references=labels
            )["accuracy"],
            "macro_f1": f1.compute(
                predictions=predictions,
                references=labels,
                average="macro"
            )["f1"],
        }
    
    args = TrainingArguments(
        output_dir="./retail-model",
        eval_strategy="epoch",
        save_strategy="epoch",
        learning_rate=2e-5,
        per_device_train_batch_size=16,
        per_device_eval_batch_size=32,
        num_train_epochs=3,
        weight_decay=0.01,
        load_best_model_at_end=True,
        metric_for_best_model="macro_f1",
        report_to="none",
    )
    
    trainer = Trainer(
        model=model,
        args=args,
        train_dataset=tokenised["train"],
        eval_dataset=tokenised["test"],
        tokenizer=tokenizer,
        data_collator=DataCollatorWithPadding(tokenizer),
        compute_metrics=compute_metrics,
    )
    trainer.train()

    API names can change between Transformers releases, so pin tested versions in requirements.txt and record the GPU, seed, model revision, and dataset version. Use early stopping, class weights, oversampling, or threshold tuning when minority categories are consistently missed. Do not optimise only for the headline metric: inspect confusion matrices and representative errors.

    Evaluate for production, not just for a leaderboard

    Run evaluation by language, category, seller, source marketplace, and text length. Test deliberately difficult examples: Hinglish, transliteration, missing brand names, misleading marketing claims, bundles, and similar pack sizes. Include an out-of-time test set if catalogue language changes frequently.

    For extraction, report entity-level precision, recall, and F1. For duplicate detection, examine false merges carefully. For search or ranking, use ranking metrics and human relevance judgements. Establish a human review queue for low-confidence predictions and monitor drift after deployment.

    Before publishing a model, add a model card covering data provenance, intended use, exclusions, metrics, languages, limitations, and privacy review. Upload only approved artefacts to the Hugging Face Hub, keep sensitive evaluation data private, and use access controls where necessary. If the model will support a production agent, review the deployment guidance in how to deploy open-source AI agents in production.

    Practical 2026 checklist

    • Define one task, label schema, and business failure threshold.
    • Audit language, category, seller, and regional coverage.
    • Remove personal data and verify dataset rights.
    • Split by product family or time to prevent leakage.
    • Establish a non-neural baseline.
    • Track macro-F1, per-class recall, and real-world error examples.
    • Version the dataset, code, model, and evaluation report.
    • Use confidence thresholds and human review for uncertain cases.
    • Monitor drift when brands, promotions, or catalogue structure change.

    Fine-tuning works best when treated as a data and evaluation project, not merely a training command. With representative Indian retail data, disciplined splits, multilingual testing, and transparent model documentation, Hugging Face provides a practical route from catalogue experiments to a maintainable production system.

    Last updated 23 September 2026

AIGI may be inaccurate. Replies seeded from the guide above.