What fine-tuning should solve
Fine-tuning is useful when a general model understands language but misses the details that matter in Indian commerce: brand and model names, regional spellings, pack sizes, ₹ pricing, product attributes, Hinglish queries, and category-specific vocabulary. The right objective is not to make a model memorise a catalogue. It is to improve a measurable task such as product classification, attribute extraction, semantic search, duplicate detection, or query-to-product relevance.
Start by defining the production decision the model will support. A classifier may predict a category or attribute; a bi-encoder may retrieve relevant products; a cross-encoder may rerank search results; and a generative model may normalise titles or create structured metadata. For many catalogue use cases, a smaller encoder fine-tuned for classification or retrieval is cheaper, faster, and easier to control than a large language model.
If you are still comparing training approaches, review these best practices for fine-tuning LLMs on custom data before choosing a model or GPU budget.
1. Secure and define the catalogue data
Use data you are authorised to process. Do not scrape or redistribute marketplace content in violation of terms, and remove customer names, addresses, phone numbers, order IDs, reviews containing personal information, and internal seller data that is not required for the task. Maintain a data sheet recording the source, licence, collection date, fields, language, and permitted use.
A useful product record might contain:
product_id: a stable internal identifier, not a customer or seller identifier.title: the original title, retained separately from a cleaned version.description: bullet points and long description, with HTML removed.brand,category,subcategory, and structured attributes.languageor language-mix labels, where available.query,click, purchase, or relevance labels for search and recommendation tasks.
Keep prices and stock status out of the training text unless the model genuinely needs them. These fields change frequently and can cause leakage. Version the raw and processed datasets so that every experiment can be reproduced.
2. Clean Indian product text without erasing useful signals
Catalogue data is usually noisy. Normalise whitespace, HTML, broken encodings, duplicated bullets, tracking parameters, and repeated promotional phrases. Preserve information that helps shoppers distinguish products: storage capacity, screen size, fabric, colour, gender, pack count, compatibility, and regional terms.
Do not blindly lowercase or remove punctuation. MI 11, Mi 11, and mi 11 may need to map to the same entity, while 10 kg and 1.0 kg require careful handling. Build validation rules for units, currency, dimensions, and pack sizes. Keep the original text for auditability and store transformations as explicit preprocessing code.
Indian catalogues often contain English, transliterated Hindi, Tamil, Bengali, Marathi, Telugu, or mixed-language text. Add representative examples rather than translating everything automatically. Test common variants such as “kurti for office”, “phone under 20k”, and regional spellings against the actual search traffic. A model that performs well only on polished English is not ready for a broad Indian audience.
3. Build task-specific training examples
The label format depends on the task:
- Classification: map a title or product text to a category, brand, or attribute label.
- Named-attribute extraction: provide spans or structured targets for colour, size, material, or compatibility.
- Semantic retrieval: pair a shopper query with a relevant product and add hard negatives from the same category.
- Duplicate detection: label whether two listings represent the same product or variant.
- Reranking: label several candidate products for one query with graded relevance.
Split by product family, seller, or time period—not only by random rows. Random splitting can place near-identical variants in both training and test sets and inflate results. Reserve a difficult test set containing new brands, spelling variations, regional language, incomplete queries, and recently added products. If you plan to fine-tune a generative model, follow the same discipline described in Indian open-source AI developer projects around reproducible datasets and transparent evaluation.
4. Choose a Hugging Face model and training method
For classification or extraction, begin with a compact encoder that supports the languages in your data. For multilingual retrieval, select a sentence-embedding model and fine-tune query-product pairs. For generation, use supervised fine-tuning only when you need controlled text output; otherwise, a structured classifier or extraction model will usually be more predictable.
Install a minimal environment:
pip install transformers datasets evaluate accelerate sentence-transformersUse the Hugging Face Hub to pin a model revision, document its licence, and keep private datasets private. A typical classification setup looks like this:
from datasets import load_dataset
from transformers import AutoTokenizer, AutoModelForSequenceClassification
model_id = "distilbert-base-multilingual-cased"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForSequenceClassification.from_pretrained(
model_id,
num_labels=number_of_labels,
)
def tokenize(batch):
return tokenizer(
batch["text"],
padding="max_length",
truncation=True,
max_length=192,
)Replace the model with one suited to your languages and task. For large models, parameter-efficient methods such as LoRA can reduce GPU memory and make experiments easier to compare. Do not assume a larger checkpoint will outperform a clean dataset and a well-designed baseline.
5. Train with reproducible, conservative settings
Tokenise only the fields needed for the task. Truncation can remove critical attributes from long descriptions, so measure how often it occurs and consider title-plus-attribute templates. Start with a low learning rate, a small number of epochs, early stopping, and a validation set. Save the tokenizer, label mapping, preprocessing version, model revision, seed, and training configuration with every run.
Monitor training and validation loss together. A widening gap usually indicates overfitting, noisy labels, duplicated examples, or an overly aggressive learning rate. Class imbalance is common: a few popular categories may dominate the catalogue while long-tail categories have very few examples. Use stratified sampling, class weights, or targeted collection of hard examples rather than duplicating rows indiscriminately.
For retrieval, evaluate negative sampling carefully. Random negatives are often too easy; add products from the same category, brand, or price range. For a production deployment, latency and memory matter as much as offline quality. Quantisation, batching, caching, and a smaller reranker may deliver more value than another training epoch.
6. Evaluate business outcomes, not just accuracy
Report metrics by category, language, device type, and catalogue freshness. Accuracy alone can hide poor long-tail performance. Use macro-F1 for imbalanced classification, precision and recall for extraction, and recall@k, MRR, or nDCG for retrieval and ranking. Always compare against a non-fine-tuned baseline and a simple lexical search baseline such as BM25.
Inspect false positives and false negatives manually. Check whether the model confuses variants, promotes out-of-stock items, drops safety or compatibility attributes, or treats an unrelated product as relevant because of shared marketing language. Run a small human review with shoppers, category experts, or support teams who understand Indian product terminology.
Before launch, test prompt injection and malicious catalogue text if a generative model is involved. Prevent the model from inventing prices, discounts, certifications, or product specifications. Retrieval systems should apply stock, policy, seller-quality, and safety filters outside the model rather than trusting generated output.
7. Publish and deploy responsibly
Push only the artefacts you are permitted to share. A private Hugging Face repository can store the model card, evaluation results, intended use, limitations, dataset provenance, licence, and known language gaps. Never upload raw customer data or unredacted internal catalogue exports.
For deployment, export a versioned model and expose a stable inference interface. Log model version, input schema, latency, and aggregate error signals without retaining unnecessary shopper text. Add a rollback path and retraining trigger for catalogue changes, new languages, seasonal products, and taxonomy updates. Teams considering an on-premise or self-hosted setup can also compare the operational trade-offs in how to deploy open-source AI agents in production, particularly around observability, access control, and rollback.
A practical launch checklist
- Define one task and one primary success metric.
- Confirm data rights, remove personal data, and document provenance.
- Deduplicate by product family and split by time or entity.
- Include multilingual, transliterated, long-tail, and hard-negative examples.
- Benchmark against a general model and a lexical baseline.
- Evaluate latency, cost, robustness, and subgroup performance.
- Publish a model card and keep private data out of public repositories.
- Monitor catalogue drift and retrain only when new evidence justifies it.
The best Hugging Face fine-tuning project is not the one with the largest model. It is the one that improves a defined commerce workflow, remains auditable, and continues to work as Indian catalogues, languages, sellers, and shopper behaviour change.