ONDC catalog data can support useful models for product classification, attribute extraction, catalogue quality checks, search, and seller tooling. But a catalog is not automatically a training dataset: fields vary by domain, languages and transliterations are common, and commercial data may carry privacy, licensing, or policy constraints.
This guide explains how to fine-tune a model using ONDC catalog data on Hugging Face in a way that is reproducible and useful for an Indian product. It focuses on supervised fine-tuning for a clearly defined task, rather than indiscriminately training a large language model on every available field.
Start with the task, not the model
Define the prediction you need before downloading data. Common ONDC-oriented tasks include:
- Product categorisation: map seller listings to a controlled taxonomy.
- Attribute extraction: identify brand, size, colour, quantity, material, or dietary attributes.
- Catalogue quality scoring: flag incomplete titles, inconsistent units, duplicate listings, or implausible prices.
- Search and matching: rank products against a buyer query or match equivalent listings.
- Translation and normalisation: convert Hinglish, regional-language text, and transliterated terms into a consistent representation.
Choose the model and label format around that task. A sequence-classification model is appropriate for categories; token classification suits span extraction; a bi-encoder or cross-encoder is better for semantic matching. For Hindi or multilingual commerce text, compare a multilingual checkpoint with a model suited to Indian languages rather than defaulting to bert-base-uncased. The guidance in best practices for fine-tuning LLMs on custom data is useful when the task involves instruction or generative models.
Audit the ONDC data before training
Obtain catalog data through an authorised source and record its provenance, collection date, network domain, seller permissions, and licence. ONDC payloads and downstream exports can differ by implementation, so inspect the actual schema instead of assuming every listing contains reviews, stable seller identifiers, or identical product fields.
Create a data dictionary covering fields such as:
item_id, seller or provider identifier, category, title, description- price, currency, quantity, unit, stock or availability status
- image URLs, attributes, brand, location, language, and timestamps
Remove credentials, contact details, precise personal information, and unnecessary identifiers. Hashing an identifier does not automatically make its use permissible. For high-stakes or sensitive applications, establish a documented review process; the principles behind data veracity infrastructure for high-stakes AI are relevant even when the immediate use case is commerce.
Expect duplicates caused by repeated seller feeds, spelling variants, and near-identical products across cities. Preserve useful regional variation, but do not let one large seller or category dominate the dataset. Keep the raw source read-only and create versioned, transformed copies.
Convert catalog records into a training dataset
For a classification task, create one JSONL record per example with a stable label:
{"text":"Aashirvaad atta 5 kg", "label":"staples/flour"}For attribute extraction, use token spans or an instruction format with a strict output schema. For retrieval, create positive query-product pairs and hard negatives that look similar but are not equivalent. Do not generate labels solely from the same noisy category field you plan to evaluate; sample records for human review and document the labelling rules.
A minimal loading workflow is:
from datasets import load_dataset
files = {
"train": "data/ondc_train.jsonl",
"validation": "data/ondc_validation.jsonl",
"test": "data/ondc_test.jsonl",
}
dataset = load_dataset("json", data_files=files)Split by seller, product family, or time period where possible. A random row-level split can place near-duplicate listings in both training and test sets, producing an inflated score. Keep a challenge set containing regional language, misspellings, long titles, missing attributes, and new categories.
Tokenise and fine-tune with Hugging Face
For sequence classification, map human-readable labels to integer IDs and tokenise without padding every record to the global maximum. Dynamic padding generally saves memory:
from transformers import AutoTokenizer
checkpoint = "ai4bharat/indic-bert"
tokenizer = AutoTokenizer.from_pretrained(checkpoint)
label_names = sorted(set(dataset["train"]["label"]))
label2id = {name: i for i, name in enumerate(label_names)}
def encode(batch):
encoded = tokenizer(
batch["text"],
truncation=True,
max_length=192,
)
encoded["labels"] = [label2id[x] for x in batch["label"]]
return encoded
tokenized = dataset.map(encode, batched=True, remove_columns=dataset["train"].column_names)Use a task-compatible checkpoint and verify its licence, language coverage, tokenizer behaviour, and commercial-use terms. Then configure training with a held-out validation set:
from transformers import AutoModelForSequenceClassification, TrainingArguments, Trainer
model = AutoModelForSequenceClassification.from_pretrained(
checkpoint,
num_labels=len(label_names),
id2label={i: x for x, i in label2id.items()},
label2id=label2id,
)
args = TrainingArguments(
output_dir="outputs/ondc-category-model",
eval_strategy="epoch",
save_strategy="epoch",
learning_rate=2e-5,
per_device_train_batch_size=16,
per_device_eval_batch_size=32,
num_train_epochs=3,
weight_decay=0.01,
load_best_model_at_end=True,
metric_for_best_model="f1",
report_to="none",
)
trainer = Trainer(
model=model,
args=args,
train_dataset=tokenized["train"],
eval_dataset=tokenized["validation"],
)
trainer.train()The exact argument names can vary across Transformers releases, so pin package versions in requirements.txt and log the checkpoint, dataset hash, seed, GPU type, and configuration. If GPU memory is limited, use gradient accumulation, mixed precision where supported, or parameter-efficient methods such as LoRA. Smaller models are often easier to serve and cheaper to retrain; compare them against a strong non-neural baseline.
Evaluate beyond one accuracy number
Report macro-F1, per-class precision and recall, and a confusion matrix. Accuracy can conceal poor performance on smaller categories. For extraction, report entity-level precision, recall, and F1. For search or matching, use recall@k and ranking metrics, then inspect false positives manually.
Evaluate slices by language, geography, category, seller size, text length, and missing-field pattern. Test robustness against code-mixed phrases such as “5 kilo chawal”, transliteration, abbreviations, and inconsistent units. Check calibration if scores drive automatic rejection or ranking. Have reviewers inspect a fixed sample of errors and classify the cause: bad label, ambiguous product, vocabulary gap, truncation, or model error.
A model should not silently invent product attributes, claim availability, or infer sensitive seller information. Keep confidence thresholds conservative and route uncertain cases to review. If your product needs Indian-language generation or multimodal catalogue understanding, compare the text model with open-source vision-language models for Indian languages rather than forcing images into a text-only pipeline.
Publish and deploy responsibly
Save the model, tokenizer, label map, preprocessing code, evaluation report, and dataset version together:
trainer.save_model("outputs/ondc-category-model")
tokenizer.save_pretrained("outputs/ondc-category-model")On Hugging Face Hub, use a private repository until data rights and model-card disclosures are complete. Document intended use, excluded uses, languages, known failure modes, training-data limitations, evaluation slices, and the licence. Never upload raw catalog exports containing restricted information merely to make an experiment reproducible.
For production, validate the model on a fresh catalogue sample, monitor category drift and seller changes, and retain a rollback version. Batch inference may be sufficient for catalogue enrichment; real-time search may require quantisation or a smaller checkpoint. For edge or cost-sensitive deployments, see the practical considerations in AI model optimisation for mobile devices.
A practical launch checklist
- Define one measurable task and a controlled label schema.
- Confirm ONDC data access, licensing, retention, and privacy requirements.
- Deduplicate records and split by seller or product family to prevent leakage.
- Build multilingual and hard-negative evaluation slices.
- Compare a baseline, a multilingual checkpoint, and a smaller production candidate.
- Track macro-F1 or ranking metrics, not only aggregate accuracy.
- Version data, code, model weights, and configuration.
- Add human review, monitoring, and a retraining trigger before automation.
Fine-tuning ONDC catalog data is most valuable when it improves a narrowly defined workflow: cleaner listings, better discovery, or less manual seller effort. Treat the catalog as governed operational data, not a generic web corpus, and Hugging Face becomes a practical path from an India-specific dataset to a model that can be tested, audited, and maintained.