0tokens

Apply for AI Grants India

Financial support for innovators building the future of AI in India.

Apply now

Chat · how to build a quantized model for ondc sellers

How to Build a Quantized Model for ONDC Sellers

  1. aigi

    ONDC sellers rarely need a large model for every task. Product categorisation, catalogue-quality checks, demand forecasting, search ranking, fraud signals, delivery estimates, and customer support often run more reliably when the model is small, fast, and inexpensive to serve. Quantization helps achieve that by representing model weights and activations at lower numerical precision, commonly INT8 instead of FP32.

    For an ONDC seller, the goal is not simply to shrink a model. It is to meet a measurable business target: respond quickly to a buyer query, classify a product accurately across Indian languages, or generate a forecast within a fixed cloud or device budget. This guide explains how to build, test, and operate a quantized model for those conditions.

    Start with the ONDC use case

    Define the task and its operating constraints before selecting a framework. A catalogue classifier, a recommendation model, and a voice-based support assistant have different latency, accuracy, and hardware requirements. If the system serves multilingual customers, review the guidance on low-resource Indic natural language processing before preparing the dataset.

    Write down:

    • Input: product title, description, image, price, location, order history, or customer message.
    • Output: category, ranking score, demand estimate, moderation label, or generated response.
    • Success metric: F1 score, recall, NDCG, mean absolute error, conversion rate, or support resolution rate.
    • Latency budget: for example, p95 inference below 100 milliseconds for an API.
    • Deployment target: CPU-only cloud instance, seller laptop, Android device, or edge gateway.
    • Failure cost: a wrong category may reduce discoverability, while a wrong fraud decision may block a legitimate order.

    Keep seller and buyer data separate from model artefacts. Remove unnecessary personal information, define retention rules, and record consent and access controls. ONDC integrations also depend on network participants and API contracts, so keep the model behind a versioned service rather than coupling business logic directly to a model file.

    Choose the smallest suitable baseline

    Train or fine-tune a full-precision baseline first. This gives you a reference for accuracy, memory, throughput, and latency. Do not quantize an untested model and assume any performance change is caused by quantization.

    Good candidates include:

    • Tree-based models for tabular pricing, demand, and fraud features. Quantization may be less important than efficient feature processing.
    • Compact CNN or vision transformers for catalogue images and quality checks.
    • Small transformer encoders for product search, classification, and multilingual text.
    • Distilled or compact language models for seller support and intent detection.

    For Indian commerce, test code-mixed text, transliterated Hindi, Tamil, Telugu, Bengali, Marathi, and spelling variations. Product names may include local brands, abbreviations, units, and regional terms. A model that performs well on clean English data can lose accuracy after deployment.

    Prepare representative training and calibration data

    Quantization depends heavily on the examples used to estimate activation ranges. Create three separate datasets:

    • Training data for fitting the model.
    • Validation data for architecture and hyperparameter decisions.
    • Calibration data for estimating quantization scales and zero points.

    Calibration data does not need labels, but it must resemble production traffic. Include high- and low-priced products, long and short descriptions, images from different phones, regional language inputs, missing fields, and seasonal categories. A few hundred to a few thousand carefully selected examples can be more useful than a large but narrow sample.

    Avoid leakage. Do not calibrate only on the easiest or most frequent category. Also preserve rare but commercially important cases, such as medicines, fresh food, jewellery, or local services, where an error has a larger impact.

    Select a quantization method

    Post-training weight quantization

    Convert weights from FP32 to INT8 or another lower precision after training. This is the quickest approach and often provides a useful reduction in model size with limited accuracy loss. It is a strong first experiment for CPU deployment.

    Dynamic quantization

    Weights are stored in lower precision while some activations are quantized during inference. It is straightforward for many language and recurrent models, but latency gains depend on the operators and hardware runtime.

    Full integer quantization

    Quantize both weights and activations using representative calibration data. This is usually better for predictable CPU or edge performance, but unsupported operations can force parts of the graph back to floating point.

    Quantization-aware training

    Simulate quantization during training so the model learns to tolerate reduced precision. Use this when post-training methods cause unacceptable degradation, particularly in sensitive classification layers or compact vision models. It requires more engineering time but can recover accuracy.

    For transformer workloads, also consider distillation, pruning, or a smaller architecture before aggressive quantization. Combining several optimisations is useful, but change one variable at a time so the effect remains measurable.

    Implement with a production runtime

    Common deployment paths include PyTorch with an appropriate backend, TensorFlow Lite for mobile and edge targets, and ONNX Runtime for portable CPU inference. Export the baseline model first, then verify that preprocessing and outputs match across frameworks.

    A practical workflow is:

    1. Train and save a reproducible FP32 checkpoint.
    2. Export it to the target format, such as ONNX or TFLite.
    3. Apply dynamic, static, or weight-only quantization.
    4. Run the same test records through both models.
    5. Compare numerical outputs, business metrics, and latency.
    6. Package preprocessing, tokenizer, labels, and model version together.

    Do not benchmark only average latency. Measure p50, p95, and p99 latency, cold-start time, memory use, batch size, CPU utilisation, and throughput. Test on the actual hardware used by sellers or the serving cluster. A model that is smaller on disk may still be slower if its runtime lacks optimised INT8 kernels.

    Evaluate accuracy and business impact

    Compare the quantized model against the FP32 baseline and a simple non-AI fallback. Review both aggregate and slice-level results:

    • Category-level precision and recall.
    • Performance by language, region, product type, and seller size.
    • Error rates for missing, noisy, or code-mixed fields.
    • Ranking quality for new and low-inventory products.
    • Forecast error during promotions and seasonal peaks.
    • API latency, memory, cost per 1,000 inferences, and failure rate.

    Set a release threshold before testing. For example, accept a maximum one-percentage-point drop in macro-F1 only if p95 latency improves by 30% and infrastructure cost falls materially. Route uncertain predictions to rules, human review, or the FP32 model. This is safer than forcing every input through the compressed model.

    Deploy progressively and monitor drift

    Start with a shadow deployment: send production-shaped requests to the quantized model without using its decisions. Then run a small canary for selected sellers, compare outcomes, and expand gradually. Keep the FP32 model and a rules-based fallback available for rollback.

    Monitor:

    • Input distribution and missing-field rates.
    • Prediction confidence and abstention rates.
    • Language and category mix.
    • Latency and infrastructure cost.
    • Seller corrections, buyer complaints, and conversion changes.
    • Quantized-versus-baseline disagreement.

    Retrain or recalibrate when catalogues, language patterns, inventory, or ONDC traffic changes. Store model, dataset, tokenizer, calibration-set, and runtime versions so every prediction can be traced to a release.

    Build for India’s next billion users

    ONDC sellers range from digitally mature brands to small businesses operating through modest phones and intermittent connectivity. That makes offline support, compact downloads, low-memory operation, and clear fallback behaviour important. For customer-facing systems, pair compact text models with carefully designed retrieval and escalation flows. If your product is voice-led, review this voice agent architecture and deployment guide and test Indian accents, code-switching, and noisy environments separately.

    Teams building broader seller tools can also learn from approaches for AI apps for the next billion users in India: minimise data transfer, make errors recoverable, and design for assisted rather than fully automated workflows.

    A practical release checklist

    Before production, confirm that:

    • The FP32 baseline and quantized model are evaluated on the same frozen test set.
    • Calibration data represents real seller and buyer traffic.
    • INT8 operators are actually used on the target hardware.
    • Accuracy thresholds are defined by business slice, not only overall average.
    • Preprocessing and postprocessing are versioned with the model.
    • Privacy, retention, and access controls are documented.
    • Canary, rollback, fallback, and monitoring paths have been tested.
    • Model updates can be reproduced and audited.

    Quantization is valuable when it improves the complete ONDC workflow—not merely the model file size. Start with a narrow seller problem, measure the baseline, validate on representative Indian data, and ship the smallest model that meets the accuracy and reliability bar.

    Last updated 23 September 2026

AIGI may be inaccurate. Replies seeded from the guide above.