0tokens

Apply for AI Grants India

Financial support for innovators building the future of AI in India.

Apply now

Chat · how to build a quantized model for return and refund queries in india

How to Build a Quantized Model for Return and Refund Queries in India

  1. aigi

    What you are building

    A return-and-refund assistant should do more than generate a polite reply. It must identify the customer’s intent, extract order details, check the applicable policy, calculate the next step, and escalate cases it cannot safely resolve. Quantization helps make this system cheaper and faster to run, but it does not replace good data, policy controls, or human review.

    A practical architecture has four parts:

    • Intent classification: Identify requests such as return eligibility, refund status, damaged product, missing item, cancellation, replacement, or failed pickup.
    • Entity extraction: Capture order ID, product, purchase date, payment method, language, and requested resolution.
    • Policy retrieval: Fetch the current seller, marketplace, category, and payment-policy rules instead of storing them only in model weights.
    • Action and escalation: Answer, request missing information, create a support ticket, or route the case to an agent.

    For a broader view of multilingual product design, see this guide to building AI apps for the next billion users in India.

    Start with an India-specific dataset

    Collect anonymised conversations from chat, email, call transcripts, WhatsApp workflows, and support tickets. Include English, Hindi, Hinglish, and the regional languages relevant to your customer base. Indian users may switch scripts within one message, use transliterated Hindi, or describe dates and amounts informally: “parcel kal aaya,” “refund kab tak,” or “UPI se paisa nahi mila.”

    Create a label taxonomy before training. A useful first version might include:

    • Return eligibility and policy explanation
    • Refund pending, partial, failed, or reversed
    • Damaged, defective, incorrect, or missing item
    • Pickup rescheduling and failed pickup
    • Cancellation before and after dispatch
    • Replacement request
    • Payment-method-specific issue, including UPI, cards, wallets, and cash on delivery
    • Fraud, abuse, legal threat, or vulnerable-customer escalation

    Keep separate fields for intent, urgency, language, sentiment, and required action. Do not use demographic attributes unless there is a documented operational need. Remove names, phone numbers, addresses, payment credentials, and unnecessary order information from training data. Maintain a locked test set containing real linguistic variation but no personally identifiable information.

    Low-resource languages need deliberate evaluation rather than assumptions. This builder’s guide to low-resource Indic NLP covers data, tokenisation, and evaluation choices that apply directly to this use case.

    Choose a small model and a retrieval layer

    For most support deployments, begin with a compact encoder or instruction-tuned language model rather than fine-tuning a large general-purpose model. A classifier can handle intent and routing, while a smaller generative model drafts the response using retrieved policy text. This separation makes errors easier to detect and policies easier to update.

    Use retrieval-augmented generation for information that changes frequently:

    • Return windows by product category
    • Seller-specific exclusions
    • Refund timelines by payment rail
    • Pickup and logistics rules by pin code
    • Marketplace policy changes
    • Tax, invoice, or replacement requirements

    Store policy documents with effective dates, geography, seller scope, and source references. Require the assistant to cite the policy record internally before making a claim. Never let a model invent a refund date or confirm that money has been credited without checking the order and payment systems.

    Apply quantization deliberately

    Quantization converts higher-precision weights and, in some approaches, activations into lower-precision representations. For a support model, compare at least three approaches:

    • Dynamic post-training quantization: A fast baseline for CPU inference, particularly for classifiers and smaller transformer models.
    • Static post-training quantization: Uses representative calibration data to quantize activations and weights; it can improve latency and memory use but requires careful calibration.
    • Quantization-aware training: Simulates lower precision during training and is useful when post-training quantization causes unacceptable accuracy loss, especially for code-switched or low-resource inputs.

    Start with INT8 for a production baseline. Test INT4 only when the memory or cost benefit justifies a more demanding quality review. Export through the runtime you plan to operate—such as ONNX Runtime, TensorRT, or a mobile/edge-compatible stack—and benchmark the actual hardware. A model that is smaller on paper may not be faster if the target runtime lacks efficient kernels.

    Calibrate with representative queries, not generic text. Include misspellings, transliteration, short messages, angry complaints, mixed languages, and messages containing order numbers. Keep a full-precision model as a reference so you can identify which intents degrade after quantization.

    Train and evaluate the complete workflow

    Split data by customer, order, and conversation thread to prevent leakage. A random message-level split can make results look strong while testing near-duplicates. Track macro-F1 for intent classification, entity-level precision and recall, calibration error, retrieval accuracy, grounded-answer rate, and escalation recall.

    Create separate slices for:

    • English, Hindi, Hinglish, and each supported regional language
    • Roman and native scripts
    • Short versus multi-turn queries
    • COD, UPI, card, and wallet payments
    • Damaged goods, high-value orders, and policy exceptions
    • New sellers and recently changed policies

    For refunds, false reassurance is more serious than a slow handoff. Set thresholds so uncertain or high-risk cases escalate. Test whether the model can say it lacks enough information, ask one useful clarifying question, and avoid exposing another customer’s data. Run adversarial tests for prompt injection, policy conflicts, fabricated order status, and attempts to bypass return controls.

    Deploy with guardrails

    Use a deterministic orchestration layer around the model. The model may classify, extract, and draft; application code should verify identity, query order systems, enforce eligibility rules, and execute refunds. Record the policy version, model version, quantization format, retrieved evidence, confidence score, and final action for every case.

    A reliable response flow is:

    1. Detect language and intent.
    2. Authenticate or request the minimum safe identifier.
    3. Retrieve order and policy data.
    4. Ask for missing information if required.
    5. Provide the next step and expected timeline only when verified.
    6. Offer escalation for exceptions or low confidence.

    Protect customer data with encryption, access controls, retention limits, and redaction in logs. Align the deployment with applicable Indian privacy and consumer-protection obligations, and provide a clear human-support route. Voice support requires additional latency and transcription checks; these voice-agent architecture guides explain how to design that layer.

    Monitor after launch

    Track resolution rate, escalation rate, containment, average handling time, p95 latency, cost per conversation, language-wise quality, refund-error rate, and customer recontact. Review a sample of conversations weekly, prioritising cases where the customer contacted support again or an agent corrected the model.

    Retrain when policies, products, payment flows, or language patterns change. Do not automatically learn from every conversation: route new examples through annotation, privacy review, and regression testing first. Maintain rollback versions for both the model and policy index.

    A practical launch plan

    Begin with three to five high-volume intents and read-only answers. Add authenticated order lookup next, then controlled ticket creation. Introduce automated refund actions only after the system demonstrates stable performance across languages, payment methods, and exception cases. Quantization should be judged by end-to-end business quality, not model size alone: a smaller model is valuable when it remains accurate, grounded, observable, and safe to operate.

    Last updated 23 September 2026

AIGI may be inaccurate. Replies seeded from the guide above.