Indian ecommerce support teams need AI that is fast, affordable, multilingual, and dependable under heavy traffic. A quantized model can reduce inference cost and latency, but quantization alone does not create a useful support system. The strongest production design combines a compact language model with retrieval, business tools, clear escalation rules, and evaluation on real Indian customer queries.
This guide explains how to build that system, with an emphasis on catalogues, order workflows, returns, COD issues, delivery exceptions, and support across English, Hindi, Hinglish, and other Indic languages.
Start with the support problem, not the model
Define the tasks the assistant must handle before selecting a base model. Common ecommerce intents include:
- Order tracking, delivery delays, and address changes
- Returns, refunds, cancellations, replacements, and warranty questions
- Cash-on-delivery eligibility, payment failures, and invoice requests
- Product availability, size or compatibility questions, and catalogue discovery
- Seller, logistics, and damaged-package complaints
- Handoff to a human agent for sensitive, ambiguous, or high-value cases
Create an intent and entity schema. For example, a return request may require order_id, product_id, purchase date, reason, and eligibility status. Keep these fields separate from the natural-language response so your backend can validate them and call the correct order or returns API.
Set measurable targets early: first-response latency, resolution rate, escalation accuracy, hallucination rate, cost per conversation, and customer satisfaction. A smaller model that correctly resolves routine requests and escalates edge cases is more valuable than a larger model that produces fluent but unsafe answers.
Build an India-ready dataset
Use historical chat and email transcripts only after removing phone numbers, addresses, payment details, order identifiers, and other personal information. Obtain the permissions required for your data sources and define retention rules before training or evaluation.
Label examples for intent, language, sentiment, entities, policy outcome, and whether the answer requires a tool call. Include difficult cases rather than only clean conversations:
- Hinglish, code-switching, spelling variation, and transliterated Hindi
- Regional language phrases and informal marketplace vocabulary
- Short messages such as “parcel nahi aaya” or “return kab tak?”
- Multiple issues in one turn
- Conflicting catalogue or order information
- Abusive, fraudulent, or prompt-injection attempts
For broader Indic coverage, review the principles in this low-resource Indic NLP guide. Do not assume that translation quality is uniform across languages. Test each target language separately, including native-script and Romanised inputs.
Split data by conversation or customer, not by individual message. Keep a private test set containing recent production-like queries. This prevents near-duplicate messages from appearing in both training and evaluation and gives you a more honest estimate of performance.
Choose the right architecture
For most support deployments, begin with a compact instruction-tuned language model rather than training from scratch. Use retrieval-augmented generation (RAG) for information that changes frequently, including product specifications, delivery promises, return policies, and order status. Store approved documents with metadata such as language, region, category, effective date, and policy version.
Use tools for facts the model must not invent. Typical tools include:
- Order-status and shipment-tracking APIs
- Return and cancellation eligibility checks
- Product catalogue and inventory search
- Refund-status and payment reconciliation services
- Agent handoff and ticket creation
A reliable request path is: classify intent, extract entities, retrieve policy or catalogue context, call an authorised business tool where needed, generate a concise response, and validate the result against policy. For complex workflows, an orchestration layer is safer than giving the model unrestricted access to internal systems. Teams designing agent-based workflows can compare this approach with distributed systems using AI agents.
For voice support, keep speech recognition, dialogue orchestration, and text-to-speech as separate components. Quantizing the language model may reduce cost, but voice quality also depends on turn-taking, latency, transcription accuracy, and interruption handling. See the voice-agent architecture guide before adding a phone channel.
Quantize the model deliberately
Benchmark the full-precision baseline first. Record response quality, peak memory, tokens per second, time to first token, and cost on the hardware you expect to use. Then compare quantization approaches:
- Post-training dynamic quantization: quick to test and useful for some CPU workloads.
- Post-training static quantization: calibrates activations using representative data and can improve predictable deployment performance.
- Weight-only quantization: often a practical choice for generative models, with formats such as 8-bit or 4-bit weights.
- Quantization-aware training: simulates lower-precision behaviour during fine-tuning when post-training loss is unacceptable.
Use a calibration set that reflects actual support traffic: long and short chats, code-switching, product names, order numbers, and policy-heavy questions. Compare FP16 or BF16, INT8, and lower-bit variants on the same hardware. Lower precision is not automatically better; a small quality loss in language understanding may create a large increase in incorrect refunds or failed escalations.
Validate numerical and operational compatibility with your serving stack. Depending on the model and hardware, practical options may include PyTorch, ONNX Runtime, TensorRT-LLM, llama.cpp, or a managed inference service. Measure end-to-end latency rather than relying on model benchmarks alone.
Evaluate support quality, safety, and cost
Create a test suite organised by business risk, not just language fluency. Track:
- Intent classification and entity-extraction accuracy
- Retrieval recall and citation or source-grounding accuracy
- Correct tool selection and argument validation
- Policy compliance for refunds, returns, cancellations, and warranties
- Hallucination and unsupported-claim rate
- Escalation precision and recall
- Performance by language, script, region, device, and customer segment
- Latency, throughput, memory use, and cost per resolved conversation
Run adversarial tests for prompt injection, fake order numbers, social-engineering attempts, personally identifiable information, and requests to bypass policy. The assistant should refuse unsupported actions and route uncertain cases to a human. Never allow a generated message alone to approve a refund, modify an address, or expose account information.
Use a shadow deployment before switching traffic. Compare the quantized model with the existing system on identical conversations, review failures with support agents, and launch gradually with feature flags. Monitor drift after catalogue, policy, logistics, or pricing changes.
Design for Indian operating conditions
Support language selection explicitly and allow customers to switch languages without restarting the conversation. Keep responses short on low-bandwidth connections, preserve order numbers exactly, and provide a clear human-support path. If the business serves multiple states, test terminology and policy differences rather than applying one translated script everywhere.
For mobile and cost-sensitive environments, a compact model can run near the user or on a lower-cost CPU service. Larger models may remain useful for difficult cases, with the quantized model handling routine intents. This tiered design supports the broader goal of building AI apps for the next billion users in India.
Document model cards, dataset sources, known failure modes, access controls, and rollback procedures. Encrypt logs, minimise stored conversation content, restrict tool permissions, and define who can approve model and policy changes.
A practical deployment checklist
Before production, confirm that you can:
- Reproduce the quantization and evaluation pipeline
- Roll back to the previous model without data migration
- Update policies and catalogue content without retraining the model
- Audit every tool call and sensitive decision
- Monitor quality by intent and language, not only aggregate accuracy
- Escalate unresolved or high-risk cases to trained agents
- Budget inference, storage, observability, and human-review costs
Quantization is best treated as an optimisation stage inside a broader support architecture. Start with a narrow set of high-volume intents, ground answers in current business data, measure failures by language and risk, and expand only when the system is reliable. That approach produces a faster and cheaper assistant without sacrificing customer trust.