Seller-support AI has to work under practical constraints: high query volumes, mixed Hindi-English conversations, regional languages, intermittent connectivity, strict latency expectations, and sensitive business data. A quantized model can reduce memory use and inference cost, but quantization is not a substitute for good product design or reliable retrieval.
The strongest approach is usually a small language model or classifier paired with a trusted knowledge base, clear escalation rules, and human review. This guide explains how to build that system for Indian marketplaces and seller platforms.
Start with a narrow seller-support job
Do not begin by training a general chatbot. Choose one measurable workflow, such as:
- Explaining listing, pricing, or catalogue requirements
- Classifying support tickets by issue and urgency
- Guiding sellers through order, return, refund, or payment workflows
- Translating or summarising a seller’s issue for a support agent
- Detecting when a query needs escalation to a human team
Define the model’s boundaries before collecting data. It should know which questions it can answer, which actions require authentication, and which cases must be transferred. For voice-first workflows, review the architecture in this voice agent guide, particularly its treatment of streaming, tool calls, and fallback handling.
Set success metrics that reflect operations rather than model benchmarks alone:
- Correct intent classification and routing
- Resolution rate without a human handoff
- Citation or policy-grounding accuracy
- Median and p95 response latency
- Cost per resolved conversation
- Escalation precision and seller satisfaction
Build an India-relevant data pipeline
Useful training and evaluation data may include anonymised chat transcripts, support tickets, policy documents, seller FAQs, product-category rules, and agent-written resolutions. Label each example with the intent, language, domain, urgency, required action, and resolution status. Retain the original message and a cleaned version so that the system is tested on real spelling, code-switching, and shorthand.
Indian seller queries often combine English with Hindi or another Indic language, use Romanised text, and contain marketplace-specific abbreviations. Create evaluation slices for Hindi-English, Tamil-English, Bengali-English, Romanised Indic text, speech transcripts, low-literacy phrasing, and noisy messages. The low-resource Indic NLP guide is useful when selecting tokenisation, transliteration, and language-identification strategies.
Before training, remove or mask phone numbers, bank details, addresses, GST-related identifiers, order IDs, and other personal or commercially sensitive information. Keep a documented consent, retention, and access policy. Never use raw support logs as a training set without checking for data leakage between train, validation, and test splits.
Choose the smallest model that meets the job
A seller-support system may need several models rather than one large model:
- A language identifier and safety filter
- An intent or urgency classifier
- A compact generative model for drafting answers
- An embedding model for retrieving relevant policy passages
- A reranker or confidence model for deciding whether to answer
For policy-heavy support, retrieval-augmented generation is generally safer than relying on model memory. Store versioned policy documents, retrieve the relevant passages, and require the response to cite or quote the source internally. Use deterministic workflows for actions such as changing bank details, issuing refunds, or modifying listings.
Select models based on Indic-language quality, licence terms, context length, hardware support, and operational cost. Test the candidate model on your own seller-support set before committing to fine-tuning. For teams building broader products for Indian users, this guide to AI apps for the next billion users in India covers reliability and access considerations that also apply here.
Establish a float32 baseline
Train or fine-tune the unquantized model first. Record its quality, memory footprint, throughput, and latency on the exact hardware you expect to use. This baseline lets you distinguish quantization loss from problems in data, prompting, retrieval, or serving.
Use separate test sets for common questions, rare intents, adversarial prompts, policy changes, and each priority language. Track more than accuracy. For generative responses, evaluate factuality, policy compliance, completeness, tone, language correctness, and whether the answer gives an unsafe or unauthorised instruction.
Create a fixed golden set of production-like conversations. Include difficult examples such as ambiguous refund requests, sellers switching languages mid-message, incomplete order references, and queries where the correct response is an escalation.
Select a quantization method
Quantization maps higher-precision values, commonly FP32, to lower-precision representations such as INT8 or INT4. The right method depends on the model and runtime.
- Dynamic post-training quantization is a quick starting point for some CPU-based models; weights are quantized while selected activations are handled at runtime.
- Static post-training quantization uses representative calibration data to quantize weights and activations. It can deliver predictable INT8 performance but requires a representative calibration set.
- Quantization-aware training (QAT) simulates lower-precision behaviour during training and is often preferable when post-training quantization causes unacceptable quality loss.
- Weight-only quantization can reduce memory for generative models while keeping activations at higher precision, depending on the serving stack.
Use real seller-support samples for calibration, balanced across languages, intents, message lengths, and frequent entities. Do not calibrate only on clean English text. Export to a runtime supported by your target hardware, such as a mobile, CPU, GPU, or edge inference stack, and verify that every operator is actually accelerated rather than silently falling back to a slower implementation.
Evaluate quality and efficiency together
Compare the quantized model directly with the float32 baseline. Report:
- Intent F1 and per-language recall
- Grounded-answer accuracy and unsupported-claim rate
- Escalation recall for high-risk cases
- First-token and end-to-end latency at p50 and p95
- Requests per second under realistic concurrency
- Peak RAM or VRAM usage
- Model size, energy use, and cost per 1,000 conversations
Pay special attention to language-specific degradation. A model can show a small overall accuracy drop while becoming materially worse for Romanised Hindi or a lower-resource language. If quality falls, try better calibration data, selective rather than universal quantization, QAT, a stronger tokenizer, shorter prompts, or routing difficult queries to a higher-precision model.
Deploy with retrieval, controls, and fallbacks
A production request path should authenticate the seller, classify the issue, retrieve current policy, generate or select a response, validate the output, and log an auditable result. Keep model inference separate from transactional systems. The model may recommend an action, but permissions and business rules should execute it.
Use confidence thresholds and abstention. Escalate when the model lacks evidence, detects sensitive financial or account issues, receives conflicting policy information, or cannot identify the seller’s intent. A human support console should show the original message, language, retrieved policy, model response, and reason for escalation.
For high-volume systems, queueing, caching, batching, and regional deployment can matter as much as quantization. If multiple specialised models and tools coordinate, principles from building distributed systems with AI agents can help—but avoid adding agent complexity to a workflow that a classifier and retrieval pipeline can solve.
Monitor after launch
Create dashboards for quality, latency, cost, and safety. Monitor drift in languages, product categories, policy versions, and new seller terminology. Sample conversations for human review, with access controls and redaction. Compare quantized and fallback-model decisions on a continuing shadow-traffic basis.
Retrain or recalibrate when policy changes, new categories launch, or error rates rise. Maintain model cards and deployment records covering training data, quantization method, calibration set, known limitations, evaluation slices, and rollback procedures. For voice support, separately monitor transcription errors, interruptions, accents, and language switching; a comparison of voice agents and IVR provides useful product trade-offs.
A practical launch checklist
- Define one seller workflow and explicit escalation boundaries.
- Build privacy-safe, multilingual, production-like datasets.
- Establish a float32 quality and performance baseline.
- Calibrate on representative Indian seller queries.
- Compare dynamic, static, weight-only, or QAT approaches.
- Test per-language quality, not just aggregate scores.
- Verify runtime acceleration on target hardware.
- Add retrieval, permissions, abstention, and human handoff.
- Run a limited pilot before full rollout.
- Monitor drift and keep a tested rollback path.
Quantization is valuable when it improves the economics and responsiveness of a well-scoped support product. For Indian seller platforms, the winning system is not simply the smallest model; it is the smallest model that remains trustworthy across languages, devices, policies, and real support conditions.