What you are building
A multilingual product-search system should do more than match words. It must understand misspellings, transliteration, code-mixed queries, local product names, units, brands, and shopping intent. A query such as “lal cotton kurti under 1000” may combine English, Hindi usage, Latin-script Hindi, and a price constraint. Quantization helps reduce inference cost, but it does not fix weak data or an unsuitable retrieval architecture.
A robust design usually has two stages:
- Candidate retrieval: encode the query and catalogue text, then retrieve a small set of likely products using an approximate nearest-neighbour index.
- Ranking: score those candidates using relevance, availability, price, delivery constraints, and business rules.
Start with a full-precision baseline, establish quality targets, and quantize only after measuring where latency and memory are actually spent. For broader context on language coverage and product decisions, see this guide to building AI apps for the next billion users in India.
Choose the model and retrieval architecture
For most Indian commerce use cases, a compact multilingual or Indic-capable encoder is a better starting point than a large generative model. The encoder should produce embeddings for queries and product fields, while a separate lightweight ranker can combine semantic similarity with structured signals.
Useful options include:
- Dual-encoder retrieval: encode queries and products independently, enabling precomputed catalogue vectors and low-latency search.
- Cross-encoder ranking: pass a query-product pair through a smaller model for better precision on the top 50–200 candidates.
- Hybrid search: combine dense vectors with BM25 or another lexical index. Exact brand names, model numbers, sizes, and spelling variants often perform better lexically.
- Distilled ranking models: train a small student model from a stronger teacher to reduce latency before applying quantization.
For an initial production target, benchmark a 384- or 768-dimensional embedding model, an INT8 vector index, and a compact ranker. Keep catalogue fields separate during experimentation: title, brand, category, attributes, seller, language, and structured specifications. Concatenating everything into one unstructured string can bury important signals.
Build an Indic and commerce-specific dataset
The most valuable training data is not generic multilingual text. It is representative search behaviour from your marketplace, including successful and failed searches.
Create a dataset containing:
- Queries in Devanagari, Bengali, Tamil, Telugu, Kannada, Malayalam, Gujarati, Marathi, Punjabi, and other target scripts.
- Romanised queries such as “sadi”, “kurti chahiye”, or “mobile cover red”.
- Code-mixed Hindi-English and regional-language-English queries.
- Product titles, attributes, category paths, brand aliases, seller text, and catalogue translations.
- Clicks, add-to-cart events, purchases, reformulations, zero-result searches, and abandoned sessions.
- Hard negatives: products from the same category that appear plausible but fail on brand, size, colour, gender, compatibility, or price.
Treat personal data carefully. Remove phone numbers, addresses, account identifiers, and free-text fields that are not needed for relevance training. Sample logs by language, category, geography, and device so that high-volume Hindi or English traffic does not hide poorer performance in lower-resource languages. The low-resource Indic NLP builder’s guide covers additional approaches for scarce labelled data.
Normalisation requires restraint. Preserve numerals, units, model codes, and meaningful punctuation. Build explicit transformations for Unicode normalisation, repeated characters, common spelling errors, transliteration, and script detection. Do not blindly lowercase or strip tokens when case distinguishes a product model or abbreviation.
Train for search relevance, not language fluency
For retrieval, train query and product encoders with a contrastive objective. Positive pairs can come from purchases, strong clicks, or human judgements; negatives should include both random products and hard catalogue alternatives. A useful batch may contain:
- The clicked or purchased product as the positive.
- Same-category products with a different brand or attribute.
- Products matching the words but violating a key constraint.
- Lexically similar products that are unavailable or irrelevant.
For ranking, combine semantic features with structured features such as stock status, price range, delivery location, seller quality, and category compatibility. Keep business rules observable instead of hiding them inside the model. This makes debugging and compliance easier.
Evaluate by language and query type, not only on one aggregate score. Track Recall@K and NDCG@K for retrieval and ranking, plus zero-result rate, reformulation rate, click-through rate, add-to-cart rate, and purchase conversion. Report slices for native script, Romanised input, code mixing, long-tail categories, and new products. A system that improves English conversion while harming Tamil or Marathi discovery is not a successful multilingual launch.
Quantize safely
Quantization maps floating-point weights and activations to lower-precision representations. INT8 is usually the practical first target for CPU or mobile inference; INT4 can reduce memory further but often needs careful calibration and may require hardware-specific kernels.
Use this sequence:
1. Export and benchmark the full-precision model.
2. Apply dynamic quantization to linear layers for a quick baseline.
3. Use static post-training quantization when activation ranges are stable and representative calibration data is available.
4. Try quantization-aware training if post-training quantization causes a material relevance drop.
5. Quantize the encoder, ranker, and vector index independently; they may have different accuracy-cost trade-offs.
Calibration data must reflect real Indian-language traffic. Include scripts, transliteration, code mixing, short queries, spelling errors, numerals, and long-tail categories. Check embedding drift by measuring cosine similarity between full-precision and quantized vectors, but do not treat similarity alone as a quality metric. Re-run retrieval and ranking evaluation after every precision change.
Keep sensitive components at higher precision if necessary. For example, an INT8 encoder with an FP16 or FP32 final scoring layer may preserve ranking quality at modest additional cost. Record the runtime, operator support, batch size, thread count, memory use, and hardware for every benchmark; otherwise results will not be reproducible.
Deploy for Indian traffic patterns
For a catalogue search API, precompute product embeddings whenever product content changes, then update only affected SKUs. Cache popular queries, but avoid caching results without considering inventory, location, price, and delivery changes. Use a hybrid index to handle both semantic discovery and exact matching.
A production service should include:
- Script and language detection with a safe fallback for mixed queries.
- Query rewriting for spelling and transliteration, while retaining the original query for auditability.
- Timeouts and fallbacks from dense retrieval to lexical search.
- Monitoring for p50, p95, and p99 latency, index freshness, CPU or accelerator utilisation, and memory pressure.
- Per-language dashboards for zero-result rate, Recall@K, relevance complaints, and conversion.
- Versioned models, calibration sets, indexes, and rollback procedures.
If you expose the model through multiple services, design clear contracts for embeddings, ranking, catalogue updates, and feedback events. Distributed-system design becomes important as traffic and catalogue size grow; the principles in building distributed systems with AI agents are also relevant to service boundaries, observability, and failure handling.
A practical launch plan
Begin with two or three high-volume languages plus English, and include Romanised input from day one. Establish a labelled test set before training, then compare four versions: lexical baseline, full-precision dense model, quantized dense model, and hybrid production candidate. Launch behind an experiment flag and measure both relevance and commercial outcomes.
Do not expand language coverage solely because the encoder supports a script. Add a language when you have enough catalogue coverage, evaluation examples, moderation support, and a plan for collecting feedback. Refresh the calibration set as new brands, categories, and user vocabulary appear.
The best result is not the smallest model. It is the lowest-cost system that maintains fair, measurable search quality across the languages and shopping behaviours your customers actually use.