Marketplace search is no longer limited to matching words in a product title. Shoppers describe outcomes, occasions, budgets, sizes, materials, delivery needs, and personal preferences in one sentence—and expect the catalogue to understand them.
Natural-language product discovery for marketplaces combines lexical search, semantic retrieval, structured filters, and learning-to-rank models to interpret that intent. For Indian marketplaces, the system must also handle spelling variation, transliterated Indian languages, code-switching, inconsistent seller data, price formats, and locality-sensitive delivery expectations.
The objective is not to replace the search bar with a chatbot. It is to help a shopper reach a trustworthy shortlist quickly, while giving marketplace teams measurable control over relevance, availability, margin, seller quality, and safety.
Why keyword search underperforms
Keyword retrieval remains valuable for exact identifiers such as brand names, model numbers, SKUs, and product codes. It struggles when the shopper’s language differs from the catalogue language.
Typical failure modes include:
- Vocabulary mismatch: “office wear for a summer wedding” may not match listings described as linen shirts, bandhgalas, or formal trousers.
- Ambiguous intent: “flat shoes for running” and “running shoes for flat feet” require different interpretations.
- Hidden constraints: budget, colour, size, material, use case, delivery deadline, and compatibility may appear anywhere in the query.
- Messy marketplace data: duplicate listings, missing attributes, exaggerated titles, and seller-specific terminology weaken retrieval quality.
- Mobile friction: shoppers do not want to open a long filter panel to express a simple requirement.
A semantic layer improves recall, but semantic similarity alone is not enough. A product can be conceptually related to a query and still be unavailable, too expensive, incorrectly sized, or unsafe to recommend.
The architecture: hybrid retrieval with structured intent
A production system usually has five layers.
1. Catalogue normalisation
Before adding an embedding model, clean the source data. Standardise units, currencies, colour names, sizes, brands, categories, and attributes. Detect duplicate listings and separate seller claims from verified specifications. Preserve the original text, but create a canonical representation for retrieval.
For Indian catalogues, include transliterated forms and common variants: “kurta” and “kurti,” “mobile cover” and “phone case,” or “laal” and “red.” Maintain a controlled vocabulary rather than relying entirely on an LLM to infer attributes at query time.
2. Query understanding
Parse the query into a structured object containing:
- category and subcategory;
- positive attributes, such as waterproof or vegetarian;
- exclusions, such as no leather or not refurbished;
- numerical constraints, including price, size, capacity, and rating;
- use case, occasion, audience, and urgency;
- location or delivery requirements.
For example, “waterproof trekking jacket under ₹5,000, women’s medium, deliver to Pune this week” should produce filters that the inventory system can verify—not merely a vector representation.
Multilingual and code-switched queries need dedicated evaluation. Teams working with Hindi, Tamil, Marathi, Bengali, or Hinglish can use techniques from low-resource Indic natural language processing, but should validate them against their own search logs and category vocabulary.
3. Candidate retrieval
Use several retrieval paths in parallel:
- lexical search for exact terms, brands, and SKUs;
- dense-vector search for semantic similarity;
- attribute and facet filters for hard constraints;
- catalogue rules for availability, compliance, and seller quality.
Merge candidates with a method such as reciprocal rank fusion, then remove products that violate hard requirements. A vector database can support dense retrieval, but it should not become the source of truth for price, stock, delivery promises, or eligibility.
4. Ranking and reranking
A lightweight ranker can combine semantic relevance with business and operational signals:
- attribute match and constraint satisfaction;
- text and image quality;
- stock and delivery confidence;
- historical engagement, adjusted for position bias;
- returns, cancellations, reviews, and seller reliability;
- price competitiveness and marketplace policy.
Use an expensive cross-encoder or language model only on a small candidate set. This keeps latency predictable and makes cost easier to manage. Do not let commercial promotion silently override relevance; label sponsored results and monitor their effect on trust.
5. Explanations and refinement
The interface should show why a result matched: “waterproof,” “under ₹5,000,” “size M,” or “available in Pune.” If the system is uncertain, ask one focused clarification question rather than inventing a preference.
Conversational refinement can then narrow the set: “show black options,” “only cotton,” or “remove products with delivery after Friday.” Keep the session state explicit, auditable, and easy for the user to reset.
Designing for Indian marketplace behaviour
India requires more than translating an English search model. Users may type in Roman script, switch languages mid-query, omit spaces, use local product names, or describe a product through an occasion rather than a category. “Shaadi mein pehenne ke liye simple saree” contains intent that may not appear in the listing title.
Build a query test set from real, consented logs. Include spelling errors, transliteration, local brands, units such as “pauna kilo,” price expressions such as “1k ke andar,” and mixed-language queries. If your product experience includes voice, pair discovery with a robust speech layer; the guidance on natural-sounding TTS for voice agents is relevant when results need to be read back naturally.
Images are equally important in fashion, furniture, beauty, and home categories. A shopper may upload a reference image and ask for a similar style, colour, or silhouette. Vision-language models can help, but every inferred attribute should be treated as a ranking signal unless it has been verified. Explore open-source vision-language models for Indian languages when data residency, cost, or customisation matters.
Metrics that matter
Do not judge discovery by click-through rate alone. Track the complete funnel:
- zero-result rate and reformulation rate;
- search exit and abandonment rate;
- add-to-cart and purchase rate after search;
- median rank of the first relevant result;
- constraint-violation rate;
- long-tail exposure and seller coverage;
- return, cancellation, and complaint rates;
- p50 and p95 latency, plus inference cost per query.
Create labelled query sets by category, language, and difficulty. Measure recall@K and nDCG, then compare offline gains with controlled online experiments. Segment results by new versus repeat users, language, device, geography, and catalogue maturity. A model that improves English fashion queries while degrading Hindi grocery searches is not a marketplace-wide improvement.
A practical 2026 implementation roadmap
Phase one: instrument and clean. Log queries, clicks, reformulations, purchases, and zero-result sessions with privacy controls. Fix catalogue schemas, synonyms, duplicate listings, and hard-filter correctness.
Phase two: launch hybrid retrieval. Keep the existing lexical engine, add multilingual embeddings, and blend candidate sets. Start with high-volume categories where relevance problems are visible and labels are available.
Phase three: add structured query understanding. Extract attributes and constraints using deterministic parsers plus models. Validate every extracted filter against a product schema, and expose confidence thresholds for fallback behaviour.
Phase four: improve ranking. Train on purchases and qualified interactions, correct for position bias, and include seller and fulfilment quality. Test reranking latency before expanding model size.
Phase five: introduce conversational and visual flows. Add iterative refinement or image search only after core retrieval is reliable. For teams deploying open models, how to deploy large language models locally offers useful considerations around privacy, hardware, and operational control.
Common mistakes to avoid
- Embedding unclean seller titles and assuming semantic search will fix them.
- Treating every extracted preference as a hard filter.
- Optimising clicks while ignoring returns and cancellations.
- Sending full catalogue data or personal history to an LLM unnecessarily.
- Launching a chat interface before fixing stock, price, and delivery accuracy.
- Evaluating only English queries or only the top product categories.
The strongest marketplace discovery systems are hybrid, measurable, and deliberately constrained. Use language models to interpret intent, vector retrieval to expand recall, structured systems to enforce facts, and ranking models to balance relevance with trust. For Indian builders, the differentiator will be reliable performance across languages, devices, sellers, and uneven catalogue quality—not a chatbot layer alone.
Build with AI Grants India
If you are building an AI-native marketplace, multilingual commerce search, or catalogue intelligence product in India, apply to AI Grants India. Strong applications should explain the user problem, proprietary data advantage, evaluation plan, infrastructure choices, and the measurable improvement your system will deliver.