Intent recognition determines what a user is trying to accomplish: checking a refund, changing an address, reporting fraud, or asking for a human agent. It is the decision layer between a message and the workflow that should handle it. When that layer is unreliable, even a capable LLM produces irrelevant answers, triggers the wrong action, or creates avoidable support escalations.
For teams asking how to improve intent recognition in conversational AI, the highest returns usually come from better intent design, representative data, calibrated confidence, and disciplined evaluation—not from simply switching to a larger model. The approach below is designed for Indian products that must handle English, Hindi, Hinglish, regional languages, spelling variation, and short or ambiguous messages.
Start with a usable intent taxonomy
An intent taxonomy should reflect the actions your system can actually take. Avoid creating separate intents for every wording variation. “Refund status”, “refund delayed”, and “when will my money arrive?” may belong to one intent if they use the same workflow and response policy. Split them only when the backend action, permissions, escalation path, or risk level differs.
For each intent, document:
- A clear name and one-sentence definition.
- Positive examples and realistic paraphrases.
- Similar or competing intents.
- Required entities, such as order ID, location, date, or account type.
- The action, API, or knowledge source that handles it.
- Safe fallback behaviour when required information is missing.
Keep intent recognition separate from entity extraction. The former identifies the user’s goal; the latter captures details needed to complete it. For short, ambiguous messages, see this guide to intent extraction in short text.
Build a representative, multilingual dataset
Model quality cannot exceed the quality and coverage of its examples. Collect utterances from support tickets, search logs, failed chatbot sessions, call transcripts, and user research—not only phrases written by the product team.
Include variation in:
- Formal and informal language.
- Typos, abbreviations, emojis, and incomplete sentences.
- Different levels of English proficiency.
- Code-switching, such as “Mera refund kab aayega?” or “card block karna hai”.
- Regional language inputs and transliterated Hindi, Marathi, Tamil, Bengali, or Telugu.
- Voice-transcribed text, which often lacks punctuation and contains ASR errors.
For voice products, intent recognition is downstream of speech recognition. Test the complete pipeline, including noisy audio, accents, code-switching, and number transcription. The resources on AI speech recognition for Indian regional languages and Hindi ASR low WER are useful starting points.
Synthetic paraphrases can expand coverage, but they should not replace real user data. Generate examples with explicit constraints—such as slang, code-mixing, or a particular customer segment—then have reviewers remove unnatural or duplicated samples. Keep synthetic and human-authored examples separately tagged so you can measure their effect.
Use a hybrid classification architecture
No single recogniser is ideal for every request. A production system can combine several layers:
1. Rules and deterministic checks: Use these for safety-critical phrases, authentication states, regulatory disclosures, and high-confidence commands.
2. A supervised classifier: A compact transformer or multilingual model is fast and economical for frequent, well-defined intents.
3. Embedding retrieval: Compare the message with labelled examples or intent prototypes to handle paraphrases and new wording.
4. LLM adjudication: Send ambiguous cases to an LLM with a constrained intent list, relevant conversation context, and a strict JSON schema.
Do not route solely on a raw model probability. Classifier scores are often poorly calibrated and cannot be compared across models without testing. Set thresholds using a validation set and consider separate thresholds for high-risk actions. A payment reversal should require more evidence than a request for store timings.
Latency and cost also matter. For Indian customer-service deployments, low-latency conversational AI explains why local caching, small models, early exits, and carefully limited LLM calls can be as important as accuracy.
Improve semantic matching with embeddings
Embeddings help match meaning rather than exact words. Store multiple approved examples for each intent, retrieve the nearest candidates, and then apply a reranker or classifier to distinguish close neighbours. For example, “I want to close my account” and “terminate my subscription” may be semantically related but require different workflows.
Use a two-stage design:
- Retrieve the top five to ten candidate intents with multilingual embeddings.
- Rerank or classify those candidates using a cross-encoder, supervised model, or constrained LLM.
Evaluate the similarity threshold on real traffic. A nearest neighbour is not necessarily a correct neighbour: every intent needs an explicit rejection boundary and competing examples. Store the matched examples and scores for debugging, while protecting personal data through masking and retention controls.
Handle context without letting it dominate
A message such as “yes”, “do it”, or “the second one” has no reliable intent in isolation. Pass structured conversation state to the recogniser: the active workflow, last question, selected entity, authentication status, and unresolved slots. Use only the context needed for the decision; sending an entire transcript can introduce irrelevant or contradictory signals.
Resolve follow-up messages against the current state, but allow users to change topics. A user may answer a verification question and then ask about a refund. Add an explicit topic-switch detector and prioritise a new, high-confidence intent over stale state.
For voice systems, turn-taking and interruption handling are part of intent quality. A practical voice agent architecture should distinguish a barge-in such as “no, cancel that” from background speech or an incomplete utterance.
Make out-of-scope detection a first-class task
A classifier forced to choose among known intents will confidently produce wrong answers. Create an out-of-scope or none-of-the-above pathway and train it with:
- Unrelated questions.
- Industry-adjacent requests your system cannot fulfil.
- Greetings and social conversation where no workflow is required.
- Ambiguous examples between two intents.
- Adversarial prompts and prompt-injection attempts.
Use an abstention policy: if confidence is low, candidates conflict, or required context is absent, ask a focused clarification question or route to an agent. The correct response to uncertainty is not a guess.
Evaluate the system by business risk
Overall accuracy hides the failures that matter. Build a test set that is frozen before each release and report:
- Precision, recall, and F1 for every intent.
- Confusion between similar intents.
- OOS precision and recall.
- Accuracy by language, script, channel, device, and customer segment.
- Clarification rate, handoff rate, task completion, and repeat-contact rate.
- Latency, token use, and cost per resolved conversation.
Review false positives separately from false negatives. A false positive can trigger an irreversible action; a false negative may only cause a clarification. For sensitive domains such as finance, healthcare, or mental health, require human review and conservative thresholds. This is especially important when building conversational AI for mental health in India.
Close the production feedback loop
Log the input, predicted intent, candidate scores, language signal, context state, final outcome, and human correction—subject to consent, masking, and retention policies. Sample conversations weekly, prioritising low-confidence cases, repeated clarifications, escalations, and high-impact errors.
Use active learning to label the most informative examples rather than randomly adding more data. Before retraining, deduplicate near-identical utterances, check label consistency, and test whether a new intent overlaps with an existing one. Release model and taxonomy changes behind versioned evaluations and monitor for drift after deployment.
Practical implementation checklist
- Define intents by workflows and policies, not by wording.
- Collect real multilingual and code-mixed examples.
- Separate intent classification from entity extraction.
- Combine deterministic rules, supervised models, embeddings, and selective LLM adjudication.
- Calibrate thresholds and support abstention.
- Include conversation state while detecting topic changes.
- Test OOS queries, adversarial inputs, ASR noise, and transliteration.
- Measure per-intent and per-language performance, not only aggregate accuracy.
- Feed corrected production examples into a governed active-learning process.
FAQ
How many examples are needed per intent?
There is no universal number. Begin with 30–50 diverse human-authored examples for common intents, fewer only when the intent is exceptionally clear, and more for confusing or multilingual categories. Use held-out tests rather than training volume as the quality signal.
Should I use an LLM or a BERT-style classifier?
Use a compact classifier when intents are stable and labelled data is available. Use an LLM for zero-shot exploration, difficult ambiguity, or adjudication—but constrain its labels, validate its outputs, and account for latency, cost, and data-governance requirements.
How should Hinglish be supported?
Treat it as a core product language, not an edge case. Collect naturally occurring code-mixed examples, test both native and Roman scripts, use multilingual embeddings or models, and report Hinglish performance separately.
What is the most important improvement to make first?
Audit the taxonomy and error logs before changing models. In many systems, overlapping labels, missing OOS handling, and weak examples cause more errors than model capacity.