0tokens

Apply for AI Grants India

Financial support for innovators building the future of AI in India.

Apply now

Chat · how to build whatsapp chatbot with small language model in hindi

How to Build a WhatsApp Chatbot with a Small Hindi Language Model

  1. aigi

    A Hindi WhatsApp bot should do more than generate fluent replies. It must handle Devanagari, Hinglish, spelling variation, intermittent connectivity, short mobile messages, and clear handoffs to a human agent. A small language model (SLM) is often the right starting point because it can reduce latency and inference cost while keeping sensitive business data under tighter control.

    This guide explains how to build a production-minded chatbot for customer support, appointments, order updates, or internal service desks. The design also applies to other Indian languages; the main challenge is usually data quality and evaluation, not model size alone. For broader context on Indic data and tokenisation, see this practical guide to low-resource Indic natural language processing.

    Define the job before choosing a model

    Start with a narrow, measurable use case. A bot that answers ten common questions reliably is more valuable than one that attempts unrestricted conversation.

    Write down:

    • The user groups: customers, field staff, students, patients, or merchants.
    • Supported languages: Hindi, Hinglish, English, and any regional fallback.
    • The top 20-50 intents, such as pricing, order status, returns, or appointment booking.
    • Which actions require authentication or human approval.
    • The target response time, escalation rate, and acceptable failure rate.

    Create representative test examples in Devanagari and Roman Hindi. For example, include “मेरा ऑर्डर कहाँ है?”, “mera order kaha hai”, abbreviations, typos, voice-transcribed text, and messages containing English product names. This dataset becomes your evaluation set; do not use all of it for fine-tuning.

    Choose a small model and a grounded architecture

    For most business bots, fine-tuning a model is not the first step. Begin with a multilingual or Hindi-capable instruction model that fits your latency and hosting budget. Benchmark candidate models on your own messages for intent classification, answer accuracy, refusal behaviour, and Hinglish handling.

    A robust architecture separates four functions:

    1. WhatsApp transport receives and sends messages.
    2. Orchestration identifies intent, applies authentication and business rules, and decides whether to answer, retrieve information, call an API, or escalate.
    3. Knowledge and tools provide current information from approved documents and systems.
    4. The language model drafts a concise response within strict limits.

    Use retrieval-augmented generation (RAG) for changing information such as policies, catalogues, and FAQs. Store cleaned Hindi and bilingual content with metadata, retrieve only relevant passages, and instruct the model to say when the answer is unavailable. Do not expect a model’s training memory to know your latest inventory or account data.

    If your workflow needs multiple specialised steps, compare the design with approaches for building distributed systems with AI agents, but avoid adding agents where a deterministic function will work.

    Set up the WhatsApp integration

    Use the official WhatsApp Business Platform through Meta’s Cloud API or an authorised provider. Your basic flow is:

    • Create or connect a Meta Business account and WhatsApp Business Account.
    • Register a phone number and configure credentials securely.
    • Expose a public HTTPS webhook for verification and inbound events.
    • Validate webhook signatures and reject malformed or replayed requests.
    • Store the sender ID, message ID, timestamp, language signal, and consent state.
    • Send replies through the Graph API, respecting conversation windows and approved message templates.

    Keep the webhook fast. Acknowledge the event, place the message on a queue, and process it asynchronously. This prevents model latency from causing retries and duplicate replies. Make message handling idempotent by recording event IDs before processing.

    Never place access tokens, customer records, or model prompts in source code or logs. Use environment secrets, encryption at rest, role-based access, and separate development and production numbers.

    Build Hindi and Hinglish understanding

    Normalise input carefully without destroying meaning. Preserve the original message for auditing, then create a processing version that handles Unicode normalisation, repeated characters, whitespace, emojis, and common transliteration variants. Detect language at the message or conversation level, not only from a single word: Hindi users often switch scripts within one sentence.

    A practical response policy is:

    • Reply in the user’s apparent script and language.
    • Ask once whether the user prefers Hindi or English when confidence is low.
    • Use familiar terms, short sentences, and Indian number/date formats.
    • Avoid inventing Hindi translations for product, legal, or technical names.
    • Offer numbered options for mobile users instead of long paragraphs.

    Do not fine-tune on scraped conversations without permission and redaction. Remove phone numbers, addresses, order IDs, Aadhaar details, payment information, and names where they are not required. Start with curated examples covering successful answers, ambiguous requests, unsafe requests, and escalation cases. Fine-tuning should improve style or intent behaviour; it should not be used to memorise private customer data.

    Add tools, authentication, and guardrails

    Use deterministic tools for actions such as checking an order, creating a ticket, calculating a bill, or booking a slot. The model may select a tool, but your backend must validate every parameter and permission before execution.

    Important controls include:

    • Confirm identity before exposing account or order information.
    • Require confirmation before irreversible actions, refunds, or payments.
    • Limit tool access by intent and user role.
    • Set timeouts, retries, rate limits, and spending limits.
    • Detect prompt injection in user-provided documents and messages.
    • Refuse medical, legal, financial, or safety-critical advice beyond the bot’s approved scope.
    • Provide a visible “talk to an agent” route with business hours and expected wait time.

    For regulated use cases, the privacy boundary matters as much as model quality. A private deployment pattern may be useful; compare the trade-offs in how to build a private AI chatbot for lawyers, even if your own domain is different.

    Evaluate before production

    Build an evaluation set of at least a few hundred realistic Hindi, Hinglish, and English messages. Measure:

    • Intent accuracy and fallback accuracy.
    • Retrieval recall and citation or source correctness.
    • Tool-call success and permission failures.
    • Unsupported-claim rate and unsafe-response rate.
    • Hindi script, grammar, and transliteration quality.
    • Median and p95 response latency.
    • Human handoff rate, resolution rate, and cost per conversation.

    Test adversarially: misspellings, code-switching, forwarded messages, prompt injection, duplicate webhooks, empty media captions, abusive content, and unavailable backend services. Use a human review queue for low-confidence conversations. Re-run the same benchmark after every prompt, model, retrieval, or policy change.

    Deploy and operate the bot

    A lean production stack can use a Python or Node.js webhook service, a managed queue, PostgreSQL for state and audit records, a vector store for approved knowledge, and a small model served locally or through a controlled inference endpoint. Containerise the service and deploy it in an Indian region when latency, residency, or contractual requirements justify that choice.

    Monitor both technical and user outcomes:

    • Webhook failures, queue depth, API errors, and model latency.
    • Token or inference cost by intent.
    • Fallbacks, escalations, repeated questions, and failed actions.
    • Language distribution and performance by script.
    • Sensitive-data exposure and policy violations.

    Roll out gradually to staff or a small customer cohort. Keep a kill switch that disables model replies while preserving human support. Review conversations with appropriate access controls, delete data according to a documented retention policy, and publish a simple privacy notice.

    A practical build sequence

    For a first release, avoid training a model from scratch. Build in this order:

    1. Launch five to ten high-value FAQ intents with deterministic fallback.
    2. Add WhatsApp templates, webhook reliability, state, and human escalation.
    3. Connect one read-only business tool, such as order-status lookup.
    4. Add bilingual retrieval and a curated Hindi evaluation set.
    5. Introduce fine-tuning only after prompts, retrieval, and rules have been measured.
    6. Expand intents based on unresolved conversations, not assumptions.

    This approach keeps the system affordable and debuggable while leaving room to support more languages and channels. If the project is part of a wider multilingual product, the principles in building AI apps for the next billion users in India are useful for designing around shared devices, low bandwidth, and varied digital literacy.

    FAQ

    Should I fine-tune a model immediately?
    Usually no. Start with prompting, retrieval, intent routing, and business rules. Fine-tune only when you have a clean dataset and a measured failure pattern that those methods cannot fix.

    Can the bot understand Hinglish?
    Yes, if Hinglish examples are included in evaluation and prompts. Benchmark Roman Hindi, Devanagari, spelling variation, and code-switching separately.

    Can I run the model on my own server?
    Often yes, depending on model licence, hardware, traffic, and latency requirements. Verify commercial usage terms and benchmark real workloads before committing.

    What is the most common production failure?
    Poorly defined scope and missing fallback paths. A concise refusal and quick human handoff are better than a confident, fabricated answer.

    How much does WhatsApp access cost?
    Costs depend on Meta’s current conversation and template pricing, provider fees, infrastructure, and model usage. Check the applicable commercial terms before launch rather than relying on old estimates.

    Last updated 23 September 2026

AIGI may be inaccurate. Replies seeded from the guide above.