What you are building
A Malayalam WhatsApp chatbot is not simply a language model connected to a messaging number. A reliable system combines WhatsApp’s Cloud API, a webhook service, Malayalam-aware text handling, business rules, a knowledge base, and a small language model (SLM) that generates or classifies responses.
The practical goal is a bot that can answer a defined set of questions, complete a few useful workflows, and hand off uncertain or sensitive conversations to a person. Start with one use case—such as appointment booking, order status, government-service guidance, or customer support—instead of attempting a general Malayalam assistant.
For broader design decisions, read this guide alongside low-resource Indic NLP techniques and the principles for building AI apps for India’s next billion users (link unavailable—use the listed topic URL below).
Why use a small language model?
An SLM can be cheaper and faster than a large hosted model, especially when the task is narrow. It may run on a modest GPU or, after quantisation, on CPU infrastructure. That matters when a startup needs predictable costs and must keep customer data within a controlled environment.
However, Malayalam quality depends on the model, prompt design, spelling variation, and the data used for evaluation. A smaller model should not be expected to know changing product information. Give it verified context through retrieval and restrict it with business rules.
Choose an SLM based on:
- Malayalam capability: Test native Malayalam, Manglish, code-mixed Malayalam-English, numbers, dates, names, and local place names.
- Context length: Ensure it can process your FAQ snippets and conversation history without excessive truncation.
- Deployment licence: Check commercial-use terms, redistribution requirements, and restrictions on fine-tuning.
- Latency and hardware: Benchmark quantised versions using your expected concurrent conversations.
- Safety behaviour: Test refusal, escalation, and prompt-injection resistance before launch.
Recommended architecture
A practical production flow looks like this:
1. A user sends a message to the WhatsApp business number.
2. Meta sends an HTTPS webhook event to your backend.
3. The backend verifies the signature, identifies the user, and records the message ID for deduplication.
4. A language router detects Malayalam, Manglish, English, or an unsupported language.
5. An intent classifier checks whether the request matches a known workflow.
6. The retrieval layer fetches approved Malayalam or bilingual content.
7. The SLM drafts a short answer using only the retrieved context.
8. A policy layer checks the answer, applies templates, and decides whether a human handoff is required.
9. The backend sends a WhatsApp reply and stores delivery status and evaluation signals.
Keep the model behind an application layer. The application, not the model, should control authentication, payments, account changes, refunds, appointment slots, and other consequential actions.
Step-by-step build plan
1. Define the first release
List the top 20–50 user intents from real support logs or interviews. For each intent, document:
- Example Malayalam, Manglish, and English utterances
- The correct answer or action
- Required account information
- Whether authentication is needed
- When the bot must escalate
Set a clear out-of-scope response. For example: “ഇത് പരിശോധിക്കാൻ ഞങ്ങളുടെ ജീവനക്കാരുമായി ബന്ധിപ്പിക്കാം” (“I can connect you to our team to check this”). Do not make users repeat their problem after handoff; pass the conversation summary and relevant metadata to the agent.
2. Set up WhatsApp Cloud API
Create a Meta developer app, add WhatsApp, configure a business phone number, and generate a production access token through your chosen business setup. Configure a public HTTPS webhook and subscribe to message and status events. Store tokens in a secrets manager, not in source code.
Your webhook should acknowledge valid events quickly—ideally within a few seconds—and process model inference asynchronously. Implement retries, dead-letter handling, rate limiting, and idempotency using the WhatsApp message ID. Follow the current WhatsApp Business Platform rules for opt-in, templates, conversation windows, and user-initiated versus business-initiated messages.
3. Build Malayalam text handling
Normalise Unicode without destroying meaningful characters. Preserve the original message for audit and store a normalised copy for search. Your pipeline should handle:
- Malayalam script and common spelling variants
- Manglish transliteration, such as “ente order evide?”
- Malayalam-English code switching
- Emojis, punctuation, phone numbers, dates, and currency amounts
- Voice notes, if supported, through a separate speech-to-text pipeline
Do not silently translate every message into English. Translation can lose names, quantities, politeness, and domain terms. Test direct Malayalam retrieval and, where useful, bilingual indexing.
4. Add retrieval before fine-tuning
Create a small, versioned knowledge base from approved FAQs, product documents, pricing rules, service areas, and escalation policies. Split documents into focused passages, attach metadata such as product and effective date, and retrieve the top relevant passages for each request.
Use citations internally so reviewers can see which source supported an answer. Add expiry dates to time-sensitive content. If retrieval returns no sufficiently relevant result, the bot should ask a clarifying question or escalate rather than invent an answer.
Fine-tuning can improve tone, intent classification, or formatting, but it should not be your first method for keeping facts current. For an overview of agent design and orchestration, see this practical guide to generative AI agents.
5. Implement controlled responses
Use structured outputs from the model, such as JSON containing intent, answer, confidence, source_ids, and handoff_required. Validate the schema before sending anything to WhatsApp. Apply message-length limits and split long answers into readable sections.
Prefer buttons, lists, and quick replies for predictable workflows. Keep Malayalam responses concise, use familiar terms, and offer English only when the user requests it or when a technical term is clearer in English. For regulated advice, medical topics, financial decisions, and identity-related requests, use approved templates and human review.
Testing and evaluation
Create a test set with at least several hundred examples across Malayalam script, Manglish, spelling errors, code-mixing, slang, and ambiguous requests. Measure:
- Intent accuracy and retrieval recall
- Factual accuracy against approved sources
- Correct language and script selection
- Safe refusal and escalation rate
- Median and p95 response latency
- Cost per resolved conversation
- Handoff rate and user re-contact rate
Have native Malayalam speakers review outputs. Automated scores alone will miss unnatural phrasing, incorrect honorifics, and subtle meaning changes. Run adversarial tests for prompt injection, personal-data extraction, policy bypasses, and attempts to trigger unauthorised actions.
Production operations and costs
Deploy the webhook and orchestration service separately from model inference so each can scale independently. Use a queue for bursts, Redis or a database for short-lived conversation state, encrypted storage for transcripts, and dashboards for latency, errors, token usage, and escalation.
Estimate costs across Meta messaging charges, hosting, model inference, storage, observability, speech services, and human support. A compact model may reduce inference cost, but poor answers increase support costs. Optimise for cost per successfully resolved conversation, not cost per generated message.
Use staged rollout: internal testing, a small opt-in cohort, then gradual expansion. Keep a rollback model and a content-version rollback. Review transcripts under a documented retention policy and remove unnecessary personal data.
Common failure modes
- Treating Malayalam as translated English: Build and evaluate with native Malayalam examples.
- Using the model as a database: Retrieve current facts and require source-backed answers.
- Ignoring Manglish: Include transliteration in intent and search tests.
- No human fallback: Escalate on low confidence, repeated misunderstanding, and sensitive requests.
- Slow synchronous webhooks: Acknowledge events quickly and process asynchronously.
- Uncontrolled actions: Require deterministic validation and authentication before account changes.
If the use case is primarily spoken rather than text-based, compare this approach with a voice agent versus chatbot and review voice-agent architecture and deployment choices.
Launch checklist
Before going live, confirm that you have:
- A documented opt-in and privacy process
- Verified webhook signatures and secret management
- Idempotent message processing and retry handling
- Malayalam, Manglish, English, and mixed-language test coverage
- Retrieval with dated, approved sources
- Human handoff with conversation context
- Monitoring for quality, latency, cost, and safety
- A rollback plan for the model, prompt, and knowledge base
A focused Malayalam chatbot with strong retrieval, explicit guardrails, and a dependable human fallback will usually outperform a more ambitious general-purpose bot. Improve it from real conversations, but make every model, prompt, and content change measurable and reversible.