0tokens

Apply for AI Grants India

Financial support for innovators building the future of AI in India.

Apply now

Chat · how to fine tune a small language model for telugu customer support

How to Fine-Tune a Small Language Model for Telugu Support

  1. aigi

    What you are building

    A Telugu support model should do more than produce grammatically correct sentences. It must identify the customer’s intent, preserve product and policy details, handle Telugu script and transliterated Telugu, and know when to hand a conversation to a human agent. A small model can be a strong choice when you need lower inference cost, predictable latency, data residency, or deployment on modest infrastructure.

    Fine-tuning is not a substitute for a knowledge base. Use retrieval or API calls for changing information such as order status, prices, stock, delivery estimates, and account records. Fine-tuning should teach the model how to respond, your support tone, intent patterns, escalation rules, and recurring Telugu phrasing.

    For broader language and data constraints, start with this guide to low-resource Indic natural language processing. It covers many of the issues that also affect Telugu systems: limited labelled data, spelling variation, code-mixing, and evaluation gaps.

    Define the support scope first

    Start with a narrow set of high-volume, low-risk intents rather than attempting to automate every ticket. Suitable first use cases include:

    • Delivery tracking and estimated arrival questions
    • Refund, return, cancellation, and replacement policies
    • Password resets and basic account guidance
    • Product FAQs and warranty information
    • Store timings, service availability, and branch details
    • Ticket creation, status checks, and agent hand-off

    Separate informational, transactional, and sensitive requests. A model may answer an FAQ, but it should call a verified backend before making an account-specific claim. It should never invent refunds, approve exceptions, expose personal data, or provide unsupported financial, medical, or legal advice.

    Write an intent catalogue before collecting data. For each intent, define sample user messages, required fields, the approved answer, backend action, confidence threshold, and escalation condition.

    Build a Telugu-first dataset

    Your most valuable training data is usually existing support traffic, not generic web text. Combine agent conversations, FAQ articles, call transcripts, chat logs, CRM tags, and carefully written examples. Obtain consent and remove phone numbers, addresses, order IDs, email addresses, and other personally identifiable information before annotation.

    Include the language patterns customers actually use:

    • Telugu script, such as “నా ఆర్డర్ ఎక్కడ ఉంది?”
    • Romanised Telugu, such as “naa order ekkada undi?”
    • Telugu-English code-mixing, product names, abbreviations, and regional spellings
    • Short, incomplete, emotional, or repeated messages
    • Dialect and politeness variation across Andhra Pradesh and Telangana
    • Speech-to-text errors if voice support is planned

    Create separate training, validation, and test sets by conversation or customer, not by randomly splitting individual messages. Otherwise, near-duplicate queries can leak into evaluation and make the model appear more capable than it is. Keep a challenging test set containing transliteration, noisy spelling, long context, ambiguous requests, and adversarial prompts.

    A useful supervised example contains the customer message, relevant conversation history, intent, structured fields, and the ideal response. Include refusal and escalation examples, not just successful answers. If you are adapting an existing base model, the best practices for fine-tuning LLMs on custom data provide a useful framework for dataset quality, split design, and training controls.

    Choose the smallest suitable base model

    Prioritise a model with demonstrated Telugu or multilingual Indic capability, a permissive commercial licence, a tokenizer that handles Telugu efficiently, and a context window large enough for your support workflow. Benchmark several candidates on your own test set; model size alone does not predict Telugu quality.

    For most teams, begin with parameter-efficient fine-tuning rather than updating every parameter. LoRA or QLoRA can reduce GPU memory requirements and make experiments easier to reproduce. A practical workflow is:

    1. Format conversations using the base model’s chat template.
    2. Start with a modest learning rate and a small number of epochs.
    3. Train on high-quality examples with a held-out validation set.
    4. Save checkpoints and compare them on the same Telugu evaluation suite.
    5. Quantise only after quality is acceptable, then measure the effect on accuracy and latency.

    Do not train on raw transcripts without filtering. Duplicates, contradictory policies, agent mistakes, and outdated offers teach the model behaviour you will later have to remove. Maintain a versioned dataset and record the base model, adapter, tokenizer, hyperparameters, and licence for every release.

    Evaluate behaviour, not just fluency

    BLEU and ROUGE are weak indicators for support because several answers can be correct. Use a combination of automated checks and human review by Telugu-speaking evaluators. Track:

    • Intent classification accuracy and confusion between similar intents
    • Required-entity extraction, such as order number or product name
    • Grounded-answer accuracy against approved policy content
    • Hallucination and unsupported promise rate
    • Correct escalation and refusal rate
    • Telugu script quality, transliteration handling, and code-mixed comprehension
    • Latency, token usage, failure rate, and cost per resolved conversation

    Create a rubric that scores factual correctness, completeness, tone, clarity, privacy, and actionability. Test both Telugu and English prompts, because users may switch languages mid-conversation. Include red-team cases involving prompt injection, requests for another customer’s information, policy exceptions, abusive language, and attempts to make the model claim that an action was completed.

    A strong production rule is fail closed: if confidence is low or a backend is unavailable, the assistant should say what it can verify, ask one focused question, or transfer the case. It should not guess.

    Connect the model to support systems

    Use retrieval-augmented generation for policy and product content, with document versioning and source citations available to the orchestration layer. Use tools or APIs for live actions such as checking an order, raising a ticket, or initiating a return. Keep these permissions narrow and require confirmation for irreversible actions.

    A deployment architecture can include a language detector, normalisation layer, intent router, Telugu support model, retrieval service, business tools, safety filter, and human hand-off queue. For voice channels, add speech recognition and text-to-speech evaluation; latency, pronunciation of names, numbers, and English product terms all matter. A voice agent platform for small businesses can help with the surrounding call workflow, but test Telugu recognition independently rather than assuming Hindi or English performance transfers.

    For WhatsApp or mobile deployments, optimise memory and response time after measuring real traffic. Quantisation, batching, caching, and shorter prompts can lower cost. If the model must run on-device or at the edge, review this guide to AI model optimisation for mobile devices.

    Monitor and improve safely

    Launch with a limited intent set and a human review queue. Log anonymised inputs, model outputs, retrieved sources, tool calls, confidence signals, escalations, and user feedback. Sample conversations weekly and label failure categories: misunderstanding, wrong policy, missing context, unsafe action, poor Telugu, or unnecessary escalation.

    Retrain only from reviewed examples. A monthly or quarterly refresh may be appropriate, but frequency should follow drift in products, policies, and language—not a fixed calendar. Maintain rollback capability for both the model and knowledge base. Publish a clear disclosure that users are interacting with an automated assistant and make human support easy to reach.

    A practical launch checklist

    Before production, confirm that you have:

    • A defined Telugu intent and escalation taxonomy
    • De-identified, consented, versioned data
    • Separate customer-level train, validation, and test splits
    • Baseline comparisons against a prompt-only model and human support
    • Telugu, transliteration, code-mixing, and adversarial test cases
    • Retrieval for changing facts and tools for account actions
    • Privacy, access control, audit logs, and retention policies
    • Human hand-off with full conversation context
    • Cost, latency, hallucination, and resolution-rate dashboards
    • A rollback plan and an owner for ongoing evaluation

    The goal is not to make a small model answer every question. It is to make common Telugu interactions accurate, respectful, inexpensive, and easy to escalate when automation is uncertain. That focused approach usually delivers more reliable customer outcomes than fine-tuning a broad model on a large but inconsistent transcript archive.

    Last updated 23 September 2026

AIGI may be inaccurate. Replies seeded from the guide above.