0tokens

Apply for AI Grants India

Financial support for innovators building the future of AI in India.

Apply now

Chat · how to fine tune a small language model for marathi customer support

How to Fine-Tune a Small Language Model for Marathi Support

  1. aigi

    Marathi customer support needs more than a generic multilingual chatbot. Customers switch between Marathi, Hindi, and English; use Devanagari, Latin transliteration, or both; and expect answers that reflect local policies, products, and payment practices. A small language model can handle this workload efficiently—but only when its data, training objective, and operating controls are designed for the job.

    This guide explains how to fine tune a small language model for Marathi customer support, with an emphasis on practical deployment for Indian businesses.

    Start with the right problem

    Fine-tuning is not a substitute for a knowledge base. It teaches a model how to respond, classify, or follow a support style; it does not reliably update changing prices, return rules, balances, or inventory. Use retrieval or API calls for facts that change frequently, and fine-tune for:

    • Marathi and mixed-language intent recognition
    • Consistent tone, formatting, and escalation behaviour
    • Product-specific conversation patterns
    • Structured outputs such as intent, language, priority, and next action
    • Common support workflows, including refunds, delivery updates, and account access

    A sensible architecture combines a compact model with retrieval, business-system integrations, and a human handoff. Teams building for several Indian languages should also study the constraints described in this low-resource Indic NLP builder’s guide.

    Define Marathi support requirements

    Before collecting data, document the conversations the model must handle. Create an intent taxonomy such as order status, payment failure, cancellation, replacement, warranty, KYC, complaint, and agent escalation. For each intent, specify:

    • Approved answer or action
    • Information the model must request
    • Claims it must never make
    • Conditions requiring an agent
    • Whether the response should be Marathi, English, Hindi, or code-mixed

    Include regional and conversational variation. Marathi customers may write “माझं ऑर्डर कुठे आहे?”, “माझी order अजून आली नाही”, or “order track kasa karu?” These are related requests but differ in script, spelling, and language mixing. Preserve such variation in training and evaluation data rather than normalising everything into formal Marathi.

    Build a high-quality dataset

    Use consented, redacted sources: resolved tickets, approved chat transcripts, FAQs, call summaries, and agent-written responses. Remove phone numbers, addresses, account identifiers, payment details, internal notes, and any unnecessary personal information. Do not train on unresolved or contradictory answers without labelling them.

    A useful supervised example contains:

    • Customer message
    • Relevant conversation context
    • Intent and escalation label
    • Grounded answer or tool action
    • Preferred language and script

    Create separate training, validation, and test sets by conversation, not by randomly splitting individual messages. Otherwise, nearly identical tickets can appear in both training and testing, producing misleading results. Keep a deliberately difficult test set containing transliteration, typos, dialect variation, long context, angry customers, and unsupported requests.

    For a small initial project, a few thousand carefully reviewed examples can be more useful than a large noisy export. Balance frequent intents with high-risk cases such as payment disputes, account recovery, and privacy requests.

    Choose the model and fine-tuning method

    Select a compact instruction-tuned model with demonstrated Devanagari and multilingual ability. Test candidate models on your own Marathi examples; model size and benchmark scores alone do not predict support quality. Check context length, commercial licence, quantisation options, inference speed, and whether the model can run within your infrastructure and data-residency requirements.

    For most teams, parameter-efficient fine-tuning is the practical starting point. LoRA or QLoRA updates a small set of adapter parameters, reducing GPU memory, training time, and storage. Full fine-tuning may be justified only when you have substantial clean data, a stable objective, and the engineering capacity to prevent catastrophic forgetting. Follow the principles in this guide to fine-tuning LLMs on custom data.

    For classification, consider a smaller encoder model and train it on intent labels rather than forcing a generative model to perform every task. A two-stage system—intent and safety classification first, response generation second—can be easier to evaluate and control.

    Prepare Marathi data for training

    Avoid aggressive text cleaning. Marathi punctuation, emoji, code-switching, honorifics, and informal spellings carry useful signals. Standardise only what is necessary, and retain the original message alongside any normalised form.

    Recommended preparation steps include:

    • Detect language and script, including Marathi written in Latin characters.
    • Deduplicate near-identical conversations.
    • Correct only verified annotation or transcription errors.
    • Mark sensitive entities and replace them with typed placeholders.
    • Keep response templates separate from dynamic fields.
    • Add negative examples where the correct action is to ask a question or escalate.

    Use a consistent chat format and explicitly teach the model when it should call a tool, cite retrieved information, or decline. Do not reward confident guesses. A short, accurate Marathi answer with a clear handoff is better than an invented resolution.

    Train with controlled experiments

    Start with conservative settings and establish a baseline before changing multiple variables. Track the base model, dataset version, adapter configuration, random seed, training steps, and evaluation results. Monitor training and validation loss, but do not select a model on loss alone.

    Useful experiments include:

    • Base model versus fine-tuned model
    • Marathi-only data versus balanced mixed-language data
    • Different LoRA ranks and learning rates
    • Response generation versus intent-plus-response architecture
    • Retrieval enabled versus disabled

    Use early stopping when validation quality stops improving. Watch for memorisation, repetitive answers, language drift, and a drop in performance on general safety or English inputs. Quantise only after confirming that quality remains acceptable; deployment optimisation should follow evaluation, not replace it. For edge or low-cost inference, review current approaches to optimising AI models for mobile devices.

    Evaluate what matters in production

    BLEU or perplexity cannot tell you whether a customer received a correct refund instruction. Build a Marathi evaluation set reviewed by native speakers and support specialists. Score at least:

    • Intent accuracy and escalation recall
    • Factual and policy adherence
    • Marathi fluency, spelling, and script handling
    • Performance on transliterated and code-mixed messages
    • Helpfulness, concision, and tone
    • Refusal quality and resistance to prompt injection
    • Latency, token use, and failure rate

    Measure high-risk errors separately. A model that answers 95% of routine questions well may still be unsuitable if it gives unsafe advice on payments or identity verification. Run red-team tests for fabricated order status, requests for OTPs or passwords, abusive language, prompt leakage, and attempts to override support policy.

    Deploy with retrieval and human fallback

    Put the model behind a service that authenticates users, filters sensitive inputs, retrieves current policy content, validates tool calls, and logs outcomes without storing unnecessary personal data. Constrain actions with allowlists and require confirmation for refunds, cancellations, or account changes.

    Start in suggestion mode, where agents approve AI drafts. Move to limited automation only for low-risk, high-volume intents. Display the source policy or retrieved record to agents, offer one-click correction labels, and route low-confidence or frustrated conversations to humans. Marathi voice channels may also benefit from a specialised voice agent for small businesses, but speech recognition and transliteration require separate testing.

    Operate and improve the system

    Monitor language mix, intent distribution, fallback rate, resolution rate, correction rate, customer satisfaction, and harmful-error incidents. Sample conversations regularly across districts, scripts, and customer segments. Create a versioned feedback pipeline so reviewed conversations become future training or evaluation examples only after privacy checks and approval.

    Retain a rollback path for every adapter and prompt version. Re-test after policy changes, model updates, retrieval-index changes, and quantisation. For Indian deployments, document data flows, access controls, retention periods, vendor terms, and consent practices; compliance is part of the product design, not a final checklist.

    Practical launch checklist

    • Define supported intents, languages, scripts, and escalation rules.
    • Collect and redact representative Marathi conversations.
    • Build a conversation-level test set with difficult edge cases.
    • Benchmark at least one base model and one parameter-efficient adapter.
    • Ground changing information through retrieval or verified APIs.
    • Evaluate native-language quality and high-risk errors separately.
    • Launch with agent review, monitoring, and a rapid rollback process.
    • Retrain only from reviewed, consented, versioned data.

    The strongest Marathi support systems are not the ones with the largest model. They are the ones that combine representative language data, narrow business scope, reliable grounding, measurable safeguards, and a clear path to a human agent.

    Last updated 23 September 2026

AIGI may be inaccurate. Replies seeded from the guide above.