0tokens

Apply for AI Grants India

Financial support for innovators building the future of AI in India.

Apply now

Chat · how to train a small language model

How to Train a Small Language Model: A Practical 2026 Guide

  1. aigi

    Small language models (SLMs) are useful when a product needs low latency, predictable cost, data control, or on-device inference. For most teams, however, “training a model” does not mean building a frontier model from random initialisation. It usually means choosing a compact open model, adapting it with domain data, and measuring whether it improves a clearly defined task.

    That distinction matters in India, where multilingual text, code-mixed queries, noisy user input, and limited GPU budgets can make a smaller, specialised model more practical than a large general-purpose API. This guide explains how to train a small language model in a way that is technically sound and commercially useful.

    1. Choose the right training strategy

    Start by deciding what “train” means for your project:

    • Prompting or retrieval-augmented generation: Best when facts change frequently and you mainly need better context handling.
    • Supervised fine-tuning (SFT): Best for teaching response formats, classification labels, instruction following, or domain-specific behaviour.
    • Continued pretraining: Best when you have a large, clean corpus in a particular domain or language and want to improve vocabulary and fluency.
    • Training from scratch: Suitable only when you have substantial data, a distinctive language or domain requirement, and the budget to run many experiments.

    For a first production system, fine-tuning a compact decoder model is usually the sensible starting point. If your work involves Hindi or other Indian languages, review guidance on open-source small language models for Hindi before selecting a base checkpoint. For broader multilingual work, low-resource Indic natural language processing offers useful considerations around scripts, transliteration, and dataset scarcity.

    2. Define the task and success criteria

    Write a one-page training brief before collecting data. Include:

    • Primary task: generation, classification, extraction, translation, summarisation, or conversational assistance.
    • Users and language mix: for example, Hindi-English code-mixed support queries from Indian customers.
    • Latency and hardware target: CPU, consumer GPU, cloud GPU, or mobile device.
    • Quality threshold: task accuracy, macro-F1, exact match, perplexity, groundedness, or human preference.
    • Safety requirements: personally identifiable information, financial advice, health claims, toxicity, and refusal behaviour.

    Avoid using a single score as your definition of success. A model can improve benchmark accuracy while becoming less reliable on real customer queries. Build a small “golden set” of representative examples, including difficult spellings, regional terms, code switching, and adversarial prompts.

    3. Build and govern the dataset

    Data quality generally matters more than adding another layer or increasing the parameter count. Useful sources include licensed internal records, public datasets with compatible licences, synthetic examples reviewed by humans, and carefully filtered web data.

    A practical preparation pipeline should:

    • Remove duplicates and near-duplicates.
    • Strip HTML, boilerplate, spam, and broken encoding.
    • Detect language and script rather than assuming the label from a source website.
    • Redact phone numbers, email addresses, Aadhaar details, account numbers, and other personal data.
    • Separate train, validation, and test records by user, document, or time period to prevent leakage.
    • Preserve meaningful punctuation, numerals, and code-mixed terms instead of normalising everything into English.
    • Record source, licence, processing steps, and quality checks in a dataset manifest.

    For instruction fine-tuning, use consistent records such as instruction, context, and response, or the chat format required by your base model. Keep answers accurate, concise, and varied. Repeated templates can make training loss look healthy while producing brittle behaviour.

    4. Select the model, tokenizer, and compute plan

    Choose a model based on licence, language coverage, context length, architecture, and deployment constraints—not parameter count alone. A 1B–3B parameter model may be adequate for extraction or support triage, while generation-heavy tasks may need a larger compact model or retrieval support.

    Check the tokenizer on your actual data. Poor tokenisation of Devanagari, Bengali, Tamil, or transliterated Hindi increases sequence length and compute cost. If the base model handles your target language badly, fine-tuning may not fully fix the problem. For serious regional-language work, compare token counts and sample generations before committing to a checkpoint. Fine-tuning Llama for Indian regional languages is a useful reference for this decision.

    Estimate compute before starting. A pilot can often run on a rented data-centre GPU or a capable local GPU using mixed precision and parameter-efficient methods. Use gradient accumulation when memory limits batch size, checkpoint regularly, and budget for evaluation runs—not only training runs.

    5. Fine-tune efficiently

    For most small teams, LoRA or QLoRA is the default approach. These methods update a small number of adapter parameters while keeping the base model largely frozen, reducing memory use and making experiments easier to compare.

    A typical workflow is:

    1. Load the base model and tokenizer.
    2. Format and tokenise the dataset using the model’s chat template where applicable.
    3. Apply LoRA or QLoRA adapters.
    4. Set a conservative learning rate, sequence length, batch strategy, and epoch count.
    5. Train while logging loss, learning rate, throughput, and validation results.
    6. Save checkpoints and compare them on the golden set.
    7. Merge adapters only when the deployment stack requires a standalone model.

    Do not train for more epochs simply because the loss continues to fall. Overfitting often appears as memorised phrasing, worse generalisation, or confident answers outside the training distribution. Start with a small pilot, inspect outputs manually, and change one variable at a time.

    Continued pretraining requires a different dataset and objective: large volumes of raw text rather than question-answer pairs. It can improve domain language, but it also risks degrading general instruction following. Keep a general-purpose regression set to detect that trade-off.

    6. Evaluate beyond loss and perplexity

    Use evaluation layers that match the product:

    • Task metrics: accuracy, macro-F1, precision, recall, ROUGE, BLEU, or exact match where appropriate.
    • Generation checks: factuality, completeness, instruction adherence, language quality, and citation or retrieval accuracy.
    • Robustness tests: spelling errors, long inputs, code mixing, unseen names, numeric values, and ambiguous requests.
    • Safety tests: personal-data leakage, unsafe advice, prompt injection, and inappropriate refusals.
    • Operational tests: tokens per second, memory use, cold-start time, and cost per request.

    Maintain separate evaluation sets for Hindi, English, and code-mixed inputs if your product serves all three. Human review remains important for regional-language quality because automated metrics can miss unnatural phrasing or culturally incorrect responses.

    7. Deploy and monitor the model

    Export the model in a format supported by your target runtime, then test quantised variants such as 8-bit or 4-bit inference. Quantisation can substantially reduce memory and latency, but validate quality on your golden set before shipping. If the endpoint must run on a phone, edge gateway, or low-cost server, follow the optimisation principles in AI model optimisation for mobile devices.

    Expose the model through a versioned API with authentication, rate limits, structured logs, and clear timeout behaviour. Store prompts and outputs only when your privacy policy permits it, and redact sensitive fields before logging. Track:

    • latency and error rates;
    • token usage and infrastructure cost;
    • language and task distribution;
    • abstention, escalation, or fallback rates;
    • user corrections and harmful outputs.

    A smaller model should not be forced to answer every question. Route complex requests to retrieval, a larger model, or a human workflow when confidence or evaluation rules indicate that it is out of scope.

    8. Common mistakes to avoid

    • Training from scratch before testing a strong open checkpoint.
    • Mixing licences or using scraped personal data without a defensible legal basis.
    • Splitting near-identical documents across train and test sets.
    • Optimising benchmark scores while ignoring real user inputs.
    • Fine-tuning on synthetic answers without human quality checks.
    • Changing the tokenizer casually after training begins.
    • Deploying quantisation without regression testing.
    • Treating model weights as the entire product; retrieval, monitoring, and fallback logic often matter more.

    A practical starting plan

    In week one, define the task, create a 200–1,000-example evaluation set, and compare two or three compact models. In week two, clean and document the training data, run a LoRA pilot, and inspect errors by language and use case. In week three, test quantisation, latency, safety, and failure routing. Ship only after the model beats your baseline on both quality and operating cost.

    For builders working with visual or multimodal inputs, a language model may be only one part of the system; related approaches are covered in open-source vision-language models for Indian languages. The same principle applies throughout: start with a narrow, measurable job, use the smallest model that meets it, and keep improving the data and evaluation loop.

    Last updated 23 September 2026

AIGI may be inaccurate. Replies seeded from the guide above.