0tokens

Apply for AI Grants India

Financial support for innovators building the future of AI in India.

Apply now

Chat · how to fine tune open source llms india

How to Fine-Tune Open-Source LLMs in India

  1. aigi

    Start with the task, not the model

    Fine-tuning is not a shortcut for every LLM problem. First define the behaviour you need: structured extraction, classification, instruction following, domain terminology, or conversational responses in a particular Indic language. If the model mainly needs current information, use retrieval-augmented generation instead of repeatedly retraining it. If it must follow a stable output format or learn a specialised writing style, supervised fine-tuning may be appropriate.

    Write an evaluation set before training. Include realistic prompts, difficult edge cases, code-mixed queries, spelling variation, and examples where the correct response is to refuse or ask for clarification. For Indian deployments, test transliteration and mixed-language input—for example, Hindi written in Roman script or queries combining English with Marathi, Tamil, Bengali, or Telugu. A clear target task will save more money than an oversized model.

    Teams new to open-source development can review Indian open-source AI developer projects for practical examples of how local builders select tools and structure contributions.

    Choose a model and licence carefully

    Select the smallest model that meets your quality, latency, and context-window requirements. A 3B–8B instruct model may be sufficient for classification, extraction, support workflows, and constrained generation. Larger models can improve reasoning and multilingual coverage, but they increase GPU memory, inference cost, and operational complexity.

    Check four things before downloading weights:

    • Indic coverage: Review performance in the target language, script, transliteration style, and code-mixed setting—not just English benchmarks.
    • Licence: Confirm whether commercial use, redistribution, hosted inference, and fine-tuned derivatives are allowed.
    • Tokenizer behaviour: Inspect token counts for representative Indian-language text. Inefficient tokenisation can raise both training and serving costs.
    • Base versus instruct model: Use a base model for continued pretraining or substantial domain adaptation; use an instruct model for supervised instruction tuning.

    For multilingual projects, compare the model against specialised resources discussed in this guide to low-resource Indic NLP. Do not assume that a model advertised as multilingual will perform equally across India’s languages.

    Build a legally usable, high-quality dataset

    Data quality usually matters more than dataset volume. Build examples from material you are entitled to use, and document its source, licence, language, domain, and processing history. Do not scrape personal data or copyrighted content indiscriminately. Remove phone numbers, email addresses, government identifiers, health information, and other sensitive fields unless you have a lawful, necessary basis and appropriate safeguards.

    A practical supervised fine-tuning record might contain:

    • A user instruction or conversation history
    • The ideal assistant response
    • Optional metadata such as language, domain, difficulty, and safety category
    • A source and consent or licence record

    Clean duplicates, boilerplate, broken encodings, prompt injection attempts, and contradictory answers. Preserve natural language variation rather than correcting every regional expression into formal Hindi or English. Keep a held-out test set that is never used during training. Deduplicate near-identical examples across train and test splits to avoid inflated scores.

    For customer or citizen-facing systems, include negative examples: requests the model should reject, uncertain questions it should escalate, and answers that must cite a source or return a structured error.

    Start with LoRA or QLoRA

    Full-parameter training is rarely the right first experiment for an Indian startup or student team. LoRA trains small adapter layers while leaving the base model frozen. QLoRA combines adapter training with quantised model weights, reducing GPU memory requirements while retaining strong performance for many instruction-tuning tasks.

    A typical workflow uses Python, PyTorch, Hugging Face Transformers, Datasets, and PEFT, with a bitsandbytes-style quantisation backend where supported. Begin with a small pilot:

    1. Convert data into the model’s chat or instruction template.
    2. Validate formatting by generating a few tokenised examples.
    3. Train a LoRA adapter on a small sample.
    4. Compare the adapter with the untouched base model on the held-out set.
    5. Adjust sequence length, learning rate, batch size, and number of epochs only after inspecting failures.

    Avoid copying generic settings from a different model. Learning rate, rank, target modules, packing, and maximum sequence length all interact with the model architecture and dataset. Use gradient accumulation and checkpointing when GPU memory is limited. Do not use batch normalisation as a default LLM fine-tuning step; transformer training generally relies on its existing normalisation layers.

    Plan India-based compute and costs

    You can prototype on a local NVIDIA GPU, a university lab, or a cloud GPU provider. The right choice depends on model size, sequence length, number of examples, and repetition of experiments. Track GPU hours, storage, data transfer, and inference costs—not only the training bill. Quantisation and smaller adapters can make deployment materially cheaper.

    Before renting a GPU:

    • Estimate tokens, training epochs, and expected experiments.
    • Confirm VRAM requirements for the quantisation and context length you plan to use.
    • Store datasets and checkpoints in encrypted, access-controlled storage.
    • Set budget alerts and automatic shutdowns.
    • Record package versions, hardware, seeds, and configuration files for reproducibility.

    For student or early-stage teams, the best open-source AI projects for beginners can help build the surrounding engineering skills before committing to expensive training runs.

    Evaluate language quality, safety, and usefulness

    Loss and perplexity are useful diagnostics, but they do not establish that a model is fit for production. Use a fixed evaluation suite with human review by fluent speakers of each target language. Measure task accuracy, exact-match or F1 for structured outputs, citation correctness, refusal quality, latency, and cost per request.

    For Indic applications, evaluate:

    • Native script, Romanised text, spelling variation, and code-mixing
    • Names, addresses, dates, currency, and local units
    • Dialect and register differences
    • Hallucinated legal, medical, financial, or public-service advice
    • Toxic, discriminatory, or culturally insensitive outputs
    • Prompt injection and attempts to extract training data

    Keep a comparison against the base model and a simple non-LLM baseline. If fine-tuning improves style but harms factuality, use retrieval, constraints, or better data rather than blindly increasing training time. Document known weaknesses and languages that were not adequately tested.

    Deploy the adapter responsibly

    Merge the adapter into the base model only when you have a reason to do so. Keeping adapters separate makes rollback, A/B testing, and model-specific deployments easier. Quantise for inference after validating quality, then serve behind an authenticated API with rate limits, logging, timeout controls, and input/output filtering.

    Production monitoring should track latency, error rates, token usage, language mix, user feedback, and unsafe or unanswered requests. Avoid storing raw prompts by default; redact sensitive fields and define retention rules. Route high-risk cases to a human instead of presenting model output as authoritative.

    The next step after training is often covered better by deployment guidance: see how to deploy open-source AI agents in production for operational patterns that also apply to LLM services.

    A practical 2026 checklist

    Before launch, confirm that you have:

    • A narrowly defined task and measurable acceptance criteria
    • A documented dataset licence, consent process, and privacy review
    • Separate training, validation, and untouched test data
    • Baseline results from the original model
    • LoRA or QLoRA experiments with reproducible configurations
    • Native-speaker evaluation for every supported language
    • Licence and model-card checks for the base model and datasets
    • Cost, latency, rollback, monitoring, and incident-response plans

    Fine-tuning open-source LLMs in India is most effective when treated as an engineering and governance process, not merely a GPU exercise. Start small, measure against real local usage, and improve the data and evaluation loop before scaling the model.

    Last updated 23 September 2026

AIGI may be inaccurate. Replies seeded from the guide above.