0tokens

Apply for AI Grants India

Financial support for innovators building the future of AI in India.

Apply now

Chat · how to fine tune small language models on github

How to Fine-Tune Small Language Models on GitHub

  1. aigi

    What fine-tuning a small language model means

    Fine-tuning starts with a pre-trained language model and adapts it to a narrower task, domain, tone, or language mix. A small model may have a few hundred million to several billion parameters. It will not match a frontier model on every capability, but it can be faster, cheaper, easier to run privately, and more dependable for a defined workflow.

    GitHub is useful because the complete training stack is visible: model code, configuration files, data-processing scripts, evaluation harnesses, issue discussions, and deployment examples. Treat a repository as engineering infrastructure rather than as proof that a model is production-ready. Check its license, recent maintenance, hardware assumptions, and whether the training scripts reproduce the published results.

    For Indian teams, compact models are particularly practical for customer-support triage, document classification, transliteration, retrieval-assisted answers, and domain-specific text in English or Indic languages. If your project involves limited training data or low-resource languages, first review this builder’s guide to low-resource Indic NLP.

    Choose the right training objective

    Do not begin by selecting a model. Begin with the behaviour you need:

    • Classification: route tickets, detect intent, identify sentiment, or assign a document label.
    • Token or span extraction: identify names, account numbers, clauses, products, or locations.
    • Instruction tuning: teach a model to produce structured responses from examples.
    • Continued pre-training: adapt the model to a specialised vocabulary or writing style using large volumes of unlabeled text.
    • Preference optimisation: improve response style after supervised examples, provided you have trustworthy preference data.

    A classification model is usually cheaper and easier to evaluate than a generative assistant. If the requirement is factual answers over changing documents, fine-tuning may be the wrong first step; use retrieval-augmented generation and reserve fine-tuning for formatting, intent detection, or domain language.

    Select a GitHub-based stack

    A practical stack in 2026 usually combines a model repository, the Hugging Face Transformers ecosystem, a dataset library, and a parameter-efficient training method such as LoRA or QLoRA. PEFT updates a small set of adapter weights instead of all model parameters, reducing GPU memory, training time, and storage requirements. Quantisation can reduce memory further, but test quality carefully because aggressive quantisation may harm small models disproportionately.

    Useful repository signals include:

    • A clear model card with licence, training data notes, known limitations, and supported languages.
    • Pinned dependency versions and a reproducible environment file.
    • Dataset validation, train-validation-test splitting, and leakage checks.
    • Evaluation scripts rather than only a single headline score.
    • Export and inference examples for the hardware you actually own.

    Use the best practices for fine-tuning LLMs on custom data as a companion checklist, and inspect promising beginner-friendly repositories through this guide to open-source AI projects on GitHub.

    Prepare data before touching the GPU

    Data quality is the main determinant of a useful fine-tune. Define the input and expected output precisely, then create examples that reflect real requests rather than idealised demonstrations. Remove duplicated records, secrets, personal information, corrupted text, and labels inferred from fields that will not exist at inference time.

    For instruction tuning, store examples in a consistent schema such as messages with system, user, and assistant roles. For classification, keep labels stable and document their definitions. Include difficult, ambiguous, and negative examples. If the model will handle Indian customer interactions, represent code-switching, spelling variation, Romanised Indic text, local names, dates, currency formats, and domain abbreviations where they occur in production.

    Create splits before training. Keep near-duplicates and conversations from the same source in one split to avoid leakage. A small but carefully reviewed validation set is more valuable than a large, noisy one. Maintain a data card recording provenance, consent or licensing basis, transformations, and known gaps.

    A minimal LoRA workflow

    A typical GitHub project should separate data preparation, training, evaluation, and inference. The exact API changes across library versions, so pin dependencies and follow the selected model’s chat template. Conceptually, the training loop looks like this:

    from transformers import AutoModelForCausalLM, AutoTokenizer
    from peft import LoraConfig, get_peft_model
    
    model_id = "your-org/your-small-model"
    tokenizer = AutoTokenizer.from_pretrained(model_id)
    model = AutoModelForCausalLM.from_pretrained(model_id)
    
    config = LoraConfig(
        r=16,
        lora_alpha=32,
        lora_dropout=0.05,
        target_modules=["q_proj", "v_proj"],
        task_type="CAUSAL_LM",
    )
    model = get_peft_model(model, config)
    # Tokenise validated examples, train with your Trainer or SFTTrainer,
    # then save the adapter and tokenizer.

    Start with a short pilot run. Confirm that loss decreases, samples follow the expected format, and the model has not memorised private examples. Only then increase the dataset, sequence length, or number of steps. Save checkpoints, configuration, random seeds, library versions, and the base-model commit so another developer can reproduce the result.

    Budget and hardware planning

    A small model with QLoRA may run on a single consumer GPU, a rented cloud GPU, or an institutional workstation, but memory needs depend on parameter count, sequence length, batch size, optimiser, and quantisation. Gradient accumulation can simulate a larger batch, while gradient checkpointing trades compute for memory. Reduce sequence length before reducing data quality; long contexts are expensive and often unnecessary for classification or short support conversations.

    For an India-based project, compare cloud GPU pricing in rupees, data-transfer costs, storage, and whether sensitive data may leave your environment. Keep a cost log containing GPU hours, failed runs, storage, and inference cost. A model that trains cheaply but requires expensive serving or constant GPU uptime may not be the best choice.

    Evaluate beyond training loss

    Use task-specific metrics and a fixed test set that was never used for model selection. Classification projects should report per-class precision, recall, F1, and a confusion matrix, especially when minority classes matter. Generative projects need structured checks for format compliance, factuality, refusal behaviour, language quality, and harmful or sensitive outputs.

    Compare three systems: the untuned base model, the fine-tuned model, and a simple non-LLM baseline where possible. Test on shifted examples such as spelling errors, mixed languages, short inputs, long inputs, and out-of-domain requests. Have Indian-language speakers review outputs when the model handles Hindi, Tamil, Bengali, Marathi, or other Indic languages; automated English metrics will not capture all failures.

    Track regressions in a versioned evaluation set. A fine-tune that improves one workflow but weakens general instruction following, privacy behaviour, or refusal boundaries may not be ready for release. For larger deployments, add monitoring, rate limits, prompt and output redaction, and a human escalation path.

    Publish and deploy responsibly

    A useful GitHub release includes the training configuration, dataset documentation, evaluation results, licence information, limitations, inference instructions, and a model card. Do not commit API keys, raw personal data, private customer conversations, or model weights whose licence prohibits redistribution. Use Git LFS or an approved model registry for large artifacts and scan commits for secrets.

    Before production, test the adapter with the exact tokenizer, prompt template, quantisation settings, and serving runtime you will use. Merge adapters only after comparing quality and memory usage. For deployment options, separate offline batch processing from low-latency APIs; the operational requirements differ. If you later need a managed Kubernetes deployment, this guide to deploying deep learning models on GKE covers the infrastructure considerations.

    Common mistakes to avoid

    • Fine-tuning before defining a measurable task.
    • Training on synthetic or duplicated examples without human review.
    • Using a random split that leaks documents or users across datasets.
    • Treating a lower training loss as proof of better production quality.
    • Ignoring the model, dataset, and dependency licences.
    • Overlooking code-switching and Indic-language edge cases.
    • Publishing checkpoints or logs that contain personal or confidential data.
    • Choosing a model that cannot meet serving latency or hardware constraints.

    FAQ

    Can I fine-tune any model found on GitHub?

    No. The repository may contain code without redistributable weights, or the model licence may restrict commercial use or adaptation. Verify the licence, architecture support, tokenizer, training objective, and hardware requirements before committing to a workflow.

    Is LoRA better than full fine-tuning?

    LoRA is often the better starting point for small teams because it is cheaper, faster, and produces compact adapters. Full fine-tuning can be useful with substantial data and compute, but it increases storage, experimentation cost, and the risk of losing useful base-model capabilities.

    How much data do I need?

    There is no universal number. A few hundred carefully reviewed examples can improve a narrow format or classification task, while robust domain adaptation may need thousands or more. Run a pilot, compare against the untuned model, and add examples that target observed failures rather than merely increasing volume.

    Should I contribute improvements back to GitHub?

    Yes, when the licence and organisational policy permit it. Share reproducible bug fixes, documentation, evaluation scripts, or anonymised examples. Read this guide on contributing to AI GitHub repositories in India before opening a pull request.

    Last updated 23 September 2026

AIGI may be inaccurate. Replies seeded from the guide above.