0tokens

Apply for AI Grants India

Financial support for innovators building the future of AI in India.

Apply now

Chat · how to fine tune small language models locally

How to Fine-Tune Small Language Models Locally

  1. aigi

    Small language models are now capable enough for many focused applications: customer-support assistants, document classifiers, extraction pipelines, code helpers, and Indic-language interfaces. Fine-tuning one locally can reduce cloud cost, keep sensitive data on your machine, and give you tighter control over latency and deployment.

    The most reliable approach in 2026 is usually parameter-efficient fine-tuning (PEFT) rather than updating every model weight. For a decoder model, QLoRA with 4-bit quantisation is a practical starting point; for classification, a compact encoder model may be simpler and faster.

    Choose the right fine-tuning objective

    Start with the product behaviour you need, not the model name.

    • Instruction or chat adaptation: use supervised fine-tuning (SFT) with examples formatted as conversations.
    • Classification: use a sequence-classification head and labelled text.
    • Extraction: train the model to return a strict JSON schema or use token classification for predictable entities.
    • Domain language adaptation: continue pretraining on clean, domain-specific text before SFT.
    • Style or terminology adaptation: use a small, high-quality instruction set rather than a large noisy corpus.

    Fine-tuning does not reliably add fresh factual knowledge. For changing information, combine the model with retrieval or a database. The guide to best practices for fine-tuning LLMs on custom data is useful when deciding whether fine-tuning is the right tool.

    Hardware and software requirements

    You can run a small experiment on a modern laptop, but a discrete NVIDIA GPU makes training substantially easier. A rough starting point is:

    • CPU-only: suitable for tokenisation and very small encoder experiments, but slow for generative-model training.
    • 8 GB VRAM: workable for small 1B–3B models with 4-bit QLoRA, short sequences, and micro-batches.
    • 12–16 GB VRAM: more comfortable for 3B–7B models with conservative sequence lengths.
    • System RAM: 16 GB is a practical minimum; 32 GB or more helps with datasets and model loading.
    • Storage: reserve space for model weights, caches, checkpoints, and logs.

    Create an isolated environment and install a current PyTorch build matched to your CUDA version:

    python -m venv .venv
    source .venv/bin/activate          # Windows: .venv\\Scripts\\activate
    pip install -U torch transformers datasets accelerate peft bitsandbytes trl

    On Apple Silicon or AMD hardware, verify backend support before choosing a training recipe. If local hardware is insufficient, prototype with a small model or use a rented GPU, while keeping the dataset and checkpoints encrypted.

    Select a model and prepare data

    Choose a model whose licence permits your intended commercial or research use. For Indian-language work, check tokenizer coverage, script support, and performance on the languages and code-mixed text you actually expect. A model that supports Hindi may still perform poorly on Marathi, Tamil, Bengali, or Romanised Hinglish.

    For Indic projects, review low-resource Indic natural language processing and compare relevant open models in this guide to small language models for Hindi. If your target is several regional languages, evaluate each language separately instead of reporting one blended score.

    Use JSONL and keep examples simple. An instruction-tuning record might look like this:

    {"messages":[{"role":"user","content":"ग्राहक की शिकायत का संक्षिप्त उत्तर लिखें: डिलीवरी देर से हुई।"},{"role":"assistant","content":"देरी के लिए क्षमा करें। हम आपका ऑर्डर जल्द से जल्द पहुँचाने की पुष्टि कर रहे हैं।"}]}

    Your dataset should include:

    • Realistic inputs, including spelling variation, code-switching, and short messages.
    • Desired outputs that are accurate, concise, and consistent in format.
    • A held-out validation set and a final test set created before training.
    • Redacted personal, financial, medical, and business-sensitive information.
    • Difficult and failure cases, not only clean examples.

    Deduplicate near-identical records, remove contradictory labels, and inspect random samples manually. Ten thousand consistent examples can outperform a much larger noisy dataset.

    Fine-tune with QLoRA

    QLoRA loads the base model in 4-bit precision and trains small adapter weights. This reduces memory use while preserving the original checkpoint. The following skeleton uses TRL's supervised fine-tuning trainer; argument names can vary between library releases, so check the installed documentation.

    import torch
    from datasets import load_dataset
    from transformers import AutoTokenizer, AutoModelForCausalLM, BitsAndBytesConfig
    from peft import LoraConfig
    from trl import SFTTrainer, SFTConfig
    
    model_id = "your-compatible-small-causal-model"
    data = load_dataset("json", data_files={
        "train": "train.jsonl", "validation": "validation.jsonl"
    })
    
    tokenizer = AutoTokenizer.from_pretrained(model_id)
    if tokenizer.pad_token is None:
        tokenizer.pad_token = tokenizer.eos_token
    
    quant = BitsAndBytesConfig(
        load_in_4bit=True,
        bnb_4bit_quant_type="nf4",
        bnb_4bit_compute_dtype=torch.bfloat16,
    )
    model = AutoModelForCausalLM.from_pretrained(
        model_id, quantization_config=quant, device_map="auto"
    )
    
    lora = LoraConfig(
        r=16, lora_alpha=32, lora_dropout=0.05,
        target_modules=["q_proj", "k_proj", "v_proj", "o_proj"],
        task_type="CAUSAL_LM"
    )
    
    config = SFTConfig(
        output_dir="./adapter",
        num_train_epochs=2,
        per_device_train_batch_size=1,
        gradient_accumulation_steps=8,
        learning_rate=2e-4,
        max_seq_length=1024,
        logging_steps=10,
        eval_strategy="steps",
        eval_steps=100,
        save_steps=100,
        bf16=torch.cuda.is_available(),
    )
    
    trainer = SFTTrainer(
        model=model, args=config,
        train_dataset=data["train"],
        eval_dataset=data["validation"],
        processing_class=tokenizer,
        peft_config=lora,
    )
    trainer.train()
    trainer.save_model("./adapter")
    tokenizer.save_pretrained("./adapter")

    If you see out-of-memory errors, lower max_seq_length, reduce the number of LoRA target modules, use gradient checkpointing, or reduce the micro-batch size. Gradient accumulation increases the effective batch size without requiring all examples in memory.

    For language-specific adaptation, the workflow in fine-tuning Llama for Indian regional languages offers a useful comparison of data and evaluation concerns.

    Evaluate beyond training loss

    A falling training loss is not proof that the model is useful. Track validation loss, but also create task-specific tests for:

    • Accuracy, macro-F1, and confusion matrices for classification.
    • Exact-match and schema validity for extraction.
    • Factuality, refusal behaviour, and instruction compliance for assistants.
    • Script, spelling, and code-mixing performance for Indic text.
    • Latency, memory use, and tokens per second on your target device.

    Compare the fine-tuned adapter with the base model on the same test set. Test prompt variations, empty inputs, adversarial instructions, and long context. Keep a small regression suite in version control so every new dataset or hyperparameter change is measurable.

    Save, quantise, and deploy locally

    Keep the base model, adapter, tokenizer, dataset version, training configuration, and evaluation results together. Adapters are smaller and easier to version than merged models. Merge only when your serving stack requires it, and retain the original adapter so you can reproduce the result.

    For CPU deployment, convert a compatible merged model to a format supported by a local runtime such as llama.cpp. For GPU serving, use a Transformers-based service or an inference server compatible with your model architecture. Add input validation, output-length limits, logging with sensitive fields removed, and a fallback for low-confidence or malformed responses.

    Common mistakes

    • Fine-tuning before establishing a base-model baseline.
    • Using too many epochs and memorising a small dataset.
    • Mixing chat templates between training and inference.
    • Evaluating only in English when the product serves Indian languages.
    • Training on unredacted customer or health data.
    • Assuming a larger model is always better than a well-curated smaller one.

    Start with one narrowly defined task, 200–1,000 carefully reviewed examples, and a reproducible baseline. Expand the dataset only after the error analysis shows what the model still gets wrong.

    Last updated 23 September 2026

AIGI may be inaccurate. Replies seeded from the guide above.