Small language models are now capable enough for many focused applications: customer-support assistants, document classifiers, extraction pipelines, code helpers, and Indic-language interfaces. Fine-tuning one locally can reduce cloud cost, keep sensitive data on your machine, and give you tighter control over latency and deployment.
The most reliable approach in 2026 is usually parameter-efficient fine-tuning (PEFT) rather than updating every model weight. For a decoder model, QLoRA with 4-bit quantisation is a practical starting point; for classification, a compact encoder model may be simpler and faster.
Choose the right fine-tuning objective
Start with the product behaviour you need, not the model name.
- Instruction or chat adaptation: use supervised fine-tuning (SFT) with examples formatted as conversations.
- Classification: use a sequence-classification head and labelled text.
- Extraction: train the model to return a strict JSON schema or use token classification for predictable entities.
- Domain language adaptation: continue pretraining on clean, domain-specific text before SFT.
- Style or terminology adaptation: use a small, high-quality instruction set rather than a large noisy corpus.
Fine-tuning does not reliably add fresh factual knowledge. For changing information, combine the model with retrieval or a database. The guide to best practices for fine-tuning LLMs on custom data is useful when deciding whether fine-tuning is the right tool.
Hardware and software requirements
You can run a small experiment on a modern laptop, but a discrete NVIDIA GPU makes training substantially easier. A rough starting point is:
- CPU-only: suitable for tokenisation and very small encoder experiments, but slow for generative-model training.
- 8 GB VRAM: workable for small 1B–3B models with 4-bit QLoRA, short sequences, and micro-batches.
- 12–16 GB VRAM: more comfortable for 3B–7B models with conservative sequence lengths.
- System RAM: 16 GB is a practical minimum; 32 GB or more helps with datasets and model loading.
- Storage: reserve space for model weights, caches, checkpoints, and logs.
Create an isolated environment and install a current PyTorch build matched to your CUDA version:
python -m venv .venv
source .venv/bin/activate # Windows: .venv\\Scripts\\activate
pip install -U torch transformers datasets accelerate peft bitsandbytes trlOn Apple Silicon or AMD hardware, verify backend support before choosing a training recipe. If local hardware is insufficient, prototype with a small model or use a rented GPU, while keeping the dataset and checkpoints encrypted.
Select a model and prepare data
Choose a model whose licence permits your intended commercial or research use. For Indian-language work, check tokenizer coverage, script support, and performance on the languages and code-mixed text you actually expect. A model that supports Hindi may still perform poorly on Marathi, Tamil, Bengali, or Romanised Hinglish.
For Indic projects, review low-resource Indic natural language processing and compare relevant open models in this guide to small language models for Hindi. If your target is several regional languages, evaluate each language separately instead of reporting one blended score.
Use JSONL and keep examples simple. An instruction-tuning record might look like this:
{"messages":[{"role":"user","content":"ग्राहक की शिकायत का संक्षिप्त उत्तर लिखें: डिलीवरी देर से हुई।"},{"role":"assistant","content":"देरी के लिए क्षमा करें। हम आपका ऑर्डर जल्द से जल्द पहुँचाने की पुष्टि कर रहे हैं।"}]}Your dataset should include:
- Realistic inputs, including spelling variation, code-switching, and short messages.
- Desired outputs that are accurate, concise, and consistent in format.
- A held-out validation set and a final test set created before training.
- Redacted personal, financial, medical, and business-sensitive information.
- Difficult and failure cases, not only clean examples.
Deduplicate near-identical records, remove contradictory labels, and inspect random samples manually. Ten thousand consistent examples can outperform a much larger noisy dataset.
Fine-tune with QLoRA
QLoRA loads the base model in 4-bit precision and trains small adapter weights. This reduces memory use while preserving the original checkpoint. The following skeleton uses TRL's supervised fine-tuning trainer; argument names can vary between library releases, so check the installed documentation.
import torch
from datasets import load_dataset
from transformers import AutoTokenizer, AutoModelForCausalLM, BitsAndBytesConfig
from peft import LoraConfig
from trl import SFTTrainer, SFTConfig
model_id = "your-compatible-small-causal-model"
data = load_dataset("json", data_files={
"train": "train.jsonl", "validation": "validation.jsonl"
})
tokenizer = AutoTokenizer.from_pretrained(model_id)
if tokenizer.pad_token is None:
tokenizer.pad_token = tokenizer.eos_token
quant = BitsAndBytesConfig(
load_in_4bit=True,
bnb_4bit_quant_type="nf4",
bnb_4bit_compute_dtype=torch.bfloat16,
)
model = AutoModelForCausalLM.from_pretrained(
model_id, quantization_config=quant, device_map="auto"
)
lora = LoraConfig(
r=16, lora_alpha=32, lora_dropout=0.05,
target_modules=["q_proj", "k_proj", "v_proj", "o_proj"],
task_type="CAUSAL_LM"
)
config = SFTConfig(
output_dir="./adapter",
num_train_epochs=2,
per_device_train_batch_size=1,
gradient_accumulation_steps=8,
learning_rate=2e-4,
max_seq_length=1024,
logging_steps=10,
eval_strategy="steps",
eval_steps=100,
save_steps=100,
bf16=torch.cuda.is_available(),
)
trainer = SFTTrainer(
model=model, args=config,
train_dataset=data["train"],
eval_dataset=data["validation"],
processing_class=tokenizer,
peft_config=lora,
)
trainer.train()
trainer.save_model("./adapter")
tokenizer.save_pretrained("./adapter")If you see out-of-memory errors, lower max_seq_length, reduce the number of LoRA target modules, use gradient checkpointing, or reduce the micro-batch size. Gradient accumulation increases the effective batch size without requiring all examples in memory.
For language-specific adaptation, the workflow in fine-tuning Llama for Indian regional languages offers a useful comparison of data and evaluation concerns.
Evaluate beyond training loss
A falling training loss is not proof that the model is useful. Track validation loss, but also create task-specific tests for:
- Accuracy, macro-F1, and confusion matrices for classification.
- Exact-match and schema validity for extraction.
- Factuality, refusal behaviour, and instruction compliance for assistants.
- Script, spelling, and code-mixing performance for Indic text.
- Latency, memory use, and tokens per second on your target device.
Compare the fine-tuned adapter with the base model on the same test set. Test prompt variations, empty inputs, adversarial instructions, and long context. Keep a small regression suite in version control so every new dataset or hyperparameter change is measurable.
Save, quantise, and deploy locally
Keep the base model, adapter, tokenizer, dataset version, training configuration, and evaluation results together. Adapters are smaller and easier to version than merged models. Merge only when your serving stack requires it, and retain the original adapter so you can reproduce the result.
For CPU deployment, convert a compatible merged model to a format supported by a local runtime such as llama.cpp. For GPU serving, use a Transformers-based service or an inference server compatible with your model architecture. Add input validation, output-length limits, logging with sensitive fields removed, and a fallback for low-confidence or malformed responses.
Common mistakes
- Fine-tuning before establishing a base-model baseline.
- Using too many epochs and memorising a small dataset.
- Mixing chat templates between training and inference.
- Evaluating only in English when the product serves Indian languages.
- Training on unredacted customer or health data.
- Assuming a larger model is always better than a well-curated smaller one.
Start with one narrowly defined task, 200–1,000 carefully reviewed examples, and a reproducible baseline. Expand the dataset only after the error analysis shows what the model still gets wrong.