Small language models (SLMs) make specialised AI more accessible to Indian builders. A 1B–8B model can run on a workstation, support private workflows, and perform better than a much larger general model when it is tuned for a narrow task. This guide explains how to train small language models locally—with an emphasis on fine-tuning existing open-weight models rather than attempting expensive pretraining from scratch.
For multilingual products, pair this workflow with guidance on low-resource Indic natural language processing and model selection for open-source small language models for Hindi.
Decide What “Training” Means
There are three distinct projects:
- Pretraining from scratch: Learning language patterns from a large corpus. This requires substantial data, compute, engineering, and evaluation. It is rarely practical on one consumer GPU.
- Continued pretraining: Adapting a base model to a domain or language using large volumes of raw text. This can work locally for smaller models, but still demands careful data curation and long runs.
- Supervised fine-tuning (SFT): Teaching a capable base model to follow your task format using labelled examples. This is the best starting point for most teams.
Most local projects should begin with LoRA or QLoRA fine-tuning. If your goal is a Hindi or regional-language assistant, review fine-tuning Llama for Indian regional languages before choosing a dataset or tokenizer.
Choose a Model and Licence
Select the smallest model that can meet your quality target. Smaller models train faster, need less VRAM, and are easier to deploy at the edge.
- 1B–3B models: Suitable for classification, extraction, short-form support, and constrained assistants.
- 3B–8B models: Better for instruction following, multilingual responses, reasoning-lite workflows, and code assistance.
- Indian-language use cases: Check native language coverage, script support, tokenizer efficiency, and performance on code-mixed text—not just English benchmarks.
Candidate families may include Qwen, Llama, Mistral, Gemma, and Phi variants, subject to their current licence terms and hardware compatibility. Read the model card before commercial use. Confirm whether fine-tuned weights can be redistributed, whether attribution is required, and whether the model has restrictions that affect your product.
Hardware Planning for Local Fine-Tuning
VRAM is the primary constraint, but system RAM, storage, cooling, and power stability also affect productivity.
- 8GB GPU: Start with 1B–3B models, short sequence lengths, 4-bit QLoRA, gradient checkpointing, and batch size 1.
- 12–16GB GPU: Practical for many 3B–7B QLoRA experiments with conservative context lengths.
- 24GB GPU: A strong single-GPU setup for 7B-class QLoRA, larger batches, and faster iteration.
- 48GB or more: Useful for full fine-tuning of smaller models, longer contexts, or multi-GPU experiments.
Use at least 32GB system RAM for small experiments and 64GB for smoother dataset processing and checkpoint management. Keep datasets and checkpoints on an NVMe SSD. Consumer GPUs such as RTX 3060 12GB, RTX 4070-class cards, RTX 3090, and RTX 4090 remain practical options, but actual capacity depends on sequence length, precision, optimiser, and framework overhead. Avoid promising that a particular model will fit without testing a short run.
For CPU-only environments, use smaller models and expect substantially slower training. Apple Silicon and integrated GPUs can be useful for inference and data preparation, but CUDA remains the most broadly supported path for local fine-tuning.
Prepare High-Quality Training Data
Fine-tuning quality is usually limited by the dataset, not the number of GPU hours. Build examples that mirror real inputs and desired outputs.
For instruction tuning, JSONL records can use a consistent structure:
{"messages":[{"role":"user","content":"Summarise this customer complaint in Hindi."},{"role":"assistant","content":"ग्राहक की शिकायत का सारांश: ..."}]}Before training:
- Remove personal data, secrets, duplicated records, and unsupported claims.
- Normalise encoding and preserve Indic scripts correctly.
- Separate train, validation, and test sets by document or customer—not by randomly splitting near-duplicate lines.
- Include difficult examples, refusals, edge cases, spelling variation, and code-mixed language.
- Record provenance, licences, consent status, and transformation steps.
For Indian deployments, test Devanagari, Bengali, Tamil, Telugu, Kannada, Malayalam, Gujarati, Marathi, and Romanised text only when relevant to your users. Tokenisation can make some languages disproportionately expensive in context length, so inspect token counts before setting sequence limits.
Set Up a Reproducible Environment
Use Linux or WSL2, a dedicated virtual environment, and pinned package versions. A typical setup includes PyTorch, Transformers, Datasets, Accelerate, PEFT, TRL, and BitsAndBytes. Install the CUDA-compatible PyTorch build first, then verify the GPU:
nvidia-smi
python -c "import torch; print(torch.cuda.is_available())"Download model files before moving to an offline or restricted environment. Keep a record of the model revision, dataset version, prompt template, hyperparameters, random seed, and hardware. This matters when a local experiment becomes a production system.
Use QLoRA as the Default Baseline
QLoRA loads the base model in 4-bit precision and trains low-rank adapter weights. It reduces memory use while retaining the original model, making it a sensible baseline for a single workstation.
A reliable first run should use:
- 4-bit NF4 quantisation for supported models.
- LoRA rank between 8 and 32 as an initial search range.
- Learning rate around 1e-4 to 2e-4 for SFT, then adjust from validation results.
- Batch size 1–4 with gradient accumulation.
- A short sequence length initially, increasing only after confirming memory use.
- One to three epochs, with checkpointing and evaluation during training.
- BF16 where hardware supports it; otherwise use FP16 with loss scaling.
Do not assume that more epochs improve quality. Monitor training and validation loss, sample outputs, repetition, memorisation, and task-specific scores. If the model starts copying examples or losing general capability, reduce epochs, learning rate, or dataset redundancy.
A Practical Training Loop
1. Run a baseline: Evaluate the untouched base model on a fixed test set.
2. Train a small pilot: Use a few hundred or thousand curated examples to validate formatting and tooling.
3. Inspect outputs manually: Check factuality, language, tone, refusal behaviour, and formatting.
4. Scale the dataset: Add high-value examples rather than indiscriminately adding volume.
5. Tune one variable at a time: Change rank, learning rate, context, or epochs—not everything together.
6. Save adapters separately: This keeps experiments small and lets you compare or merge adapters later.
Gradient accumulation simulates a larger effective batch. Gradient checkpointing reduces memory at the cost of speed. Flash Attention can improve throughput when the GPU and software stack support it. These techniques solve different problems; enable them incrementally and profile each change.
Evaluate for the Real Indian Use Case
MMLU or similar public benchmarks are useful for broad comparison, but they do not establish product readiness. Build a held-out evaluation set that reflects your users and measure:
- Accuracy, exact match, F1, or extraction validity for structured tasks.
- Human preference and rubric scores for generation.
- Hindi and regional-language fluency, script correctness, and code-mixing behaviour.
- Hallucination, unsafe advice, privacy leakage, and refusal quality.
- Latency, peak VRAM, tokens per second, and cost per request.
Keep the test set private and versioned. For sensitive domains such as health, finance, or public services, have qualified reviewers assess high-risk outputs before deployment.
Deploy the Result Locally
Merge the adapter only when necessary; keeping it separate makes rollback easier. For local inference, convert to a supported format such as GGUF when appropriate, then test with Ollama, llama.cpp, or another runtime. GPU-serving stacks can provide higher throughput, while CPU or edge deployment may require aggressive quantisation.
Measure the deployed model—not just the training checkpoint. Quantisation can change accuracy, especially for multilingual generation and longer contexts. If your application needs vision or document understanding, a text-only SLM may not be sufficient; compare it with open-source vision-language models for Indian languages.
Common Failure Modes
- Out-of-memory errors: Reduce sequence length, use 4-bit loading, enable checkpointing, or lower micro-batch size.
- Gibberish output: Confirm the tokenizer, chat template, special tokens, and label masking.
- Training loss falls but quality worsens: Check for leakage, duplicates, overfitting, and a weak validation split.
- Poor Indic performance: Inspect token counts, add native-script examples, and evaluate each target language separately.
- Unstable experiments: Pin versions, save configurations, and avoid changing data and hyperparameters simultaneously.
- Thermal throttling: Monitor temperature, clocks, power draw, and fan curves during long runs.
A Sensible Starting Plan
For a first project, choose a 3B–7B instruction model with a permissive licence, prepare 1,000–5,000 clean examples, and run QLoRA on a 12GB–24GB GPU. Establish a baseline, train a short pilot, evaluate on a held-out Indian-language or domain-specific set, and only then expand the dataset or model size. This approach produces evidence quickly and avoids spending weeks optimising an unsuitable setup.
If you are building a local-first product, an Indic-language model, or a privacy-sensitive AI workflow in India, AI Grants India can help you explore funding, compute support, and ecosystem opportunities.