Small language models (SLMs) are often the better engineering choice when an application needs predictable latency, lower inference cost, offline capability, or tighter data controls. Fine-tuning adapts an existing model to a defined task or domain; it is not a substitute for a larger model, better product design, or retrieval when the problem requires fresh knowledge.
For Indian startups, the strongest use cases are usually narrow and measurable: classifying support tickets, extracting fields from invoices, detecting abusive content in regional languages, routing voice transcripts, or generating structured replies. Start with the task and success metric, then decide whether fine-tuning is justified.
Decide whether fine-tuning is the right approach
Fine-tuning is useful when the model must consistently learn a behaviour, format, label set, tone, or domain vocabulary. It is usually less suitable when the main problem is missing or changing information. In that case, use retrieval-augmented generation, a database query, or tool calling.
A practical decision rule:
- Use supervised fine-tuning for classification, extraction, rewriting, and tightly specified response formats.
- Use continued pretraining when the model lacks exposure to a domain’s language, such as legal, agricultural, or Indic-language text.
- Use preference tuning only when you have reliable preference data and a clear quality rubric.
- Use prompting or retrieval first if you have little labelled data or need current facts.
If your project involves Hindi or another Indian language, review the considerations in this guide to open-source small language models for Hindi and the broader low-resource Indic NLP builder’s guide. Tokenisation quality, script variation, transliteration, code-switching, and dialect coverage can matter more than parameter count.
Choose a base model and adaptation method
Select a model by task, licence, language coverage, context length, hardware needs, and community support—not by benchmark score alone. Encoder models such as BERT variants remain strong for classification and token labelling. Decoder-only models are more suitable for instruction following, structured generation, and conversational workflows.
For a first experiment, compare two or three small checkpoints on a fixed validation set. Record parameter count, memory use, tokenizer behaviour, licence obligations, and inference speed on the hardware you expect to deploy. A model that performs slightly better but cannot meet an Indian mobile or on-premise latency target may be the weaker product choice.
Full fine-tuning updates every parameter and can be expensive. Parameter-efficient methods are often preferable:
- LoRA: learns small low-rank adapter matrices while freezing the base model.
- QLoRA: combines quantised base weights with LoRA to reduce GPU memory.
- Adapters: keep task-specific weights separate, allowing multiple capabilities on one base model.
- Prompt or prefix tuning: modifies a small learned input representation, though results vary by task.
Use the recommendations in best practices for fine-tuning LLMs on custom data when deciding between full and parameter-efficient training.
Build a trustworthy training dataset
Data quality determines the ceiling of the result. Define the input, expected output, and acceptable edge cases before collecting examples. For a support classifier, for example, specify whether an ambiguous request receives one primary label or multiple labels. For extraction, define how missing, uncertain, and conflicting fields should be represented.
Aim for representative examples rather than a large volume of near-duplicates. Include:
- Realistic spelling mistakes, abbreviations, emojis, and code-switching.
- Hindi-English and other transliterated inputs where users actually write that way.
- Different devices, document layouts, accents, and regional terminology.
- Difficult negatives and examples where the correct answer is to abstain.
- A small, carefully reviewed set of high-impact failure cases.
Remove personal data unless it is essential and lawfully processed. Hash or redact phone numbers, Aadhaar details, financial identifiers, and unnecessary names. Keep dataset versions, annotation instructions, reviewer decisions, and source permissions in a reproducible record.
Split data by user, document, conversation, or time—not merely by random rows. Otherwise, near-duplicate examples can leak into validation and produce an inflated score. Hold out a final test set that is not used for model or hyperparameter decisions.
Configure and run the first training experiment
Use a reproducible environment with a pinned Python version, framework versions, model revision, random seed, dataset hash, and training configuration. Hugging Face Transformers with PyTorch is a common starting point, while tools such as PEFT and bitsandbytes support LoRA and quantised training.
Begin conservatively. Typical starting ranges are:
- Learning rate: around 1e-5 to 5e-5 for full fine-tuning; LoRA may tolerate higher rates, but validate rather than assume.
- Epochs: one to three for a large or repetitive dataset; more is not automatically better.
- Effective batch size: increase through gradient accumulation if GPU memory is limited.
- Sequence length: set it from the real input distribution; unnecessary context wastes memory.
- Warm-up and scheduling: use a short warm-up and monitor whether loss and task metrics move together.
Use mixed precision where supported, gradient checkpointing for memory pressure, and early stopping when validation quality stops improving. Save checkpoints and log training loss, validation metrics, learning rate, throughput, and peak memory. For generative tasks, train on the exact prompt-and-response structure used in production, including delimiters and output schema.
Do not assume that lower training loss means better product performance. A model can memorise templates while failing on spelling variation, long inputs, or unfamiliar users.
Evaluate for accuracy, safety, and cost
Choose metrics that reflect the decision the product makes. Accuracy can hide poor minority-class performance; use macro-F1, per-class precision and recall, confusion matrices, and calibration for classification. For extraction, measure field-level exact match and partial match. For generation, combine automated checks with blinded human review against a rubric.
Create an evaluation slice for Indian deployment conditions:
- Each supported language and script.
- Transliteration and code-switching.
- Urban and rural terminology where relevant.
- Low-quality scans, noisy speech transcripts, or short messages.
- Sensitive, ambiguous, and adversarial inputs.
Track abstention and escalation quality. In finance, healthcare, education, and public-service workflows, a safe refusal or human handoff may be more valuable than a confident incorrect answer. Test for memorisation and unintended disclosure by probing with training examples and sensitive-looking prompts.
Measure production constraints before launch: p50 and p95 latency, tokens per second, RAM or VRAM use, model download size, concurrency, and cost per request. For mobile or edge use, follow this 2026 guide to AI model optimisation for mobile devices.
Deploy, monitor, and iterate
Export the adapter or merged model only after confirming that the deployment runtime supports the chosen format. Quantisation can reduce memory and latency, but test quality after quantisation on the same evaluation slices. Package the tokenizer with the model, validate input limits, and expose structured errors rather than silently truncating requests.
A production service should include:
- Versioned model and dataset identifiers in every prediction log.
- Input validation, rate limits, authentication, and redaction of sensitive logs.
- Confidence thresholds and human review for high-risk cases.
- Drift monitoring by language, label, customer segment, and input length.
- A rollback path and a small regression suite run on every release.
For teams serving multiple Indic languages, avoid treating translation as a complete evaluation strategy. Review native-language outputs and include local annotators who understand cultural and domain context. If your application also uses speech, vision, or multimodal inputs, test the full pipeline rather than evaluating the language model in isolation.
A practical starter plan
For a first two-week experiment, define one task and one primary metric, collect and review a few thousand representative examples if available, establish a prompt or unfine-tuned baseline, then train a LoRA adapter on a small checkpoint. Compare against the baseline on a fixed multilingual and edge-case test set. Only proceed to larger training if the gain is material and survives latency, cost, and safety checks.
Fine-tuning works best as a disciplined product experiment: narrow scope, clean data, reproducible runs, meaningful evaluation, and a deployment plan from the beginning. Small models can deliver excellent results, but only when the task boundaries and failure handling are designed as carefully as the training loop.