Small language models (SLMs) are no longer merely compressed alternatives to large models. With the right data, objective, and deployment constraints, a model with millions or a few billion parameters can deliver reliable results for classification, extraction, summarisation, retrieval assistance, and focused conversational workflows. The goal is not to make a small model behave like a frontier model at every task. It is to build the smallest model that meets a clearly defined quality, latency, cost, and privacy target.
For Indian builders, this often means supporting noisy user input, code-mixed English and Indic languages, regional terminology, and deployment on modest cloud or edge infrastructure. The following practices provide a practical path from problem definition to production.
Start with a narrow, testable job
Before selecting a base model, define what the model must do and what it should refuse to do. “Build a chatbot” is too broad. “Extract invoice fields from Hindi-English WhatsApp messages with 95% field-level accuracy” is measurable.
Document:
- Input languages, formats, maximum context length, and expected traffic.
- Required outputs, including a strict schema where possible.
- Latency, memory, cost, and offline or on-device requirements.
- Accepted error rates for each user segment.
- Safety boundaries and escalation paths for uncertain cases.
Create a small but representative evaluation set before training. Include ordinary examples, difficult cases, spelling errors, code-mixing, abbreviations, adversarial prompts, and examples from different Indian regions. This prevents optimisation against an overly clean benchmark.
Curate data before increasing its volume
Data quality usually matters more than adding another large, weakly filtered corpus. Build a data pipeline that records source, licence, language, date, processing steps, and quality signals. Remove duplicates at both document and near-duplicate levels; otherwise, validation scores may look strong because the model has seen almost identical text during training.
Prioritise:
- Task relevance: examples should resemble production requests and desired responses.
- Language coverage: preserve Devanagari and other Indic scripts rather than transliterating everything into Latin characters.
- Representation: include dialects, accents in transcripts, varied names, and different levels of literacy.
- Privacy: remove personal information, secrets, account numbers, and unnecessary conversation history.
- Licensing: verify that training and redistribution rights cover the intended use.
For low-resource Indic applications, the low-resource Indic NLP builder’s guide offers useful context on script variation, annotation, and language-specific evaluation. Do not rely on synthetic data alone: generated examples can expand coverage, but they may repeat the teacher model’s errors and produce unnatural language.
Choose the smallest suitable architecture
Start with an existing open-weight decoder or encoder model that has a compatible licence, tokenizer, framework, and community support. Parameter count is only one consideration. A smaller model with a poor tokenizer for Marathi, Tamil, or mixed Hindi-English text may be less efficient than a slightly larger model with better token coverage.
Choose the training route deliberately:
- Prompting or retrieval: use this when the task is mainly knowledge access and the model already understands the required format.
- Supervised fine-tuning: use labelled examples to teach behaviour, style, extraction, or domain terminology.
- Continued pretraining: use domain text when the model lacks foundational vocabulary or writing patterns.
- Distillation: train the compact model on carefully filtered outputs, rationales where appropriate, or task labels from a stronger teacher.
- Adapters such as LoRA: use parameter-efficient updates when GPU memory, iteration speed, or multiple domain variants matter.
Fine-tuning does not replace retrieval for frequently changing facts. For a deeper implementation workflow, see these best practices for fine-tuning LLMs on custom data.
Design the training set for behaviour, not just language
Instruction examples should show the exact input-output contract. Include valid responses, concise refusals, uncertainty handling, and malformed inputs. If the production application requires JSON, train and evaluate valid JSON rather than judging free-form prose.
Keep training, validation, and test sets separated by user, document, organisation, or time period—not only by random rows. A random split can leak templates and memorised content. Maintain slices for:
- Each supported language and code-mixed combination.
- Short and long inputs.
- High-frequency and rare intents.
- Noisy spelling, speech transcription, and informal messages.
- Safety-sensitive or high-impact scenarios.
Use augmentation carefully. Back-translation, paraphrasing, and spelling variation can improve robustness, but manually inspect a sample for meaning drift. Synthetic data should be labelled, versioned, and balanced against real examples.
Train efficiently and track reproducible experiments
Small models still benefit from disciplined training. Begin with conservative runs and change one major variable at a time. Track the base checkpoint, data version, tokenizer, sequence length, learning rate, batch size, warm-up, optimiser, precision, random seed, and evaluation results.
Practical techniques include:
- Use gradient accumulation when memory limits the effective batch size.
- Apply mixed precision where hardware supports it, while checking numerical stability.
- Pack compatible short sequences to reduce padding waste.
- Use early stopping when validation loss and task metrics stop improving.
- Compare checkpoints on fixed evaluation slices, not only aggregate loss.
- Save frequently enough to recover from infrastructure interruptions.
Do not select a model solely because its training loss is lower. Overfitting can appear as fluent but less faithful output, memorisation of sensitive text, or declining performance on minority languages.
Evaluate quality, safety, and cost together
Perplexity is useful for language modelling, but it is not a sufficient product metric. Combine automated measures with human review and task-specific scoring. Depending on the application, measure exact match, field-level accuracy, F1, calibration, groundedness, refusal quality, latency, memory use, and cost per request.
Build an evaluation harness that runs on every model candidate. Include regression tests for previously fixed errors and compare performance by language, script, intent, and user type. For generative tasks, use rubric-based human assessment or a carefully validated judge model, with periodic manual audits to detect evaluator bias.
Test for:
- Hallucinated facts and unsupported citations.
- Prompt injection and instruction-conflict behaviour.
- Leakage of training examples or personal data.
- Toxic, discriminatory, or culturally inappropriate outputs.
- Failure under long, noisy, or mixed-language input.
For sensitive sectors, route low-confidence cases to a human rather than forcing the SLM to answer.
Compress only after quality is stable
Quantisation, pruning, distillation, and efficient runtimes can reduce memory and latency, but each may alter output quality. Establish a full-precision baseline first, then benchmark 8-bit, 4-bit, or hardware-specific formats on the same test suite. Measure end-to-end latency—including tokenisation, retrieval, and post-processing—not just model inference time.
For Indian deployments, test on the actual target hardware and scripts. A model that performs well on a high-end GPU may be uneconomical on a CPU instance or mobile device. Keep a larger fallback model for difficult cases if the product can tolerate routing complexity.
Operate the model as a monitored system
Production quality depends on more than weights. Version prompts, retrieval indexes, adapters, tokenisers, safety filters, and data-processing code. Log privacy-safe inputs and outputs, latency, token counts, refusal rates, and escalation frequency. Monitor drift by language and intent rather than relying on one overall score.
Set a retraining policy based on observed failures. Add reviewed examples to a governed dataset, rerun regression tests, and release updates gradually. Maintain rollback capability and document known limitations for support and operations teams.
A practical build sequence
A lean implementation can follow this order:
1. Define the task, constraints, and unacceptable failures.
2. Create a representative test set before training.
3. Establish a prompt or retrieval baseline with an existing model.
4. Clean, deduplicate, license-check, and label the training data.
5. Fine-tune or distil a suitable base model with reproducible tracking.
6. Evaluate by language, task slice, safety category, latency, and cost.
7. Quantise and benchmark on target infrastructure.
8. Launch gradually with monitoring, human escalation, and rollback.
Teams building language products for India can also study open-source small language models for Hindi when comparing tokenisation, licensing, and Indic-language readiness. The strongest SLM is rarely the largest available checkpoint; it is the one whose data, evaluation, and operating design match the job.
FAQs
How small should a small language model be?
There is no universal threshold. Choose the smallest model that meets quality and latency targets on your real evaluation set. A compact model may be ideal for extraction, routing, and classification, while a larger model may be needed for open-ended reasoning.
Should I train from scratch?
Usually not. Fine-tuning, continued pretraining, or distillation from an open-weight base is more practical unless you have substantial domain data, compute, and a strong reason to control the full pretraining stack.
Is synthetic data safe to use?
Yes, if it is reviewed, deduplicated, balanced with real examples, and checked for factual and cultural errors. Treat it as an augmentation strategy, not a substitute for production data.
What should Indian teams test first?
Test language and script coverage, code-mixing, transliteration, noisy inputs, regional terminology, privacy leakage, and performance on the hardware you intend to operate.
Apply for AI Grants India
If you are building an efficient language model, Indic-language application, or privacy-preserving AI product in India, explore funding opportunities through AI Grants India.