Small language models are becoming a practical choice for teams that need reliable domain performance without the cost, latency, and infrastructure burden of a frontier model. For an Indian startup, hospital network, lender, public-service platform, or regional-language product, a compact model can be easier to deploy, audit, and operate close to users.
The objective is not to build the smallest model possible. It is to build the smallest model that meets a defined quality, latency, privacy, and maintenance target.
Start with a narrow, measurable job
“Build a domain model” is not a sufficient product brief. Begin with one workflow and define what success means. Common starting points include:
- Extracting fields from invoices, claims, contracts, or application forms
- Classifying customer messages and routing them to the right team
- Summarising clinical, legal, or operational documents
- Answering questions over a controlled knowledge base
- Translating or normalising Indian-language text
- Drafting responses that a trained employee reviews before sending
Specify the users, input formats, acceptable response time, languages, and consequences of an error. A model for internal document search can tolerate a different error profile from one that influences a credit decision or clinical workflow.
For multilingual products, language coverage must be explicit. Hindi, Tamil, Bengali, Marathi, and other Indian languages differ in script, morphology, code-mixing, spelling variation, and available training data. The guide to low-resource Indic natural language processing is useful when estimating data and evaluation requirements for these settings.
Decide whether you need a model, fine-tuning, or retrieval
Many projects should not begin with training. Compare three approaches:
- Prompting and structured output: suitable for a quick baseline or low-volume workflow.
- Retrieval-augmented generation (RAG): useful when answers must reflect changing policies, product catalogues, regulations, or internal documents. The model retrieves relevant passages instead of memorising everything.
- Fine-tuning or continued pre-training: worthwhile when the model must consistently follow a domain style, recognise specialised patterns, classify local terminology, or operate in an underrepresented language.
A compact model paired with good retrieval can outperform a larger model given poor documents and vague instructions. Build a baseline with an existing open-weight model before committing to a training pipeline.
For Hindi and regional-language work, compare available open models using your own test set rather than relying only on published benchmarks. The resources on open-source small language models for Hindi and fine-tuning Llama for Indian regional languages provide relevant starting points.
Build a trustworthy domain dataset
Data quality usually matters more than adding another layer or increasing parameter count. Create a data inventory covering source, licence, language, date, sensitivity, and intended use. Potential sources include:
- Public regulatory documents, manuals, reports, and government publications
- De-identified business records and support tickets
- Expert-written examples and difficult edge cases
- Opt-in user interactions, with consent and retention controls
- Synthetic examples used only to supplement, not replace, authentic data
Remove duplicates and near-duplicates, corrupted text, irrelevant boilerplate, and evaluation examples that may have leaked into training. Preserve domain expressions, abbreviations, spelling variants, and code-mixed language when they occur in real use.
For supervised tasks, write annotation guidelines before labelling. Define ambiguous cases, escalation rules, and the difference between “unknown” and “not applicable”. Have domain experts review a sample of annotations and measure agreement. In healthcare, finance, legal services, and public services, de-identification and access control are part of model development—not post-launch administration.
Keep a held-out test set that reflects production conditions. Split by customer, document, time period, or case—not just random rows—so repeated templates do not make performance look better than it is.
Choose an efficient model and training method
Select the architecture based on the task:
- Encoder models are strong candidates for classification, search, entity extraction, and reranking.
- Decoder models support generation, rewriting, structured extraction, and conversational interfaces.
- Smaller instruction-tuned models are useful where the application needs controlled responses.
- Embedding models may be sufficient when the main requirement is semantic search or clustering.
Use transfer learning wherever possible. Parameter-efficient methods such as LoRA and other adapter techniques reduce memory use and allow multiple domain adapters to share one base model. Quantisation can lower inference cost and make local or edge deployment practical, but test its effect on Indian scripts, long context, and structured output rather than assuming quality will remain unchanged.
Training from scratch is justified only when you have substantial, legally usable domain or language data and a reason existing tokenisation or representations are inadequate. For most product teams, continued pre-training or supervised fine-tuning is a more defensible investment.
Track experiments systematically: base checkpoint, dataset version, tokenizer, context length, learning rate, adapter settings, quantisation method, hardware, and evaluation results. Reproducibility matters when a model will support regulated or high-impact decisions.
Evaluate usefulness, safety, and cost
Do not rely on one score. Build a task-specific evaluation suite containing normal examples, rare terminology, noisy input, multilingual and code-mixed queries, adversarial prompts, and cases where the correct response is to abstain.
Use metrics appropriate to the task:
- Precision, recall, and F1 for classification and extraction
- Exact match or field-level accuracy for structured outputs
- Retrieval recall and ranking quality for search and RAG
- ROUGE or similar measures only as supporting evidence for summarisation
- Human ratings for factuality, completeness, tone, and usefulness
- Latency, memory use, throughput, and cost per request for operations
Test factuality separately from fluency. A polished answer containing invented policy, dosage, eligibility, or financial information is a failure. Add confidence thresholds, citations, constrained schemas, human review, or refusal behaviour where the risk warrants it.
Deploy for Indian production conditions
A model that works in a notebook may fail when exposed to unstable networks, low-end devices, variable document quality, or sudden traffic. Decide whether inference belongs in a managed cloud service, a private environment, or on-device. Smaller models can support lower latency and data-local processing, especially for sensitive workflows.
Expose the model through a versioned API with authentication, rate limits, logging, timeouts, and fallback behaviour. Store prompts and outputs only when the data policy permits it, and redact personal information from observability systems. Monitor drift by language, customer segment, document type, and confidence—not merely aggregate accuracy.
If the model handles voice, invoices, or customer conversations, errors may originate upstream in speech recognition, OCR, or formatting. The model should be evaluated as part of the complete pipeline. For example, teams building multilingual interfaces can also examine open-source vision-language models for Indian languages when documents or images are part of the workflow.
Create a maintenance and governance plan
Domain knowledge changes. Schedule dataset refreshes, regression tests, security reviews, and model rollback procedures. Maintain a model card that records intended use, limitations, languages, training data categories, known failure modes, and evaluation results.
For high-impact applications, define who approves training data, who can change prompts or adapters, who investigates incidents, and when a human must override the system. Obtain legal review for personal data, copyrighted material, sector-specific obligations, and cross-border processing. India-focused products should align their data practices with applicable privacy and sectoral requirements rather than treating compliance as a generic checkbox.
A practical build sequence
A sensible 2026 delivery plan is:
1. Define one workflow, user group, and measurable quality threshold.
2. Establish a prompt-only, retrieval, or conventional ML baseline.
3. Assemble and review a representative dataset with a clean test split.
4. Benchmark two or three compact base models on the same tasks.
5. Fine-tune with adapters only if the baseline cannot meet requirements.
6. Test quality, safety, latency, memory, and cost under realistic load.
7. Launch to a limited group with human review and clear feedback capture.
8. Monitor failures, refresh data, and promote new versions through regression tests.
The strongest small language model projects are disciplined product systems, not merely model-training exercises. Narrow scope, representative Indian-language data, transparent evaluation, and operational controls will usually create more value than parameter count alone.
Apply for AI Grants India
Indian founders building domain-specific language technology can explore AI Grants India for funding opportunities and support. A clear problem definition, credible dataset plan, measurable evaluation framework, and responsible deployment strategy will strengthen any grant application.