Indian-language small language models (SLMs) do not fail because they are small. They fail when their data is narrow, duplicated, poorly licensed, or disconnected from how people actually communicate. A model trained mostly on formal Hindi, English-heavy code-mixed text, or clean news articles may perform well on a benchmark and still struggle with a farmer’s voice note, a government form, or a regional spelling variant.
The right dataset strategy combines coverage, quality, permission, and task relevance. This guide explains what to collect, what to measure, and how to build a defensible data pipeline for Indian-language SLMs in 2026.
Start with a language and deployment specification
Before collecting data, define the model’s intended users and operating conditions. “Indian languages” is not a single dataset category: Hindi in Devanagari, Hindi typed in Roman script, Marathi voice queries, and Tamil-English customer support require different data.
Write down:
- Target languages, scripts, dialects, and regional varieties.
- Expected input modes: typed text, speech, OCR, or mixed input.
- Use cases: classification, retrieval, translation, summarisation, tutoring, or conversation.
- Deployment constraints such as mobile inference, intermittent connectivity, and memory limits.
- Safety requirements for health, finance, education, government, or legal applications.
This planning is especially important for teams working in low-resource Indic NLP, where a small but carefully designed corpus can be more useful than a large, noisy crawl. See the low-resource Indic natural language processing guide for an overview of language-specific constraints.
The core dataset layers
1. Clean monolingual text
A pre-training corpus should contain varied, representative text rather than the maximum possible number of tokens. Useful sources include:
- Government information and public-service content, after checking reuse rights.
- Wikipedia and other openly licensed encyclopaedic material.
- Public-domain books and literary archives.
- News, magazines, blogs, and educational resources with explicit licensing.
- Product documentation, FAQs, and support content supplied by partners.
- Carefully filtered web text in both native scripts and Romanised forms.
Deduplicate at document and paragraph level. Remove boilerplate, navigation menus, malware pages, excessive advertisements, machine-generated spam, and repeated translations. Keep metadata such as language, script, source, licence, date, and region so that problematic slices can be removed later.
Do not treat token count as the primary success metric. Track the proportion of each language, script, domain, dialect, and time period. A corpus with 100 million repetitive sentences may be less valuable than 10 million diverse, well-documented sentences.
2. Code-mixed and transliterated text
Indian users frequently mix English with Hindi, Tamil, Telugu, Bengali, Marathi, and other languages. They also type Indic languages in Roman script. These forms should be collected deliberately, not treated as noise.
Include:
- Native-script and Romanised versions of common phrases.
- Multiple spellings and phonetic variants.
- Code-mixed customer queries and informal conversations collected with consent.
- Named entities, abbreviations, numbers, emojis, and regional expressions.
- Normalised and original forms, stored as separate fields.
Annotate the language or language span where possible. This supports language identification, transliteration, search, and robust downstream generation. Avoid aggressive normalisation: a model should learn variation, while a production system can apply a separate normalisation layer when needed.
3. Parallel and comparable corpora
Translation and cross-lingual retrieval require aligned examples between Indian languages and English, as well as between Indian languages themselves. Prioritise quality over indiscriminate alignment.
A useful parallel dataset records:
- Source and target text with sentence-level alignment.
- Translation direction and domain.
- Human, professional, community, or machine-generated origin.
- Quality ratings and known terminology.
- Script, dialect, and regional context.
Include comparable documents—such as the same policy topic written independently in Hindi and Kannada—even when exact translations are unavailable. These help multilingual retrieval and representation learning, but should not be labelled as parallel data. Public collections such as OPUS can be starting points, subject to licence and quality review.
4. Speech, phonetic, and spoken-language data
Voice interfaces need more than transcripts. Build datasets with recordings, accurate transcripts, speaker metadata, and environmental labels. Cover phone microphones, noisy streets, homes, classrooms, and varied speaking speeds.
For each sample, capture where consent permits:
- Language, dialect, district, and speaker age range.
- Code-switching and disfluencies.
- Background noise and recording device.
- Word-level or utterance-level timestamps.
- Named entities, numbers, addresses, and domain terminology.
Do not infer sensitive attributes unnecessarily. Use clear consent forms in languages participants understand, specify retention and commercial use, and provide deletion mechanisms. Speech data is particularly valuable for builders developing voice agents for Indian businesses, but it also carries higher privacy and re-identification risks.
5. Instruction, conversation, and task datasets
Pre-training teaches language patterns; instruction data teaches the model how to respond. Create examples for the exact tasks users will perform:
- Question answering grounded in approved documents.
- Summarisation at different reading levels.
- Classification of intents, complaints, or documents.
- Translation and transliteration.
- Form filling and information extraction.
- Educational explanations and practice questions.
- Safe refusal and escalation for high-risk requests.
Each example should include the language, script, domain, user intent, expected answer, and—where relevant—the source evidence. Use trained annotators and native-language reviewers. For conversations, include realistic turn-taking, clarification, uncertainty, and corrections rather than only polished question-answer pairs.
6. Domain and safety data
A general corpus will not prepare a model for healthcare, law, finance, agriculture, or public services. Build domain-specific datasets with expert review and provenance. For high-risk uses, include examples where the correct response is to ask for missing information, cite a source, or refer the user to a qualified professional.
Safety sets should test privacy leakage, abusive content, caste and religious stereotyping, gender bias, political persuasion, misinformation, and harmful medical or financial advice. Balance harmful examples with safe, helpful alternatives so the model learns behaviour rather than merely memorising offensive text.
Evaluation data must be separate
Reserve a private test set before training. It should represent real deployment conditions, including dialects, code-mixing, spelling variation, low-quality OCR, noisy speech, and long-tail names. Never let benchmark or evaluation examples enter the training corpus.
Measure more than perplexity:
- Language identification and script accuracy.
- Translation adequacy and terminology accuracy.
- Factuality and citation correctness.
- Intent and entity extraction performance.
- Speech word error rate by language and dialect.
- Safety, refusal, and escalation behaviour.
- Latency, memory use, and performance on target devices.
Report results by language and slice. An average score can hide a model that works for Hindi but fails for smaller language communities.
Governance, licensing, and quality control
Maintain a dataset card for every source. Record owner, licence, collection method, consent basis, geography, dates, transformations, known gaps, and permitted uses. Do not scrape private groups or assume that publicly visible content is automatically suitable for training.
A practical quality workflow is:
1. Language and script identification.
2. Malware, spam, and personal-data filtering.
3. Exact and near-duplicate removal.
4. Human sampling by language and domain.
5. Licence and provenance review.
6. Toxicity, bias, and sensitive-content audits.
7. Versioned release with hashes and change logs.
For open projects, publish documentation and reproducible preprocessing wherever legally possible. Indian developers can also learn from Indian open-source AI projects and student-led efforts that make local-language tooling easier to inspect and improve.
A practical minimum viable dataset plan
A small team does not need to collect everything at once. Start with one or two deployment languages and build:
- A balanced monolingual corpus with documented sources.
- A code-mixed and transliterated slice.
- A modest, high-quality instruction set reviewed by native speakers.
- Speech or OCR data only if the product needs it.
- A domain glossary and retrieval test set.
- A private evaluation suite covering real user conditions.
Expand after measuring failures. If the model confuses names, collect named-entity examples; if it fails on Romanised input, improve transliteration coverage; if it gives unsafe medical answers, add expert-reviewed refusal and referral examples. This feedback loop is more efficient than indiscriminate data accumulation.
FAQ
What is the most important dataset for an Indian-language SLM?
There is no universal winner. Start with clean, diverse monolingual text for the target language, then add task-specific instruction, code-mixed, speech, or domain data based on the product.
Should Romanised Indian-language text be included?
Yes, when users type that way. Store it as a distinct, labelled slice so it improves robustness without obscuring native-script performance.
Can synthetic data replace human-created data?
Synthetic data can expand coverage and create controlled examples, but it should be labelled, filtered, and validated by native speakers. It should not replace authentic language variation or expert review.
How many languages should a small model support?
Support fewer languages well before adding many superficially. Language balance, tokenizer coverage, evaluation quality, and deployment needs should determine the scope.
For Indian founders building language products, the strongest dataset is not simply the largest. It is the one that is representative, traceable, legally usable, and tested against the situations your users actually face. Teams building education products can also review the AI tutor landscape for Indian competitive exams to see why domain, language, and safety evaluation must be designed together.