A small language model can be trained with far less data than a frontier model—but “small” does not mean data-light by default. The right target depends on whether you are pretraining from scratch, continuing pretraining an existing model, or fine-tuning it for a narrow task. It also depends on the model’s parameter count, tokenizer, language mix, domain, and quality requirements.
For most Indian builders, the practical question is not simply how many documents to collect. It is: how many clean, relevant tokens can the model learn from without memorising, overfitting, or losing performance across languages and use cases?
Start with the training method
Data requirements change dramatically depending on your approach.
- Fine-tuning: Usually needs the least data. A few hundred to several thousand high-quality examples may be enough for classification, extraction, style adaptation, or a tightly defined assistant workflow.
- Continued pretraining: Requires substantially more text, often millions to billions of tokens, because the model is adapting its vocabulary and domain knowledge.
- Pretraining from scratch: Requires the most data and the strongest infrastructure. You must teach the model language patterns, facts, formatting, and reasoning behaviour from the beginning.
If your goal is an Indian-language chatbot, document assistant, or speech-adjacent application, starting with an existing multilingual or Indic model is usually more efficient than training from zero. See the practical guidance on open-source small language models for Hindi before committing to a new base model.
Practical token ranges for small models
Parameter count provides a useful starting point, but it is not a complete data plan. For a decoder-only model trained from scratch, an initial planning range could look like this:
- 1–10 million parameters: 10 million to 500 million clean tokens for experimentation and narrow domains.
- 10–100 million parameters: 100 million to several billion tokens, depending on the target language and desired generality.
- 100–300 million parameters: Often benefits from billions of tokens, especially for multilingual or general-purpose use.
These are planning ranges, not guarantees. A 20-million-parameter model trained on 500 million excellent, domain-relevant tokens may outperform a larger model trained on duplicated, noisy web text. Conversely, a multilingual model may need much more data because its token budget is divided across languages.
A useful rule is to count tokens, not sentences. One sentence can contain a few tokens or several hundred, depending on the language, script, punctuation, and tokenizer. Indian languages can also be inefficiently tokenised by models built primarily for English, increasing sequence length and training cost. Measure tokenisation on representative Hindi, Tamil, Bengali, Marathi, or code-mixed samples before estimating the dataset size.
How much data does fine-tuning need?
For a focused application, fine-tuning is usually the sensible first experiment.
- Classification or intent detection: 500–5,000 labelled examples can establish a useful baseline.
- Information extraction: 1,000–10,000 carefully annotated examples may be sufficient, depending on schema complexity.
- Instruction tuning: 2,000–50,000 diverse, high-quality instruction-response pairs can produce meaningful behavioural improvements.
- Specialised language or terminology adaptation: Start with tens of millions of domain tokens for continued pretraining, then fine-tune on task examples.
Quality matters more than a large example count. Ten thousand near-duplicate prompts teach less than two thousand examples covering realistic inputs, edge cases, refusals, spelling variation, code mixing, and regional terminology. For medical, legal, financial, or public-service applications, include expert review and traceable source material rather than relying on synthetic examples alone.
Read best practices for fine-tuning LLMs on custom data for decisions around instruction format, validation splits, and avoiding contamination.
A better way to estimate your dataset
Use a staged data plan instead of collecting everything upfront.
1. Define the deployment task
Write down the inputs, expected outputs, supported languages, latency target, and acceptable error rate. A model for customer-support routing needs a different dataset from a generative assistant for government documents.
2. Build a clean pilot set
Create a small, representative corpus first. Include production-like queries, not only polished benchmark text. For an Indian deployment, test spelling variants, transliterated text, mixed English, local names, abbreviations, and low-bandwidth or mobile-typed inputs.
3. Train a baseline
Run a small model with a fixed token budget. Record training loss, validation loss, task metrics, inference latency, and failure categories. This gives you evidence about whether more data, better labels, a different tokenizer, or a larger model is the real bottleneck.
4. Scale based on learning curves
Add data in controlled increments—such as 2x or 4x—and measure whether validation performance continues to improve. Stop expanding the corpus when additional data produces negligible gains or introduces quality regressions.
5. Preserve a clean evaluation set
Keep a held-out test set that is never used for training, deduplication decisions, or prompt development. For high-stakes systems, maintain a separate challenge set reviewed by domain experts.
Data quality controls that matter
Before training, create a reproducible pipeline for:
- Deduplication: Remove exact and near-duplicate documents, boilerplate, repeated posts, and copied answers.
- Language identification: Detect the actual language and script, especially in code-mixed and transliterated content.
- PII removal: Redact phone numbers, Aadhaar-like identifiers, addresses, account details, and other sensitive information.
- Source filtering: Exclude malware, spam, scraped navigation text, low-quality machine translations, and unlicensed material.
- Balance checks: Track token counts by language, domain, source, and document type so one dominant source does not define the model.
- Versioning: Store dataset manifests, filtering rules, hashes, and licence information for every training run.
For systems where incorrect information creates material harm, ordinary cleaning is not enough. Plan a verification layer using principles from data veracity infrastructure for high-stakes AI. Medical projects should also account for ICMR-compliant medical AI data verification in India.
Indian-language considerations
Indic model development needs deliberate coverage. Public web data is uneven across languages, and high-resource English content can overwhelm smaller language corpora during multilingual training. Set explicit token or sampling targets for each supported language, then evaluate each language separately rather than reporting only an aggregate score.
Include regional vocabulary, formal and conversational registers, script variation, transliteration, and code-mixing. For example, a Hindi support assistant may need Devanagari, Romanised Hindi, and English product terms. A dataset that contains only standard newspaper Hindi will not represent how users type on WhatsApp or mobile apps.
For low-resource languages, a smaller but carefully curated corpus can be more valuable than indiscriminate scraping. The guide to low-resource Indic natural language processing covers collection, annotation, transfer learning, and evaluation choices.
Cost, compute, and stopping criteria
Data is only one part of the budget. Storage, tokenisation, GPU time, annotation, evaluation, and data governance can dominate a small project’s cost. Run a pilot on a modest corpus before purchasing extended compute. Track tokens processed per rupee, validation improvement per training hour, and inference quality at the target model size.
Stop when the model meets the product threshold—not when the dataset reaches an arbitrary number. A compact model with strong retrieval, evaluation, and guardrails may be more useful than a larger model trained on uncontrolled data. For many Indian startups, the winning architecture is a capable open model plus focused fine-tuning and retrieval, not full pretraining.
Bottom line
There is no universal sentence count for training a small language model. For fine-tuning, start with hundreds to tens of thousands of high-quality examples. For continued pretraining, plan in millions or billions of tokens. For pretraining from scratch, begin with a measured pilot and expect the data requirement to grow with model size, language count, and generality.
The best dataset is representative, legally usable, deduplicated, privacy-aware, and evaluated against real user behaviour. Build a baseline, measure learning curves, and expand only when the evidence shows that more data—not better data or a better training method—is the constraint.