0tokens

Apply for AI Grants India

Financial support for innovators building the future of AI in India.

Apply now

Chat · how to build a small language model from scratch

How to Build a Small Language Model from Scratch

  1. aigi

    What you are actually building

    A small language model is a next-token prediction system trained on a focused text corpus. It is not a miniature version of a frontier model in every respect. The practical goal is usually narrower: learn a domain vocabulary, generate predictable text, classify or extract information through prompting, or run inference on affordable hardware.

    For most builders in 2026, the best first project is a decoder-only Transformer with roughly 10 million to 300 million parameters, a carefully selected corpus, and a measurable use case. Training from random weights teaches you the full pipeline. It does not automatically produce a production-grade assistant. If your goal is a useful product rather than a learning exercise, compare this route with adapting an existing open model through fine-tuning or retrieval.

    Language technology for Indian users adds practical constraints. English-only tokenisation can be inefficient for Hindi, Tamil, Bengali, Marathi, and code-mixed text. A focused corpus and a tokenizer tested on Indic languages may deliver more value than simply increasing parameter count. For broader context, see this guide to low-resource Indic natural language processing.

    Define the target before writing code

    Write a one-sentence specification before collecting data:

    • Input: What will the model receive—plain text, chat turns, documents, or code?
    • Output: Should it complete text, answer questions, classify intent, or extract fields?
    • Language: Which Indian languages, scripts, transliteration systems, or code-mixed combinations matter?
    • Latency and hardware: Must it run on a laptop, an Indian cloud GPU, a CPU server, or an edge device?
    • Success metric: What result would justify the project?

    A model trained on government circulars may be useful for document completion but poor at general conversation. A model trained on customer-support transcripts may memorise sensitive information. Narrow scope improves evaluation, but it also makes data governance more important.

    Prepare a defensible dataset

    Data quality is usually the largest source of model quality. Start with public-domain, licensed, or organisation-owned text. Record the source, licence, language, date, and preprocessing decisions in a dataset manifest.

    A practical pipeline is:

    1. Ingest: Convert HTML, PDF, DOCX, and plain-text files into UTF-8 text while retaining source identifiers.
    2. Normalise: Standardise Unicode, whitespace, punctuation, and line endings. Do not blindly lowercase Indic scripts or remove diacritics.
    3. Filter: Remove navigation menus, boilerplate, corrupt files, extremely short records, and documents outside the target domain.
    4. Deduplicate: Use exact hashes and near-duplicate detection. Duplicates can inflate validation scores and cause memorisation.
    5. Remove sensitive data: Detect phone numbers, email addresses, Aadhaar-like identifiers, financial details, and private conversations. Establish an escalation process for uncertain cases.
    6. Split correctly: Create train, validation, and test sets by document or source—not by random lines—so related passages do not leak across splits.

    Keep a small, human-reviewed test set that never enters training. For Indian-language projects, annotate a sample for script, language, transliteration, and code mixing; automatic language identification is not always reliable on short text.

    Train a tokenizer that fits the corpus

    The tokenizer maps text to integer IDs. For a first implementation, use a byte-level BPE or SentencePiece-style subword tokenizer with special tokens such as beginning-of-sequence, end-of-sequence, padding, and unknown tokens.

    Choose vocabulary size based on corpus diversity and memory. A vocabulary of 16,000–64,000 tokens is a reasonable experimental range. Measure tokens per character or word across each target language. If Hindi or Tamil text expands into many fragments while English remains compact, the tokenizer is wasting context on the languages you care about.

    Reserve token IDs explicitly and save the tokenizer with every model checkpoint. A model and tokenizer are a single artefact; changing tokenisation after training invalidates comparisons.

    Implement a compact Transformer

    A decoder-only Transformer predicts the next token from all previous tokens. The core components are:

    • Token and positional embeddings
    • Causal self-attention, which prevents access to future tokens
    • Feed-forward layers
    • Residual connections and normalisation
    • A linear output head projecting hidden states to vocabulary logits

    A sensible starter configuration might use 6–12 layers, a hidden size of 384–768, 6–12 attention heads, and a context window of 512–2,048 tokens. These are starting points, not universal settings. Keep the model small enough that you can run several experiments rather than one expensive training job.

    Use PyTorch with mixed precision where supported. Implement a causal mask and verify tensor shapes with tiny synthetic inputs before training. A useful smoke test is to train on a few repeated examples until the model nearly memorises them; failure usually indicates a bug in masking, labels, token shifts, or the loss calculation.

    Build the training loop

    For each sequence, create input tokens x = tokens[:-1] and target tokens y = tokens[1:]. The standard objective is cross-entropy over the target token at every position:

    • Optimiser: AdamW
    • Learning-rate schedule: warm-up followed by cosine decay
    • Gradient clipping: useful for preventing unstable updates
    • Gradient accumulation: simulates larger batches on limited VRAM
    • Checkpoints: save model, optimiser, scheduler, tokenizer, configuration, and training step
    • Logging: record loss, learning rate, throughput, memory use, and validation loss

    Do not choose training duration by a fixed “10–50 epochs” rule. Language-model training is better described by tokens processed. Monitor training and validation loss: falling training loss with rising validation loss indicates overfitting. Stop, reduce the learning rate, add data, or reduce model capacity as appropriate.

    For compute-constrained teams, begin with a small corpus and a single GPU or CPU experiment. Once the pipeline is correct, scale data and model size. Distributed training adds operational complexity; use it only after profiling shows that a single device is the bottleneck. If your product needs multiple coordinated model or tool calls, that is a separate systems problem covered in building distributed systems with AI agents.

    Evaluate more than perplexity

    Perplexity is useful for comparing checkpoints on the same held-out distribution, but it does not prove that users will find the model helpful. Build an evaluation set that reflects real tasks:

    • Next-token loss by language and domain
    • Exact-match or F1 scores for extraction tasks
    • Instruction-following tests with fixed prompts
    • Factuality and refusal checks
    • Memorisation tests using canary strings and sensitive examples
    • Human ratings for relevance, fluency, and harmful or stereotyped output
    • Latency, peak memory, and cost per generated token

    Evaluate English, each target Indic language, transliterated input, and code-mixed prompts separately. Include difficult cases such as spelling variation, noisy OCR, and long context. Save prompts and model settings so every release can be compared fairly.

    Fine-tune, compress, and deploy

    Pretraining from scratch teaches general continuation behaviour. To make the model follow instructions, prepare high-quality prompt-response examples and apply supervised fine-tuning. Keep a separate test set and check that fine-tuning has not damaged the base model’s language coverage.

    For deployment, export a stable checkpoint and serve it behind an authenticated API. Quantisation to 8-bit or 4-bit precision can reduce memory and improve affordability, but test quality and Indic tokenisation after quantisation. Add request limits, prompt and output logging with privacy controls, timeouts, structured error responses, and a model version in every response. A private deployment may be preferable for legal, health, or financial material; the considerations in how to build a private AI chatbot for lawyers are relevant even outside legal workflows.

    If the model is part of a voice product, separate language-model latency from speech recognition and synthesis latency. See how to build a voice agent for the broader architecture.

    A practical eight-week build plan

    • Week 1: Define the task, licence requirements, languages, and evaluation criteria.
    • Weeks 2–3: Collect, clean, deduplicate, and document the corpus.
    • Week 4: Train and inspect the tokenizer; create leakage-safe splits.
    • Weeks 5–6: Implement the Transformer, smoke tests, and baseline training run.
    • Week 7: Run evaluations, error analysis, and targeted fine-tuning.
    • Week 8: Quantise, package, serve, monitor, and document limitations.

    The most valuable deliverable is not only a checkpoint. It is a reproducible repository containing data documentation, preprocessing code, tokenizer files, configuration, evaluation prompts, licences, and deployment instructions. That package makes the work auditable and easier for another Indian builder to extend.

    Common mistakes to avoid

    • Training on scraped data without checking rights or personal information
    • Randomly splitting passages and reporting inflated validation performance
    • Increasing parameter count before fixing data and tokenisation
    • Treating perplexity as a complete measure of usefulness
    • Changing the tokenizer between experiments
    • Deploying without rate limits, privacy controls, and rollback checkpoints
    • Calling a model “from scratch” while omitting the data, code, or evaluation needed to reproduce it

    A small model can be an excellent research and product asset when it solves a narrow problem reliably. Start with a transparent baseline, measure it against a real Indian-language workflow, and scale only when the evidence supports the cost.

    Last updated 23 September 2026

AIGI may be inaccurate. Replies seeded from the guide above.