0tokens

Apply for AI Grants India

Financial support for innovators building the future of AI in India.

Apply now

Chat · building customizable large language models from scratch

Building Customizable Large Language Models from Scratch

  1. aigi

    Building a language model from randomly initialized weights is a serious engineering project, not a weekend fine-tuning exercise. It requires a defensible data strategy, tokenizer design, distributed training, evaluation, safety controls, and a deployment plan. For most teams, the right goal is not the largest possible model, but a smaller, auditable model that performs reliably on a defined set of tasks and languages.

    This guide explains when training from scratch makes sense, how to structure the work, and where India-specific constraints—Indic scripts, code-mixed text, limited high-quality data, and cost-sensitive inference—change the design.

    Decide whether “from scratch” is necessary

    Start by comparing three paths:

    • Prompting or retrieval-augmented generation: Best when the knowledge changes often and a capable base model already understands the task.
    • Parameter-efficient fine-tuning: LoRA or other adapters are usually the fastest route to domain style, terminology, and structured outputs.
    • Continued pre-training: Useful when a model has the right broad capabilities but weak coverage of a domain or language.
    • Training from scratch: Justified when you need control over the tokenizer, training corpus, license, architecture, privacy boundary, or language coverage.

    A new model may be warranted for Marathi-heavy customer support, code-mixed Hindi-English workflows, offline public-service applications, or research requiring transparent data provenance. Before committing, benchmark a few open models on a representative internal test set. If retrieval or an adapter closes most of the gap, spend the budget on data and product reliability instead.

    Teams building for India should also read the practical considerations in building AI apps for the next billion users in India, particularly around connectivity, device constraints, and user experience beyond English.

    Define the target before choosing the model

    Write a one-page model specification covering:

    • Primary tasks: generation, classification, extraction, translation, summarisation, or tool use.
    • Languages and scripts: include language, script, dialect, transliteration, and code-mixing requirements.
    • Context length: estimate the real input size rather than selecting a large window by default.
    • Latency and hardware: identify whether inference must run on a CPU, a single GPU, or a cloud cluster.
    • Quality thresholds: set task-level targets for factuality, instruction following, safety, and structured-output validity.
    • Data and licence constraints: record what may be used for training, evaluation, and commercial deployment.

    A 300-million-parameter model that runs affordably on local infrastructure can be more valuable than a multi-billion-parameter model that cannot meet latency or cost targets. Establish a baseline with existing models and define a clear “stop” condition for the scratch-training project.

    Build a defensible data pipeline

    Data quality is usually the largest determinant of performance. Collect data legally and document every source, licence, language label, filtering decision, and transformation. Do not treat scraped web text as automatically usable; it may contain copyrighted material, personal information, malware, duplicated pages, or machine-generated low-quality text.

    A robust pipeline should include:

    • Ingestion: store immutable raw data and version every processing stage.
    • Language identification: detect Indic languages, mixed-language samples, transliteration, and script variants rather than applying an English-only classifier.
    • Cleaning: remove boilerplate, spam, broken encodings, personal data, and unsafe content according to written policies.
    • Deduplication: use exact and fuzzy methods at document and passage level to prevent memorisation and inflated validation scores.
    • Quality scoring: rank text using source reliability, linguistic quality, length, formatting, and repetition signals.
    • Contamination checks: ensure validation and test examples do not appear in pre-training data.

    For low-resource languages, aggressive filtering can remove the very text you need. Combine automated filters with native-speaker review and maintain balanced sampling so high-resource English does not overwhelm Indic data. The methods in the builder’s guide to low-resource Indic NLP are particularly relevant here.

    Design the tokenizer for real users

    Tokenization affects cost, context capacity, and language quality. A tokenizer trained mainly on English may split Devanagari, Bengali, Tamil, or transliterated text into excessive fragments. That increases sequence length and can make multilingual training inefficient.

    Train and inspect a tokenizer on a representative mixture of all target languages, scripts, numbers, punctuation, code, and common domain terms. Measure fertility—the average number of tokens per word or character—by language. Check whether names, addresses, abbreviations, and mixed-script messages remain usable. Reserve special tokens for chat roles, documents, tool calls, and structured formats only when the training objective requires them.

    Keep a fixed tokenizer version once serious training begins. Changing it later changes the meaning of every embedding and invalidates comparisons.

    Choose a practical architecture and training scale

    For most text-generation projects, a decoder-only Transformer is the simplest starting point. Select depth, hidden size, attention configuration, context length, and vocabulary size together; do not copy a popular configuration without matching it to your compute and data.

    Use a staged plan:

    1. Prototype: train a tiny model on a small, representative corpus to validate the data loader, tokenizer, loss calculation, checkpointing, and evaluation.
    2. Pilot: train a small model long enough to compare data mixtures, sequence lengths, and architecture choices.
    3. Production run: scale only after the pilot shows predictable loss and useful downstream performance.

    PyTorch, FSDP, DeepSpeed, and comparable open-source tooling can support mixed precision, gradient accumulation, checkpoint sharding, and distributed data parallelism. Track tokens processed, tokens per second, GPU utilisation, memory, loss by data slice, and failed jobs. Checkpoint frequently and test restoration; an interrupted multi-GPU run should not erase weeks of work.

    If your model must coordinate tools or services, treat orchestration as a separate system problem. Patterns covered in building distributed systems with AI agents can inform queues, retries, observability, and permission boundaries, but they do not replace model evaluation.

    Train, adapt, and align in stages

    Pre-training teaches language patterns and broad knowledge. It does not automatically produce a helpful assistant. After pre-training, use carefully curated instruction data for supervised fine-tuning. Include realistic Indian names, addresses, dates, currencies, multilingual conversations, and domain terminology where appropriate.

    Preference optimisation can improve helpfulness and style, but only if preference data is consistent and the reward signal is not gaming superficial traits. For many teams, parameter-efficient adapters are safer and cheaper than modifying all weights. Keep domain adapters separate when customers, departments, or policies require independent updates.

    Protect sensitive data throughout training. Mask personal information, restrict access to raw corpora, encrypt artefacts, and define retention periods. A model can memorise rare strings even when aggregate metrics look strong.

    Evaluate capability, safety, and cost

    Perplexity is useful for training diagnostics but is not a product metric. Build a held-out evaluation suite that reflects actual use:

    • Exact-match or F1 scores for extraction and classification.
    • Translation quality by language pair, plus native-speaker review.
    • Factuality and citation correctness for knowledge tasks.
    • Instruction following, refusal quality, and prompt-injection resistance.
    • Robustness to spelling variation, code-mixing, transliteration, and noisy mobile input.
    • Latency, throughput, memory use, and cost per million tokens.

    Create slices by language, gendered names, geography, domain, and input difficulty. Compare against strong open baselines and a retrieval-augmented version. Red-team harmful content, privacy leakage, stereotypes, and unsafe tool calls before deployment. Publish a model card describing training data categories, known limitations, intended use, evaluation results, and licence terms.

    Deploy with an operations plan

    Quantisation, batching, caching, and smaller distilled variants can make inference viable on Indian cloud or edge infrastructure. Use a model server with authentication, rate limits, request logging that respects privacy, and configurable safety filters. Separate model weights from prompts and business rules so policies can change without retraining.

    Monitor quality in production using sampled, consented interactions, user feedback, latency, token usage, refusal rates, and language-specific failures. Establish rollback procedures for model, tokenizer, prompt, and data changes. For voice or multimodal products, pair the language model with specialised components; for example, building a voice agent with Whisper and ElevenLabs illustrates how speech layers sit around—not inside—the core language model.

    A realistic 2026 checklist

    Before starting a large run, confirm that you have:

    • A written reason not to use an existing model or adapter.
    • A licensed, versioned, deduplicated corpus with language balance.
    • A tokenizer benchmarked across target Indic languages and scripts.
    • A reproducible training configuration and tested checkpoint recovery.
    • Held-out evaluations, native-speaker reviewers, and safety tests.
    • A deployment budget, monitoring plan, and model update policy.

    The strongest custom LLM projects are disciplined rather than oversized. Build the smallest model that satisfies the requirement, invest in representative data and evaluation, and keep every decision reproducible. That approach gives Indian builders more control over language coverage, privacy, operating cost, and product behaviour than chasing scale alone.

    Last updated 23 September 2026

AIGI may be inaccurate. Replies seeded from the guide above.