0tokens

Apply for AI Grants India

Financial support for innovators building the future of AI in India.

Apply now

Chat · how to create a small language model for santali

How to Create a Small Language Model for Santali

  1. aigi

    Santali is spoken across Jharkhand, West Bengal, Odisha, Assam and neighbouring regions, and is one of India’s constitutionally recognised languages. Yet digital resources remain limited compared with Hindi, Bengali or English. A small, carefully scoped language model can support keyboards, search, translation, education, speech interfaces and community archives without requiring the budget or infrastructure of a frontier model.

    The most important design choice is not model size. It is data quality, script coverage and community ownership. This guide presents a practical 2026 workflow for building a compact Santali model, especially for teams working with limited compute.

    Define the task before choosing a model

    “Build a Santali model” can mean several different projects. Start with one measurable use case:

    • Next-token prediction for text generation or autocomplete
    • Text classification for moderation, topic labelling or intent detection
    • Translation between Santali and an Indian or international language
    • Retrieval and question answering over approved Santali documents
    • Spelling assistance, transliteration or keyboard suggestions

    A model for autocomplete needs low latency and reliable short-form predictions. A translation system needs aligned sentence pairs. A question-answering assistant may need a smaller language model combined with search rather than a larger generator. Teams new to low-resource NLP should review this low-resource Indic NLP builder’s guide before committing to architecture or infrastructure.

    Define success in user terms: suggestions accepted by speakers, translation adequacy, classification F1, response latency, memory use and harmful-error rate. Perplexity is useful during training but should not be your only metric.

    Build a lawful, representative Santali corpus

    Santali data may appear in Ol Chiki, Devanagari, Bengali or Odia scripts, alongside Romanised text. Treat each script as a meaningful part of the dataset rather than silently converting everything into one representation. Record the script, source, date, licence and consent status for every document.

    Useful sources include:

    • Public-domain books, newspapers and educational materials
    • Government publications with clear reuse terms
    • Open datasets and community repositories
    • Voluntary contributions from fluent speakers
    • Transcribed and consented conversations, stories and oral histories
    • Parallel material created specifically for translation training

    Do not scrape private groups, copyrighted books or social media and assume that public visibility equals permission. For community-contributed material, explain the purpose, retention policy, risks and withdrawal process in accessible language. Credit contributors and agree in advance whether the dataset and resulting model may be redistributed commercially.

    Create a data card describing collection methods, dialect coverage, scripts, licences, known gaps and prohibited uses. Remove phone numbers, addresses and other personal information. Keep a held-out evaluation set private or access-controlled so benchmark results remain meaningful.

    Normalise carefully without erasing language features

    Santali preprocessing requires more than lowercasing. Build a Unicode-aware pipeline that can:

    • Normalise equivalent Unicode sequences
    • Detect Ol Chiki, Bengali, Devanagari, Odia and Latin text
    • Preserve meaningful punctuation, spacing and sentence boundaries
    • Handle spelling variation without rewriting speaker identity
    • Separate code-mixed segments for analysis
    • Deduplicate near-identical documents and repeated web content

    Avoid aggressive stemming or lemmatisation unless linguists validate the rules. These steps can damage morphology, names and dialectal forms. Keep both the original text and a normalised training view, with reversible transformations wherever possible.

    Split data by document, author or source—not random lines—to prevent near-duplicates from leaking into validation and test sets. Maintain separate test slices for each script, region, genre and code-mixing pattern. This will reveal whether the model works for real users rather than only for the dominant source.

    Choose tokenisation and a compact architecture

    For a new project, compare three baselines:

    1. A word- or character-level n-gram model for autocomplete and sanity checks.
    2. A small Transformer trained from scratch when you have a sizeable, clean corpus.
    3. An existing multilingual or Indic checkpoint adapted through continued pretraining or fine-tuning.

    A subword tokenizer is usually a practical starting point, but its vocabulary must be tested on Ol Chiki and other scripts. Inspect average token length, unknown-token rates and fertility—the number of tokens required per word or sentence. A tokenizer that fragments Santali heavily can make a nominally small model expensive and less useful.

    Train a custom SentencePiece or byte-level tokenizer on a balanced sample. Reserve vocabulary capacity for scripts and common morphemes rather than allowing high-volume English or Hindi text to dominate. Compare a shared multilingual vocabulary with a Santali-focused vocabulary on held-out text.

    For efficient experimentation, begin with a compact Transformer—roughly tens to a few hundred million parameters depending on corpus size and compute. More parameters cannot compensate for duplicated, unlicensed or script-imbalanced data. If adapting an existing model, fine-tuning Llama for Indian regional languages offers relevant design considerations, but verify the checkpoint’s licence and actual Santali coverage first.

    Train efficiently on limited compute

    Prepare train, validation and test splits before launching expensive runs. Use mixed precision, gradient accumulation, sequence packing and checkpoint resumption where supported. Track:

    • Training and validation loss
    • Per-script perplexity
    • Tokenisation statistics
    • GPU memory, training time and energy use
    • Checkpoint size and inference latency

    Continued pretraining on carefully filtered Santali text may outperform full training when an appropriate base model exists. For supervised tasks, use parameter-efficient methods such as adapters or low-rank fine-tuning. Keep a small baseline and change one variable at a time: tokenizer, data mixture, context length or learning rate.

    Do not train on test examples, synthetic text generated by the same model, or translated material without marking it clearly. Synthetic augmentation can help with intent classification or formatting, but native-speaker review is essential for grammar, meaning and cultural safety.

    Evaluate with speakers, not only scores

    Create a benchmark with native or highly proficient Santali speakers from relevant regions and scripts. Include natural prompts, spelling variation, code-mixing, names, cultural references and adversarial requests. Ask reviewers to score:

    • Fluency and grammatical acceptability
    • Faithfulness and omission
    • Script correctness
    • Helpfulness for the target task
    • Stereotypes, unsafe content and fabricated facts

    Report results by slice instead of publishing one headline number. Compare against a simple n-gram baseline and an existing multilingual system. For translation, use human adequacy and preference ratings alongside automated metrics. For generation, measure repetition, memorisation and whether the model reproduces private or copyrighted material.

    Use an error log that records the prompt, output, category, script, dialect context and reviewer confidence. This turns evaluation into a prioritised training plan.

    Deploy for Indian users

    A compact model can run through an API, on a local server or on Android-class hardware. Quantisation and distillation can reduce memory and latency, but validate quality after compression. This AI model optimisation guide for mobile devices is useful when building offline or low-connectivity experiences.

    For keyboards and education tools, local inference may offer stronger privacy and lower operating costs. For larger models, use a simple API with rate limits, logging controls and clear data-retention policies. Design for intermittent connectivity, inexpensive devices and multilingual input. Include script selection, correction controls and a way for speakers to report errors.

    Publish responsibly and iterate

    Release a model card, dataset card, evaluation protocol and known limitations. Document dialect and script coverage, intended uses, excluded uses, licences, safety testing and contact details for takedown or correction requests. Do not present a model as representing all Santali speakers if the corpus is geographically or socially narrow.

    A credible first release may be a tokenizer, benchmark, autocomplete model or curated dataset rather than a general chatbot. Open-source small language models for Hindi provide a useful comparison for packaging and evaluation, but Santali needs its own data governance and speaker-led review. If your project has a clear public benefit, infrastructure plan and responsible data process, consider applying to AI Grants India for support.

    FAQ

    Is Ol Chiki required?

    No. Santali is also written in Bengali, Devanagari, Odia and Latin scripts. If your users work in Ol Chiki, however, it should receive deliberate tokenizer, training and evaluation coverage rather than being treated as an afterthought.

    How much data is needed?

    There is no universal threshold. A focused classifier or autocomplete system can start with a modest, well-labelled corpus; a general generative model needs substantially more diverse text. Begin with a baseline and publish learning curves showing how quality changes as data grows.

    Should I train from scratch?

    Only if you have sufficient clean data, compute and engineering capacity. Otherwise, adapt a compatible multilingual or Indic checkpoint, then measure whether its tokenizer and representations actually support Santali.

    What is the best first project?

    A script-aware tokenizer, autocomplete model, parallel-corpus benchmark or intent classifier is often more achievable than a general chatbot. Each can produce useful evidence and reusable infrastructure for later work.

    Last updated 23 September 2026

AIGI may be inaccurate. Replies seeded from the guide above.