Sindhi is spoken by millions across India and Pakistan, yet its digital representation remains uneven. A focused small language model can support search, keyboards, translation, education, accessibility, and local-language assistants without requiring the budget or infrastructure of a frontier model.
The most reliable approach in 2026 is not to train a large model from scratch. Start with a strong open small language model, improve its Sindhi data and tokenizer coverage, then fine-tune and evaluate it for a defined task. This produces a useful system faster and makes quality easier to measure.
Define the job before choosing the model
“An AI model for Sindhi” is too broad to guide engineering decisions. Define one primary outcome first:
- Next-token prediction or text completion
- Sindhi question answering over a trusted document collection
- Translation between Sindhi, Hindi, English, or Urdu
- Text classification, such as sentiment or topic tagging
- Spelling correction, transliteration, or keyboard suggestions
- A conversational assistant for a narrow domain
A small model optimised for classification may need only a modest encoder, while a generative assistant needs a decoder model and instruction data. If your priority is low-resource Indic NLP, use the practical recommendations in this low-resource Indic NLP guide before committing to an architecture.
Build a lawful, representative Sindhi dataset
Data quality will matter more than adding another few million noisy tokens. Assemble sources that reflect how people actually write Sindhi, while documenting permission and provenance for every collection.
Useful sources include:
- Public-domain books, newspapers, dictionaries, and educational material
- Licensed digital archives and government publications
- Open subtitles, parallel corpora, and language-learning resources
- Volunteer-contributed text collected through a clear consent process
- Carefully selected public web content, subject to site terms and copyright
Do not scrape private messages or assume that publicly viewable text is automatically reusable for training. Keep a dataset card containing source, licence, date, dialect, script, cleaning steps, and known gaps. Remove personal information, account handles, phone numbers, and copied credentials before training.
Sindhi presents several coverage challenges. Indian Sindhi is commonly written in Devanagari, while Perso-Arabic Sindhi is widely used in Pakistan and in diaspora communities. Your project should state which script and dialects it supports. If both scripts matter, retain a script label and create balanced splits rather than allowing the larger corpus to dominate.
Clean text without destroying linguistic information
Aggressive cleaning can make a corpus look tidy while removing the signals a model needs. Preserve punctuation, sentence boundaries, diacritics where meaningful, and legitimate code-switching. Normalise only known variants, such as inconsistent Unicode forms, spacing, zero-width characters, and duplicated punctuation.
A practical preprocessing pipeline should:
1. Apply Unicode normalisation and record every transformation.
2. Detect the script and language of each document.
3. Remove boilerplate, navigation text, duplicate pages, and malformed OCR.
4. Deduplicate exact and near-duplicate documents.
5. Filter extremely short, corrupted, or machine-generated passages.
6. Split documents into train, validation, and test sets by source or document—not random lines.
Keep a clean, immutable original corpus and generate processed versions with reproducible scripts. For OCR-heavy books, sample pages manually and calculate character error rates before investing in large-scale processing.
Choose a tokenizer and base model
For a first prototype, fine-tune an open small language model rather than training all parameters from zero. Models in the roughly 100-million to 1-billion parameter range can be practical on a single modern GPU, especially with parameter-efficient methods and quantisation. Compare candidates for licence, context length, Indic-script coverage, inference cost, and community support—not just parameter count.
Inspect the tokenizer directly. Encode representative Sindhi sentences in both supported scripts and measure token count. Poor coverage creates long sequences, raises memory use, and weakens learning. If Sindhi is split into excessive fragments, consider training a SentencePiece or BPE tokenizer on a balanced multilingual corpus, then pretraining or continued-pretraining a compatible model. Changing a tokenizer after full fine-tuning is costly, so test this decision early.
For a useful comparison, study the workflow in this guide to open-source small language models for Hindi. Hindi is not Sindhi, but the lessons on Indic data, tokenizer checks, and evaluation transfer well. If you already have a compatible Llama-family checkpoint, fine-tuning Llama for Indian regional languages offers a practical starting point.
Train efficiently
Begin with continued pretraining on unlabeled Sindhi text using a causal language-modelling objective. This teaches vocabulary, spelling, syntax, and local style. Use a low learning rate, packed sequences, gradient accumulation, mixed precision, and checkpointing. Track training and validation loss separately; a falling training loss with rising validation loss indicates overfitting or leakage.
Then create a smaller instruction dataset for the target use case. Examples should contain natural Sindhi prompts and answers, with script and dialect labels where relevant. Include refusal examples for unsafe or unsupported requests, and avoid presenting fabricated cultural, medical, or legal information as fact.
Parameter-efficient fine-tuning methods such as LoRA or QLoRA can reduce hardware requirements. Keep a held-out evaluation set untouched, save configuration files with every checkpoint, and record the exact base model, tokenizer, data version, random seed, and hardware. A reproducible experiment is more valuable than an unexplained “best” score.
Evaluate with metrics and native speakers
Perplexity is useful for tracking language-model training, but it does not establish that a Sindhi assistant is helpful. Build an evaluation set across scripts, dialects, domains, and sentence lengths. Include naturally written examples, not only translated English prompts.
Measure:
- Character and word-level error for spelling or correction tasks
- BLEU, chrF, or COMET-style metrics for translation, alongside human review
- Accuracy, macro-F1, and calibration for classification
- Exactness, citation quality, and refusal behaviour for question answering
- Script fidelity, grammaticality, factuality, and relevance for generation
Recruit Sindhi speakers from the communities you intend to serve. Ask reviewers to rate fluency, meaning preservation, dialect appropriateness, harmful stereotypes, and whether the answer is genuinely useful. Report results separately for Devanagari and Perso-Arabic Sindhi instead of hiding weak performance behind one average score.
Deploy a narrow, observable service
Expose the model through a small API with authentication, rate limits, input-length limits, logging controls, and an explicit privacy policy. For CPU or mobile use, export a quantised model and benchmark latency, memory, and quality together. The techniques in this mobile model optimisation guide are relevant when the model must run on low-cost devices or unreliable connections.
Add retrieval for factual or changing content rather than expecting the model to memorise everything. Store source documents in Sindhi and return citations where possible. Monitor failed queries, script imbalance, hallucinations, and abusive inputs—but minimise retained personal data. Publish a model card that states intended use, known limitations, training sources, licence, evaluation results, and unsupported dialects.
A realistic first milestone
A strong initial project might deliver a 300–500 million parameter model or adapter, trained on a carefully documented Sindhi corpus, with support for one or two clearly defined tasks. In four to eight weeks, a small team can often produce a baseline, tokenizer analysis, evaluation set, and demo if data access is already available. The next improvement should come from better coverage and human feedback, not automatically from a larger checkpoint.
For India-focused builders, this work can also become a grant-ready project: define the community served, publish measurable language-access outcomes, and show how the system will remain available after the pilot. Sindhi language technology is most valuable when speakers, educators, linguists, and engineers shape the dataset and evaluation together.