Rajasthani is not a single uniform text stream. It covers related varieties such as Marwari, Mewari, Dhundhari, Shekhawati and Harauti, with substantial variation in vocabulary, spelling, pronunciation and script usage. A useful model therefore starts with a clear scope—not an assumption that one dataset represents every speaker.
For most independent builders, the best first product is not a general-purpose chatbot. Start with a measurable task such as text completion, spelling assistance, translation support, search, transcription post-correction or a retrieval assistant for Rajasthani documents. A narrow goal reduces data and compute requirements while making quality easier to test.
Define the use case and language scope
Write a one-page model specification before collecting data. Record:
- Target varieties: for example, Marwari-first, or a balanced mix of several varieties.
- Script: Devanagari, Romanised Rajasthani, or both. Do not silently merge scripts without tagging them.
- User context: education, public-service information, cultural archives, customer support or creative writing.
- Safety boundaries: avoid presenting the model as an authority for legal, medical or government advice without verification.
- Success metrics: text quality, dialect coverage, factuality, latency, memory use and cost per request.
A strong project can borrow methods from this low-resource Indic NLP builder’s guide, especially its emphasis on data quality, language identification and evaluation beyond a single score.
Build a lawful, representative dataset
Data volume matters, but usable and well-documented data matters more. Combine sources carefully:
- Public-domain or permissioned books, folk literature, newspapers and educational material.
- Opt-in contributions from native speakers, teachers, writers and language researchers.
- Licensed web text, with robots.txt, terms of use and copyright restrictions respected.
- Parallel Hindi–Rajasthani or English–Rajasthani material for translation and alignment tasks.
- Speech transcripts only when speaker consent and recording rights are documented.
Maintain a data card for every source: licence, creator, collection date, variety, script, domain and permitted uses. Remove personal information, private conversations, phone numbers and copied copyrighted text without permission. Keep a held-out test set from a separate source so that the model cannot succeed through memorisation.
Balance the corpus by variety and domain rather than allowing the easiest-to-find variety to dominate. Tag each document with metadata such as variety, script, domain and source_type. These tags support targeted evaluation later.
Clean and normalise without erasing linguistic detail
Rajasthani text may contain inconsistent spacing, punctuation, spellings, Devanagari variants and code-switching with Hindi or English. Build preprocessing as a reproducible pipeline, not a one-time spreadsheet operation.
Useful stages include:
1. Detect and remove boilerplate, duplicate pages, navigation text and corrupted files.
2. Run language and script identification at document or sentence level.
3. Deduplicate exact and near-identical passages to limit memorisation.
4. Standardise Unicode and normalise only known orthographic variants.
5. Preserve punctuation, numerals, named entities and meaningful code-switching.
6. Segment text into documents, paragraphs and sentences while retaining source metadata.
7. Create train, validation and test splits by source—not random lines from the same document.
Do not automatically convert every word to Hindi spelling or strip diacritics. Ask native speakers whether a proposed normalisation improves readability. If you support Romanised input, retain the original form and record the transliteration convention used.
Choose a practical model strategy
Training a transformer from scratch is rarely the best first move. In 2026, a more efficient path is to start with an open small language model that supports Devanagari, then adapt it with continued pretraining or supervised fine-tuning. Compare this approach with open-source small language models for Hindi, while checking each model’s licence and actual script coverage.
Select the strategy based on the task:
- Continued pretraining: best when you have a substantial, clean Rajasthani corpus and want better next-token prediction.
- Instruction fine-tuning: useful for a narrow assistant with examples of high-quality prompts and answers.
- Parameter-efficient fine-tuning: use LoRA or related adapters when GPU memory is limited.
- Retrieval-augmented generation: preferable when answers must cite a changing document collection.
- Tokenizer adaptation: consider adding frequent Rajasthani terms if the base tokenizer fragments them excessively; test this before changing the vocabulary.
For an initial experiment, use a small base model, a low learning rate, short context lengths and frequent checkpoints. Quantised inference can reduce serving costs, but validate quality after quantisation rather than assuming the output is unchanged. If the model must run on low-cost phones or edge hardware, follow the principles in this guide to optimising AI models for mobile devices.
Train with disciplined experiments
Keep configuration, dataset version, random seed, code revision and evaluation results for every run. Begin with a baseline: the untouched base model, a simple retrieval system and a fine-tuned model. This tells you whether training is delivering real gains.
Monitor training and validation loss, but do not treat loss as a complete measure of language quality. Stop when validation performance deteriorates or generated text becomes repetitive. Use gradient accumulation, mixed precision and checkpointing to work within modest GPU limits. Never place test examples in prompts used during model selection.
Evaluate with native-speaker review
Perplexity can compare experiments, but it does not reveal whether a model uses the wrong variety or produces unnatural phrasing. Build an evaluation set covering:
- Next-token prediction and sentence completion.
- Spelling and grammatical correction.
- Translation in both directions, where relevant.
- Named entities, dates, numbers and code-switched text.
- Cultural references and region-specific vocabulary.
- Unsafe, fabricated or overconfident answers.
Use bilingual and native Rajasthani reviewers, ideally from multiple varieties. Ask them to score fluency, faithfulness, dialect appropriateness and usefulness, and to flag Hindi answers presented as Rajasthani. Publish examples, limitations and error categories with the model card. If the model handles images, scanned books or handwritten material, pair the text system with methods covered in open-source vision-language models for Indian languages.
Deploy responsibly and improve through feedback
Expose the model through a simple API with authentication, rate limits, logging controls and clear data-retention policies. For public applications, show when text is AI-generated and provide a correction mechanism. Do not use user prompts for retraining unless users have explicitly agreed.
Track latency, failure rates, token usage, dialect performance and harmful outputs after launch. Route sensitive queries to human review or verified sources. Release datasets, adapters or evaluation scripts where licences permit, but withhold personal or restricted material.
A credible first release might include a model card, data card, reproducible preprocessing code, a small benchmark and ten to twenty native-speaker test cases. That package is more valuable than a large model with unclear provenance—and gives future contributors a reliable foundation.