Start with the right problem definition
Learning how to create a small language model for Haryanvi begins with choosing a narrow use case. A model for intent classification, text normalization, search, or assisted translation needs far less data and compute than a general-purpose conversational model. Define the target users, dialect coverage, script, latency requirement, and acceptable error rate before collecting data.
Haryanvi is commonly written in Devanagari, but speakers may also use Latin transliteration, Hindi spelling, English words, emojis, and code-switching. Do not erase these patterns during cleaning. They are part of how people communicate on WhatsApp, search boxes, voice interfaces, and local commerce applications.
For background on the engineering constraints involved, see this builder’s guide to low-resource Indic natural language processing.
Choose a practical model strategy
There are three sensible starting points:
- Classifier or tagger: Use this for intent detection, moderation, entity extraction, or routing. It is the cheapest option and often delivers the clearest business value.
- Continued pre-training: Start with a multilingual or Hindi-capable causal language model and train it on clean Haryanvi text using a small learning rate. This adapts vocabulary, spelling, and style.
- Parameter-efficient fine-tuning: Use LoRA or QLoRA on instruction examples when you need a chatbot, rewriting assistant, or question-answering interface.
Do not train a transformer from scratch unless you have substantial, legally usable text and a strong reason to own the entire stack. For most Indian teams, a compact multilingual base model plus carefully prepared Haryanvi data is more affordable and easier to evaluate. Compare this approach with the practical trade-offs discussed in the guide to open-source small language models for Hindi and the overview of fine-tuning Llama for Indian regional languages.
Build a permissioned, representative dataset
Data quality matters more than raw document count. Assemble sources across geography, age groups, occupations, and communication settings. Possible sources include:
- Publicly licensed stories, folk literature, dictionaries, and educational material
- Opt-in recordings and transcriptions from Haryanvi speakers
- Local media, provided its terms permit machine-learning use
- Synthetic prompts reviewed and corrected by native speakers
- De-identified customer-support or search data collected with explicit consent
Maintain a dataset register containing the source, licence, collection date, speaker or author consent, dialect region, script, and preprocessing history. Never scrape private groups or publish identifiable conversations. Remove phone numbers, addresses, account details, and other personal information before annotation.
Create separate splits by source and speaker, not random lines alone. Otherwise, near-duplicate phrases can leak from training into evaluation and inflate results. Reserve a challenge set containing slang, idioms, spelling variation, code-switching, abusive language, and ambiguous phrases.
Normalize without destroying Haryanvi identity
A useful preprocessing pipeline should preserve meaning while making variation measurable. Store the original text, then create one or more derived versions rather than overwriting it.
Recommended steps include:
1. Detect and remove duplicates, boilerplate, spam, and corrupted Unicode.
2. Normalize Devanagari encoding and punctuation while retaining meaningful emphasis.
3. Keep a transliteration field for Latin-script inputs; do not silently convert every example into Hindi.
4. Mark code-switched words instead of deleting them.
5. Segment long documents and filter extremely short or unusable samples.
6. Run personally identifiable information detection and manual quality checks.
7. Create train, validation, test, and challenge sets by speaker, source, and region.
Ask native reviewers to label whether each example is natural Haryanvi, Hindi with Haryanvi influence, mixed speech, or unusable. This label is valuable for both training and future error analysis.
Tokenization and training setup
Inspect how the base model tokenizes representative Haryanvi sentences. If common words are split into many fragments, inference will be slower and the model may struggle with spelling variation. You can first test an existing tokenizer; replacing it is a larger intervention that requires retraining or substantial adaptation.
For continued pre-training, begin conservatively:
- Use a small learning rate and short runs to avoid losing the base model’s Hindi and multilingual capability.
- Mix Haryanvi with a controlled amount of Hindi and English if the product handles code-switching.
- Track validation loss separately for Devanagari, Latin transliteration, and mixed examples.
- Save checkpoints and compare them on the challenge set, not only on perplexity.
For supervised fine-tuning, write examples that reflect the actual product: short user messages, noisy spelling, local references, and requests that require refusal or clarification. Use LoRA or QLoRA when GPU memory is limited. Keep a held-out set that is never used for prompt design or checkpoint selection.
Evaluate usefulness, not just fluency
Perplexity can show whether the model predicts held-out text, but it does not prove that responses are helpful or culturally appropriate. Build an evaluation suite with native Haryanvi reviewers and report results by category:
- Language identification and dialect classification
- Intent accuracy and entity extraction F1
- Translation or rewriting adequacy
- Factuality and instruction following
- Toxicity, stereotyping, and unsafe advice
- Code-switching and transliteration robustness
- Response latency, memory use, and cost per request
Use pairwise human preference tests with clear rubrics. Ask reviewers whether a response sounds natural, preserves the user’s meaning, and avoids falsely claiming local knowledge. Record disagreement rather than forcing reviewers into a single answer; dialect variation is itself an important product signal.
Deploy for Indian users
For a production prototype, expose the model behind an API with authentication, rate limits, logging controls, and versioned prompts. Quantization can reduce memory requirements, while batching and caching help throughput. If the application runs on low-cost phones or edge hardware, benchmark actual latency rather than relying on parameter count alone. This 2026 guide to AI model optimization for mobile devices covers quantization and other deployment decisions.
Keep human escalation available for high-impact uses such as benefits, health, employment, or legal guidance. Display uncertainty, allow users to correct the output, and use approved corrections to improve later evaluation. For voice products, treat speech recognition, language generation, and text-to-speech as separate components; a fluent text model cannot compensate for poor Haryanvi speech data.
A lean 30-day implementation plan
- Week 1: Define one use case, collect permissions, recruit native reviewers, and create an evaluation rubric.
- Week 2: Assemble and clean a small pilot corpus; label dialect, script, intent, and sensitive content.
- Week 3: Benchmark a multilingual base model, run LoRA or continued pre-training experiments, and compare tokenization.
- Week 4: Test with real users, measure failure categories, quantize the best checkpoint, and document limitations.
Ship only when the model beats a simpler baseline, such as Hindi routing plus a rules layer, on the task that matters. A smaller, transparent system with strong human review is often better than a larger model that produces impressive but unreliable Haryanvi text.
FAQs
Can I fine-tune a Hindi model for Haryanvi?
Yes. A Hindi or multilingual model is a practical starting point, but evaluate it on native Haryanvi examples and preserve dialect-specific vocabulary rather than assuming Hindi performance transfers automatically.
How much data do I need?
The answer depends on the task. A classifier may work with thousands of labelled examples; conversational adaptation needs substantially more varied text and carefully written instruction pairs. Quality, consent, and speaker diversity matter more than an arbitrary token target.
Should I train from scratch?
Usually not. Start with an open model whose licence permits your use, then measure continued pre-training or parameter-efficient fine-tuning against a task-specific baseline.
Where can Indian founders seek support?
Teams building responsible regional-language infrastructure can apply for AI Grants India and use the application to explain their data permissions, evaluation plan, community partnerships, and deployment impact.