Nepali is a strong candidate for local-first language AI. It is used across Nepal and in Indian communities in Sikkim, West Bengal, Assam and elsewhere, yet many general-purpose models remain inconsistent with Nepali grammar, register, spelling and cultural context. The right response is not always training a model from scratch. For most teams, a carefully selected base model plus high-quality Nepali data, parameter-efficient fine-tuning and rigorous human evaluation offers a faster and more affordable path.
This guide explains how to approach local language model fine-tuning for Nepali language in 2026, from data collection to production deployment.
Start with the task, not the model
Define the product behaviour before choosing a checkpoint. A Nepali document summariser, voice assistant, tutoring model and government-service chatbot need different data and evaluation criteria.
Write a narrow task specification covering:
- Users and setting: students, farmers, officials, customer-support agents or researchers.
- Language register: formal Nepali, conversational Nepali, regional usage, or mixed Nepali-English input.
- Output format: free-form answers, JSON, summaries, translations or citations.
- Risk level: education and entertainment tolerate different error rates from healthcare, law or public services.
- Deployment constraints: cloud API, Indian data centre, private server, laptop or mobile device.
For broader language-engineering decisions, the low-resource Indic NLP builder’s guide is a useful companion. Nepali-specific work benefits from the same principles: maximise data quality, preserve linguistic diversity and measure performance on real user tasks.
Select a practical base model
Choose an open-weight model with a permissive licence, documented tokenizer and demonstrated multilingual or Indic capability. Candidate families may include Llama, Mistral, Qwen and Indic-focused checkpoints, but benchmark them on your own Nepali test set rather than relying on model popularity.
Check four things before training:
- Tokenizer coverage: inspect how common Nepali words, suffixes and punctuation are split. Excessive fragmentation raises memory use and may limit fluency.
- Existing Nepali ability: test formal prose, colloquial dialogue, translation and code-switching.
- Context length and hardware fit: a smaller model with reliable long-context behaviour may outperform a larger model that cannot run within your budget.
- Licence and data restrictions: confirm whether commercial deployment, redistribution and derivative checkpoints are allowed.
Models trained heavily on Hindi or other Devanagari languages can provide useful transfer, but Hindi proximity is not a substitute for Nepali data. Also test Devanagari normalisation, honorifics and vocabulary specific to Nepal and Nepali-speaking Indian communities.
Build a trustworthy Nepali dataset
Data quality is the main determinant of fine-tuning quality. Combine several sources rather than copying a single web corpus:
- Open Nepali news, literature, educational material and public-domain documents.
- Government notices, forms and service information, where licensing permits.
- Voluntarily contributed conversations and task demonstrations.
- Carefully translated instruction data, reviewed by native Nepali speakers.
- Domain documents supplied by partners under explicit consent and usage agreements.
Separate pre-training-style text from instruction data. Raw text can improve language modelling, while instruction examples teach the model how to follow requests. Do not treat machine-translated Alpaca- or ShareGPT-style records as finished data. Review terminology, politeness, factual content and unnatural literal translations.
A robust cleaning pipeline should include:
- Unicode normalisation and consistent Devanagari representation.
- Removal of boilerplate, navigation text, spam and broken OCR.
- Near-duplicate detection using hashes or MinHash.
- Language identification that retains legitimate Nepali-English code-switching.
- Personal-information detection and removal.
- Licence, consent and provenance records for every source.
- Train, validation and test splits made by document or source—not random lines—to prevent leakage.
Keep dialectal and informal material labelled rather than silently “correcting” it into formal Nepali. A model intended for public use should know when a variation is acceptable instead of treating one register as universally correct.
Fine-tune with LoRA or QLoRA
For most teams, LoRA is the default starting point. It freezes the base model and trains small adapter matrices, reducing storage and compute requirements. QLoRA adds low-bit quantisation to make 7B- to 8B-class models feasible on a single high-memory consumer GPU or modest cloud instance.
A sensible experiment plan is:
1. Establish a zero-shot baseline on a fixed Nepali evaluation set.
2. Run a small supervised fine-tune with 1,000–5,000 carefully reviewed examples.
3. Compare LoRA ranks, target modules and learning rates.
4. Add data in stages, tracking gains by task and register.
5. Merge or serve adapters only after testing regression and licence implications.
Typical starting ranges—not universal prescriptions—are a learning rate around 1e-5 to 2e-4, one to three epochs, gradient accumulation to achieve a stable effective batch size, and sequence lengths based on actual documents. Monitor validation loss, but do not optimise it at the expense of useful outputs. Over-training on a narrow corpus can cause repetition, memorisation and loss of general capability.
Use the workflow in best practices for fine-tuning LLMs on custom data to structure experiment tracking, checkpoint selection and rollback. Record the base model, dataset version, tokenizer, hyperparameters, GPU type and evaluation results for every run.
Evaluate Nepali quality properly
BLEU and ROUGE can help with translation and summarisation, but they are weak proxies for open-ended assistance. Build a Nepali evaluation suite with native-speaker review and task-specific tests.
Measure:
- Fluency and grammar: agreement, word order, spelling and natural phrasing.
- Register and politeness: appropriate use of formal and informal forms, including *tapai* and *timi* where relevant.
- Factuality: especially for public information, health, agriculture and law.
- Instruction following: correct format, scope and refusal behaviour.
- Robustness: Romanised Nepali, spelling variants, code-switching, noisy input and long documents.
- Cultural grounding: local geography, institutions, festivals and social conventions without stereotyping.
- Safety: privacy leakage, harmful advice, discrimination and confident fabrication.
Use blind comparisons between the base model and fine-tuned model. Ask multiple native reviewers to score the same examples, report agreement, and retain difficult disagreements as future test cases. A small, well-designed benchmark is more valuable than a large benchmark that measures only translation overlap.
Deploy for Indian and Nepali infrastructure
Quantise the final model to formats such as GGUF for local runtimes or use an inference server such as vLLM when serving multiple users. Test throughput and latency with realistic Nepali prompts; tokenisation differences can make English-centric capacity estimates misleading.
For edge or mobile use, model size, RAM, context length and battery draw matter as much as raw quality. The AI model optimisation guide for mobile devices covers quantisation and deployment trade-offs that apply to Nepali assistants as well.
Production safeguards should include:
- Retrieval from verified Nepali sources for changing facts.
- Clear uncertainty and escalation paths for high-risk requests.
- Logging that protects personal data and supports incident review.
- Human feedback channels for spelling, terminology and cultural errors.
- Versioned adapters and the ability to revert quickly.
If the application handles sensitive government, education or health data, assess whether inference should remain within India or Nepal, and document retention, access control and cross-border processing requirements.
Build a sustainable Nepali AI project
The strongest Nepali models will come from partnerships among developers, universities, translators, communities and institutions—not from indiscriminate scraping. Pay language experts for annotation and review, publish dataset documentation, and share evaluation protocols even when proprietary data cannot be released.
For teams also working across Indian languages, fine-tuning Llama for Indian regional languages offers useful transfer-learning patterns, while AI tools for local Indian dialects highlights the product and data challenges created by linguistic variation.
FAQ
Can English-only data produce a good Nepali model?
No. Cross-lingual transfer helps, but Nepali supervision and evaluation are necessary for grammar, vocabulary, cultural context and safety.
How much data is enough?
A few thousand high-quality instruction examples can improve a narrow task. Broader assistants need substantially more diverse data, while quality, coverage and deduplication matter more than an arbitrary row count.
Is a 24 GB GPU sufficient?
It can be sufficient for QLoRA on many 7B- to 8B-class models, depending on sequence length, batch configuration and implementation. Profile memory before committing to a training run.
Should I train a new tokenizer?
Usually not for an initial adapter. Consider tokenizer changes only after measuring fragmentation and confirming that continued pre-training or a new model justifies the added complexity.
How can I support Nepali-English code-switching?
Include representative, consented examples, label the intended behaviour, and evaluate both Devanagari and Romanised inputs. Do not erase code-switching during cleaning if it reflects real users.
AI Grants India supports builders developing language technology for South Asia. Explore AI Grants India for potential funding, cloud credits and ecosystem support for responsible local-language AI.