Marathi applications do not need a cloud-sized model to be useful. A compact, locally stored model can power drafting, classification, search assistance, education tools, and support workflows while keeping user data on the device. This guide explains how to run a Marathi small language model offline in a way that is reproducible for Indian developers, schools, NGOs, and early-stage product teams.
The focus is inference—not training a foundation model from scratch. In most cases, download an appropriate open model once, verify its licence, cache every required file, and run it with a local runtime. For background on data, tokenisation, and evaluation in Indic languages, see this low-resource Indic NLP builder’s guide.
Choose the right model and runtime
Start with the task rather than the model name:
- Text generation or chat: use a causal language model with Marathi coverage.
- Classification: prefer a smaller encoder model; it will usually be faster and more predictable.
- Translation: select a multilingual translation model rather than a general chat model.
- Speech applications: pair a local text model with Marathi speech-to-text and text-to-speech components.
Check the model card for Marathi examples, supported scripts, training data, licence, context length, and known limitations. A model that handles Devanagari technically may still produce Hindi-heavy vocabulary, poor Marathi grammar, or unreliable factual answers. The open-source small language models for Hindi guide is useful for comparing regional-language model trade-offs, but do not assume Hindi performance transfers directly to Marathi.
For hardware, 8 GB RAM is workable for a small quantised model, while 16 GB gives more room for longer context windows and multiple processes. A CPU is sufficient for testing. An NVIDIA GPU, Apple Silicon device, or compatible accelerator improves throughput but is not mandatory. For phones and low-power edge hardware, review this guide to AI model optimisation for mobile devices.
Prepare a genuinely offline environment
Use an internet-connected machine only for downloading packages and model files. Then move the complete environment to the target computer through an approved USB drive, internal network, or software distribution process.
Create a virtual environment and install the libraries before disconnecting:
python -m venv .venv
source .venv/bin/activate # Linux/macOS
# .venv\\Scripts\\activate # Windows
pip install torch transformers sentencepiece safetensorsPin versions in a requirements file so that a later reinstall does not silently change behaviour:
pip freeze > requirements.txtDownload the tokenizer and model into a local directory. A typical Hugging Face Transformers model directory may contain config.json, tokenizer files, model weights, and generation settings. Test the directory while connected, then run with offline flags:
export HF_HUB_OFFLINE=1
export TRANSFORMERS_OFFLINE=1On Windows PowerShell:
$env:HF_HUB_OFFLINE="1"
$env:TRANSFORMERS_OFFLINE="1"These settings prevent accidental network calls. Also inspect your application code for telemetry, remote logging, external APIs, and package update checks. “Model runs locally” is not enough if prompts are still sent to a hosted service.
Load a local Marathi model
Replace the path below with the directory containing your verified model. local_files_only=True makes missing files fail clearly instead of triggering a download attempt.
from transformers import AutoTokenizer, AutoModelForCausalLM
import torch
MODEL_DIR = "./models/marathi-small"
tokenizer = AutoTokenizer.from_pretrained(
MODEL_DIR,
local_files_only=True,
)
model = AutoModelForCausalLM.from_pretrained(
MODEL_DIR,
local_files_only=True,
torch_dtype=torch.float32,
)
model.eval()
prompt = "मराठीमध्ये शेतकऱ्यांसाठी पावसाचे महत्त्व थोडक्यात समजावून सांगा."
inputs = tokenizer(prompt, return_tensors="pt")
with torch.no_grad():
output = model.generate(
**inputs,
max_new_tokens=80,
do_sample=True,
temperature=0.7,
top_p=0.9,
repetition_penalty=1.1,
)
print(tokenizer.decode(output[0], skip_special_tokens=True))Use max_new_tokens rather than a large max_length when possible; it gives clearer control over output size. For deterministic workflows such as classification-style prompts, set do_sample=False. For creative drafting, sampling can help, but always impose output limits.
Reduce memory with quantisation
If the model is too slow or exhausts RAM, use a quantised format supported by your runtime. Four-bit or eight-bit weights can substantially reduce memory, though quality and compatibility vary. Test Marathi outputs after quantisation rather than assuming the smaller file is equivalent.
For CPU deployments, llama.cpp-compatible GGUF models are often practical. For Transformers deployments, use a quantisation configuration supported by your hardware and installed PyTorch version. Keep the original model archived so you can compare accuracy, latency, and memory use. Never quantise before establishing a baseline.
Evaluate Marathi quality before shipping
A short English prompt test is not meaningful. Build a Marathi test set covering the actual users and domains you serve:
- Devanagari spelling, punctuation, and numerals
- Formal and conversational Marathi
- Regional vocabulary and code-switching with English
- Names, places, dates, currency, and government terms
- Negation, questions, honorifics, and gender agreement
- Unsafe or uncertain requests requiring refusal or escalation
Have Marathi-speaking reviewers score factuality, fluency, relevance, and unwanted Hindi substitution. Include prompts from Maharashtra’s public-service, agriculture, healthcare, education, and small-business contexts. If your product handles documents, measure performance on the scripts and layouts users actually submit; language support alone does not guarantee document understanding. For multimodal workflows, compare suitable open-source vision-language models for Indian languages.
Track practical metrics: median response time, tokens per second, peak RAM, model load time, and failure rate when disconnected. Repeat tests on the exact target device, not only a developer laptop.
Common offline deployment problems
- Tokenizer errors: copy every tokenizer file and preserve the original directory structure.
- Unexpected downloads: set offline environment variables and
local_files_only=True. - Poor Marathi output: try a better Marathi or multilingual checkpoint, improve prompting, or fine-tune on licensed domain data. This guide to fine-tuning Llama for Indian regional languages covers the main decisions.
- Out-of-memory failures: reduce context length, use a smaller model, quantise weights, or process one request at a time.
- Hallucinated answers: add retrieval from a local, curated knowledge base and show uncertainty instead of presenting generated text as verified fact.
- Licence risk: record the model licence, dataset terms, attribution requirements, and any restrictions before distributing the application.
Package the application responsibly
Ship a self-contained bundle containing the runtime, pinned dependencies, model files, licence notices, configuration, and a health-check script. Add a clear indicator that processing is local. Encrypt sensitive data at rest where appropriate, restrict access to cached prompts and logs, and provide a deletion mechanism.
For production use, separate model inference from the user interface and enforce request limits, timeouts, and maximum prompt lengths. Keep a Marathi error message for unavailable hardware or corrupted model files. Offline does not eliminate operational risk: updates, security patches, model replacement, and evaluation still need a controlled release process.
A small Marathi model can be a strong foundation for privacy-preserving, low-connectivity products—but only when language quality, licences, hardware limits, and offline guarantees are tested together.