0tokens

Apply for AI Grants India

Financial support for innovators building the future of AI in India.

Apply now

Chat · how to run a quantized marathi model offline

How to Run a Quantized Marathi Model Offline

  1. aigi

    Running a Marathi language model offline is practical when privacy, predictable latency, or unreliable connectivity matters. Quantization makes that deployment more accessible by reducing model size and the cost of inference. The right approach, however, is not simply to download a model and convert it to 8-bit numbers: tokenisation, Marathi script handling, runtime compatibility, and evaluation data all determine whether the result works in production.

    This guide explains how to run a quantized Marathi model offline on a local computer, Android device, Raspberry Pi-class edge machine, or other CPU-first environment. It focuses on text-generation and language-understanding models, while the same principles apply to Marathi translation, classification, summarisation, and speech pipelines.

    Choose the model and quantization format

    Start by defining the task and hardware before selecting a checkpoint. A small encoder model may be sufficient for sentiment analysis or intent classification; a compact decoder model is more appropriate for generation. Check the model card for Marathi coverage, licence, tokenizer files, context length, and supported export formats.

    For India-focused applications, test performance on the actual Marathi varieties and domains you expect: formal Devanagari, colloquial chat, code-mixed Marathi-English, names, place names, and government or healthcare terminology. If your application needs dialect coverage, review guidance on fine-tuning AI models for Marathi dialects before choosing a base model.

    Common deployment formats include:

    • GGUF: A practical choice for CPU inference with llama.cpp-compatible runtimes. It supports several integer and floating-point quantization levels.
    • ONNX: Useful when you need a portable graph and hardware-specific execution providers through ONNX Runtime.
    • TFLite: Suitable for Android and embedded deployments, especially when the model is already supported by TensorFlow tooling.
    • PyTorch export formats: Convenient during experimentation, but generally heavier than an optimised mobile or CPU runtime.

    As a rule, 8-bit quantization is a safer starting point when accuracy is important. 4-bit quantization substantially reduces memory use but can affect Marathi spelling, long-context consistency, and less common vocabulary. Benchmark at least two quantization levels rather than assuming the smallest file is the best deployment.

    Prepare an offline environment

    Download the model, tokenizer, runtime wheels, and any required language resources while you have connectivity. Then test the system in airplane mode or on a machine with network access disabled. This catches hidden dependencies such as remote tokenizer downloads, telemetry, model hubs, or package installation at runtime.

    Create an isolated Python environment for evaluation:

    python -m venv .venv
    source .venv/bin/activate
    pip install --upgrade pip
    pip install transformers torch sentencepiece

    Use the runtime appropriate to the export format. For GGUF models, install a llama.cpp-compatible Python binding or use the native command-line binary. For ONNX, install onnxruntime or onnxruntime-gpu only if the target actually has a supported accelerator. For mobile deployment, follow the Android or iOS runtime’s supported operator and packaging requirements.

    Keep the following files together in an application-controlled directory:

    • Quantized model weights
    • Tokenizer model, vocabulary, configuration, and special-token files
    • A pinned runtime version
    • Marathi preprocessing and postprocessing code
    • A local evaluation set and expected outputs

    This packaging prevents an offline application from silently falling back to a different tokenizer or attempting an internet request.

    Run a quantized model locally

    The exact command depends on the runtime. A Transformers-compatible checkpoint can be loaded locally by disabling remote access:

    import os
    import torch
    from transformers import AutoTokenizer, AutoModelForCausalLM
    
    os.environ["HF_HUB_OFFLINE"] = "1"
    os.environ["TRANSFORMERS_OFFLINE"] = "1"
    
    model_dir = "./marathi-model"
    tokenizer = AutoTokenizer.from_pretrained(model_dir, local_files_only=True)
    model = AutoModelForCausalLM.from_pretrained(
        model_dir,
        local_files_only=True,
        torch_dtype=torch.float32,
    )
    
    prompt = "मराठीमध्ये गावातील पाण्याच्या समस्येचा संक्षिप्त सारांश द्या."
    inputs = tokenizer(prompt, return_tensors="pt")
    with torch.inference_mode():
        output = model.generate(**inputs, max_new_tokens=80, do_sample=False)
    print(tokenizer.decode(output[0], skip_special_tokens=True))

    For a GGUF model, use a llama.cpp-based runtime and set conservative context and thread values first. A typical command-line invocation is:

    ./llama-cli \
      -m ./marathi-model.gguf \
      -t 4 \
      -c 2048 \
      -n 80 \
      -p "मराठीमध्ये तीन वाक्यांचा परिचय लिहा."

    The flags should be tuned to the device. More threads do not always mean lower latency, particularly on thermally constrained phones and single-board computers. For a broader local deployment workflow, see how to deploy large language models locally.

    Handle Marathi text correctly

    Many apparent model failures are preprocessing failures. Preserve Unicode Devanagari text end to end and avoid normalising away combining marks without testing. Marathi uses characters and vowel signs that can be represented in multiple Unicode sequences, so compare normalised and raw inputs during evaluation.

    Also verify:

    • The tokenizer does not split common Marathi words into excessive fragments.
    • Marathi punctuation, danda marks, numerals, and whitespace are preserved as intended.
    • User text is not accidentally transliterated into Latin script.
    • Code-mixed Marathi-English inputs are included if they occur in the product.
    • Output decoding removes only unwanted special tokens, not meaningful Devanagari characters.

    If the model will process scanned documents or images, it needs an OCR or vision component before language inference. The relevant constraints differ from text-only inference; review open-source vision-language models for Indian languages when designing that pipeline.

    Benchmark accuracy, memory, and latency

    Compare the original and quantized versions on a fixed Marathi test set. Include task metrics and human review. For generation, inspect factuality, repetition, grammar, spelling, and code-mixing; for classification, report per-class precision and recall rather than only overall accuracy.

    Record:

    • Model file size and peak resident memory
    • Cold-start and warm inference latency
    • Tokens per second or requests per second
    • Battery or thermal behaviour on the target device
    • Context length at which performance degrades
    • Failure rates for empty, very long, malformed, and mixed-script input

    Use deterministic decoding while comparing versions: set a fixed seed, do_sample=False, and the same prompt and token limit. Then test realistic sampling settings separately. Quantization can change token probabilities enough to increase repetition or alter rare-word selection, even when broad benchmark scores appear stable.

    For mobile and edge devices, pair model quantization with practical optimisation such as shorter context windows, prompt templates, batching limits, and memory mapping. The AI model optimisation for mobile devices guide is useful for planning these trade-offs.

    Package for production use

    A production offline application should fail safely when the model cannot load. Check available memory before initialising the runtime, expose a clear “model unavailable” state, and avoid downloading replacement files without explicit user consent. Store model files in an application directory with integrity checks such as SHA-256 hashes.

    Separate the inference layer from the user interface so that the same model can be tested through a command line, API, or mobile wrapper. Log latency and error categories locally without recording sensitive Marathi prompts by default. If the application handles Aadhaar, health, financial, or legal content, define retention and encryption rules before field deployment.

    Troubleshooting checklist

    • Tokenizer or config not found: copy every tokenizer and configuration file locally and use local_files_only=True.
    • Out-of-memory errors: choose a smaller checkpoint, lower quantization size, or shorter context; close competing applications.
    • Garbled Devanagari: verify UTF-8 handling, tokenizer files, and Unicode normalisation.
    • Very slow output: check thread count, runtime build, memory mapping, and whether the device is swapping.
    • Quality drop after quantization: move from 4-bit to 8-bit, use a better calibration set, or consider quantization-aware fine-tuning.
    • Repetitive generation: reduce the prompt length, adjust repetition controls, and compare against the unquantized model.

    Final deployment checklist

    Before shipping, confirm that the complete bundle works with networking disabled, the licence permits redistribution, and the Marathi test set passes agreed quality thresholds. Pin the runtime and document the exact model hash, quantization method, tokenizer version, context limit, and hardware used for benchmarks. Re-test after every model or runtime update.

    Offline quantized inference is most valuable when it is measurable and maintainable. Select a model that fits the device, protect tokenizer fidelity, benchmark Marathi-specific behaviour, and package every dependency locally. That turns a promising checkpoint into a dependable Indian-language feature.

    Last updated 23 September 2026

AIGI may be inaccurate. Replies seeded from the guide above.