0tokens

Apply for AI Grants India

Financial support for innovators building the future of AI in India.

Apply now

Chat · how to create marathi instruction tuning data from indian public documents

How to Create Marathi Instruction-Tuning Data from Indian Public Documents

  1. aigi

    Marathi AI systems need more than translated prompts. They need examples grounded in the language people use across Maharashtra and Marathi-speaking communities: government services, education, agriculture, public health, civic administration, and everyday support. Indian public documents can provide that foundation, but only if you turn them into carefully reviewed task examples rather than copying documents into a training folder.

    This guide presents a practical workflow for creating Marathi instruction-tuning data from Indian public documents. It is designed for founders, research teams, student developers, and public-interest technology projects working with open-weight or hosted language models.

    Define the dataset before collecting documents

    Start with the behaviours you want the model to learn. A useful dataset usually contains several task types:

    • Information extraction: identify dates, eligibility rules, fees, offices, or required documents.
    • Question answering: answer questions using a supplied Marathi passage and cite the relevant section.
    • Summarisation: produce short Marathi summaries for citizens, officials, or students.
    • Procedure guidance: explain how to complete an application without inventing steps.
    • Classification: label documents by department, service, audience, or urgency.
    • Translation and rewriting: convert formal administrative Marathi into plain Marathi while preserving meaning.
    • Refusal and escalation: state when the document is unclear, outdated, incomplete, or insufficient to answer.

    Write a dataset brief containing the target users, domains, expected answer length, acceptable sources, safety boundaries, and evaluation criteria. If you plan to fine-tune a model, align the collection plan with best practices for fine-tuning LLMs on custom data rather than deciding the format after annotation has begun.

    Select sources and record provenance

    Prioritise sources that are authoritative, accessible, and legally usable. Potential categories include:

    • Central and state government portals, departmental circulars, schemes, forms, and citizen charters.
    • Municipal corporation notices, district administration pages, and public service instructions.
    • Public university material, open educational resources, and government textbooks.
    • Public health, agriculture, disaster-management, and consumer-awareness material.
    • Legislative, judicial, and policy documents where the licence and reuse terms are clear.
    • Public-domain Marathi literature for language variety, kept separate from factual service data.

    Do not assume that “available online” means “free to reuse.” For every source, store the URL, publisher, publication date, access date, language, licence or reuse statement, document version, and any restrictions. Government ownership does not automatically resolve every copyright, privacy, or database-rights question. Exclude material containing personal identifiers, confidential case details, or unnecessary names and contact information.

    A provenance record should travel with each example. This is essential for correcting outdated scheme rules and for building data veracity infrastructure for high-stakes AI, especially when the model may influence health, welfare, finance, or legal decisions.

    Build a Marathi-first preprocessing pipeline

    Public documents often arrive as PDFs, scans, web pages, tables, and images. Process them in stages rather than relying on one automated conversion:

    1. Ingest and fingerprint: save the original file, calculate a checksum, and assign a stable document ID.
    2. Extract text: use native PDF extraction where possible; apply OCR to scans and verify Devanagari output.
    3. Preserve structure: retain headings, numbered clauses, tables, footnotes, dates, and page references.
    4. Normalise cautiously: standardise Unicode, whitespace, punctuation, and common OCR errors without destroying meaningful formatting.
    5. Segment by meaning: split at sections, clauses, or procedure steps instead of arbitrary character counts.
    6. Filter sensitive content: remove phone numbers, addresses, identity numbers, signatures, and personal case information unless strictly necessary and lawfully processed.

    Marathi text requires special attention to Unicode normalisation, danda punctuation, numerals, abbreviations, spelling variants, and code-mixed English. Keep the original text alongside the cleaned version so reviewers can inspect every transformation. OCR confidence scores can help prioritise manual review, but they are not a substitute for it.

    Convert documents into instruction examples

    A strong example contains an instruction, optional context, an answer, and traceable metadata. For instance:

    {
      "instruction": "या योजनेसाठी अर्ज करण्यासाठी आवश्यक कागदपत्रे कोणती आहेत?",
      "context": "[verified Marathi passage from the source document]",
      "response": "अर्जदाराने ... सादर करणे आवश्यक आहे. दस्तऐवजात उत्पन्न प्रमाणपत्राची वैधता स्पष्ट केलेली नाही; संबंधित कार्यालयाकडून पुष्टी घ्या.",
      "source_id": "mh-dept-2026-014",
      "language": "mr",
      "task": "grounded_qa",
      "review_status": "approved"
    }

    Create examples from the document, not from general model knowledge. Include positive and negative cases: questions answered directly, questions requiring a citation, ambiguous questions, and requests for information absent from the source. Teach the model to say “the document does not specify” rather than fill gaps with plausible details.

    Use Marathi prompts written by native speakers, not only English prompts translated mechanically. Include formal Marathi, plain-language Marathi, common spelling variants, Marathi-English code-mixing where users genuinely use it, and regional terminology. Avoid manufacturing dialect claims from small samples. For voice or conversational applications, pair this dataset work with research on voice agent services for Indian businesses, while keeping speech transcripts and text instruction data clearly separated.

    Review quality, safety, and factual grounding

    Automated checks can catch duplicate examples, malformed JSON, empty responses, language mismatches, excessive length, and prompt leakage. Human review must assess:

    • Marathi grammar, spelling, naturalness, and register.
    • Whether the answer is supported by the supplied passage.
    • Preservation of numbers, dates, names of schemes, conditions, and exceptions.
    • Clear handling of uncertainty and conflicting versions.
    • Harmful advice, discriminatory wording, privacy exposure, and unsafe medical or legal claims.
    • Whether the instruction is genuinely useful to the intended audience.

    Use at least two reviewers for high-impact domains. Track disagreement instead of hiding it; disagreement often reveals ambiguous source language or weak annotation rules. Medical examples require heightened controls and should follow applicable institutional review and domain guidance, including approaches discussed in ICMR-compliant medical AI data verification in India.

    Split, balance, and evaluate the dataset

    Prevent leakage by splitting at the document or source level, not randomly at the example level. Otherwise, near-identical clauses may appear in both training and test sets and produce misleading scores. Balance the dataset across departments, document formats, task types, answer lengths, and formal versus plain Marathi.

    Create a held-out evaluation set that includes:

    • Exact factual questions with verifiable answers.
    • Multi-step procedure questions.
    • Out-of-scope and unanswerable requests.
    • OCR-noisy passages and tables.
    • Dates, numeric values, eligibility exceptions, and negative conditions.
    • Code-mixed and naturally phrased Marathi queries.

    Measure grounded accuracy, citation or evidence precision, refusal quality, Marathi fluency, and robustness to paraphrasing. Have native Marathi evaluators rate usefulness on a defined rubric. Test the base model, fine-tuned model, and retrieval-augmented baseline; fine-tuning is not automatically better when documents change frequently.

    Release responsibly and maintain the corpus

    Version the dataset, annotation guide, scripts, source register, and evaluation results together. Publish a datasheet covering collection dates, licences, exclusions, OCR methods, known gaps, reviewer expertise, and intended use. Keep an update queue for schemes, forms, and policies that expire or change. When a source is withdrawn or corrected, identify every affected example through its document ID.

    For many public-service use cases, retrieval plus a smaller instruction-tuning set is more maintainable than training the model to memorise changing rules. Store effective dates and require the deployed system to display source references where practical. Monitor real queries for recurring Marathi variants, unanswered intents, and unsafe overconfidence—without retaining personal data unnecessarily.

    A well-built Marathi dataset is not the largest dataset. It is a traceable, legally defensible, linguistically natural collection of tasks that teaches a model when to answer, how to explain, and when to defer. Teams looking for reusable engineering approaches can also review Indian open-source AI developer projects and adapt their documentation and evaluation practices.

    Last updated 23 September 2026

AIGI may be inaccurate. Replies seeded from the guide above.