0tokens

Apply for AI Grants India

Financial support for innovators building the future of AI in India.

Apply now

Chat · how to create hindi instruction tuning data from indian public documents

How to Create Hindi Instruction-Tuning Data from Indian Public Documents

  1. aigi

    Hindi instruction-tuning data should do more than convert documents into question-answer pairs. It must teach a model to follow requests accurately, preserve meaning, handle India-specific terminology, and refuse unsafe or unsupported claims. This guide explains how to create a defensible dataset from Indian public documents, from source selection to release and model evaluation.

    Start with a clear dataset brief

    Before downloading documents, define what the model should learn. A dataset for public-service information will look very different from one for education, legal search, or customer support.

    Write down:

    • Target users: citizens, students, researchers, field workers, or business customers.
    • Hindi profile: formal Hindi, conversational Hindi, Hindi with English technical terms, or a controlled mixture.
    • Task types: summarisation, question answering, extraction, classification, rewriting, comparison, and refusal.
    • Risk level: low-risk informational content versus medical, legal, financial, or welfare guidance.
    • Evaluation criteria: factuality, completeness, citation quality, language naturalness, and safety.

    If the project will fine-tune an open model, review best practices for fine-tuning LLMs on custom data before choosing a format or training strategy. Instruction tuning cannot repair unclear objectives or weak source material.

    Choose Indian public documents carefully

    “Publicly accessible” does not automatically mean “free to reuse.” Build a source register containing the publisher, URL, retrieval date, document language, licence or terms of use, document type, and permitted uses. Preserve the original files and record checksums so every training example can be traced back to its source.

    Useful source categories include:

    • Central and state government notifications, schemes, FAQs, annual reports, and public information material.
    • Parliamentary and legislative documents, subject to the applicable reuse terms.
    • Public court judgments and orders, with careful handling of personal information.
    • University repositories, open educational resources, and publicly released research.
    • Official statistical reports, manuals, standards, and public-sector datasets.

    Prefer stable, authoritative sources over scraped aggregators. Avoid documents containing private citizen data, leaked material, or content whose copyright and reuse rights are unclear. For sensitive projects, maintain a documented legal review rather than relying on assumptions about government content.

    A data veracity infrastructure approach is valuable here: provenance, versioning, source confidence, and audit trails should be part of the dataset design—not an afterthought.

    Extract and normalise the Hindi text

    Indian public documents commonly arrive as scanned PDFs, image-heavy files, HTML pages, spreadsheets, and bilingual publications. Use format-specific extraction and retain the original layout where it carries meaning.

    A practical pipeline is:

    1. Download the source and store its metadata.
    2. Detect whether the document contains selectable text or requires OCR.
    3. Run Devanagari-aware OCR on scans and preserve page references.
    4. Remove repeated headers, footers, page numbers, navigation text, and boilerplate.
    5. Repair broken Unicode, spacing, punctuation, and line-wrap errors.
    6. Segment the document by headings, clauses, tables, lists, and paragraphs.
    7. Compare extracted text against random page images through human sampling.

    Do not over-normalise. Hindi spelling variants, numerals, abbreviations, transliterated names, and English terms may be meaningful. Keep a raw layer, a cleaned layer, and a training layer. This makes it possible to correct an extraction error without rebuilding the entire corpus.

    Tables require special treatment. Convert them into structured records or carefully worded descriptions; do not flatten columns into misleading prose. For legal or policy documents, preserve section numbers and qualifying phrases such as “may,” “subject to,” and “except.”

    Convert documents into instruction examples

    A source paragraph is not automatically an instruction example. Create examples that test useful behaviours and require answers grounded in the supplied context.

    A robust record can include:

    {
      "instruction": "इस योजना के लिए पात्रता की मुख्य शर्तें क्या हैं?",
      "context": "<source passage>",
      "response": "<Hindi answer grounded in the passage>",
      "source": {
        "title": "<document title>",
        "url": "<canonical URL>",
        "section": "<heading or page>"
      },
      "language": "hi",
      "task": "question_answering",
      "risk_level": "medium"
    }

    Useful patterns include:

    • Direct answers: ask for facts explicitly stated in the passage.
    • Structured extraction: request dates, eligibility criteria, amounts, or responsible departments.
    • Summarisation: specify audience, length, and register.
    • Rewriting: convert formal administrative Hindi into plain Hindi without changing meaning.
    • Comparison: compare two clauses or versions while citing both sources.
    • Unanswerable questions: require the model to say that the passage does not contain enough information.
    • Safety and uncertainty: teach the model not to turn general public information into personalised legal, medical, or financial advice.

    Synthetic instructions can accelerate coverage, but a human should verify every generated answer. Never allow a model to generate both question and answer and then accept the pair without checking it against the source.

    Build quality controls into the pipeline

    Use multiple review stages rather than one final spot check. At minimum, inspect for:

    • OCR errors and incorrect Devanagari characters.
    • Hallucinated facts, omitted conditions, and altered numbers.
    • Unnatural Hindi, excessive Sanskritisation, or inappropriate Hinglish.
    • Duplicates and near-duplicates across documents.
    • Contradictions between older and newer versions of a policy.
    • Personal data, defamatory content, or unsafe instructions.
    • Biased representation of regions, communities, genders, and social groups.

    Have Hindi-first reviewers assess language and meaning, and subject-matter reviewers assess domain accuracy. Use a shared rubric with pass, revise, and reject outcomes. For high-stakes domains, apply a stricter protocol similar to ICMR-compliant medical AI data verification, including expert review and escalation rules.

    Split data by source document, not randomly by individual examples. Otherwise, nearly identical passages may appear in both training and test sets, producing inflated scores. Keep a challenge set with long clauses, tables, code-mixed queries, misspellings, ambiguous questions, and requests that the model should decline.

    Evaluate the tuned model in real Hindi use

    Measure more than loss. Track factual accuracy, answer completeness, citation or passage attribution, refusal quality, and instruction adherence. Ask reviewers to compare the model response with the source, not with an ideal answer alone.

    Test variations such as:

    • Formal Hindi versus conversational Hindi.
    • Devanagari numerals versus Arabic numerals.
    • Hindi questions containing English names or abbreviations.
    • Regional references and government department names.
    • Follow-up questions that change a date, condition, or audience.
    • Prompts asking for information absent from the source.

    Release a dataset card describing sources, dates, licences, extraction methods, known gaps, annotator instructions, and intended uses. Version the corpus whenever a source document changes or a correction is made. A smaller, traceable dataset is usually more valuable than a large collection with uncertain provenance.

    Common mistakes to avoid

    • Treating every government PDF as reusable without checking terms.
    • Training directly on OCR output without page-level sampling.
    • Removing all English words, even when official names or technical terms require them.
    • Mixing obsolete and current policy versions without date labels.
    • Publishing personal information copied from judgments or public registers.
    • Measuring only fluency while ignoring factual errors.
    • Using random train-test splits that leak passages across sets.

    Final checklist

    Before training, confirm that each example has a source, a task label, a reviewed answer, and a clear language tag. Confirm that sensitive content is minimised, licences are recorded, duplicates are removed, and evaluation data is held out by document. Then run a small pilot, review failures, and expand only where the data improves a measurable capability.

    Teams building Hindi systems can also study Indian open-source AI developer projects for practical tooling and community patterns. The goal is not simply to produce more Hindi tokens; it is to create instruction examples that are accurate, culturally legible, auditable, and useful to people in India.

    Last updated 23 September 2026

AIGI may be inaccurate. Replies seeded from the guide above.