Indian public documents can help you build language models that understand local schemes, regulations, administrative language, and multilingual user needs. But downloading a large document collection is not the same as creating useful instruction-tuning data. A strong dataset converts authoritative source material into clear tasks, grounded answers, refusal examples, and evaluation cases—while preserving provenance and respecting copyright, privacy, and access conditions.
This guide presents a practical workflow for teams building datasets in 2026, whether you are creating a domain assistant, a public-service chatbot, a legal research tool, or a multilingual support system.
Start with the model behaviour you need
Define the intended behaviour before collecting documents. Instruction tuning should teach a model how to respond, not merely expose it to more text. Write a short dataset specification covering:
- Users: citizens, entrepreneurs, students, officials, clinicians, or internal staff.
- Tasks: answer questions, summarise circulars, extract eligibility criteria, compare schemes, classify requests, translate, or cite sources.
- Languages: English, Hindi, a regional language, code-mixed input, or several languages.
- Risk level: informational, financial, legal, health, or administrative use.
- Grounding requirement: whether every answer must cite a document, section, page, or date.
- Out-of-scope behaviour: what the model must refuse, qualify, or route to a human.
For model adaptation, review best practices for fine-tuning LLMs on custom data. Instruction tuning is usually most effective after you establish a clean retrieval or knowledge layer; it should not be used to make a model memorise rapidly changing scheme details.
Choose authoritative Indian sources
Prioritise sources with clear ownership, update dates, stable URLs, and identifiable issuing authorities. Useful collections include:
- Central and state government department portals
- Gazette notifications, acts, rules, circulars, and official FAQs
- Parliamentary reports and ministry scheme guidelines
- Court judgments and tribunal decisions, where legally reusable
- Public procurement documents and regulatory orders
- University, examination-board, and public research publications
- Official datasets and reports published through government data portals
- Public-service forms and procedural manuals
Do not assume that “available online” means “free to reuse”. Record the source URL, publisher, licence or terms of use, retrieval date, document version, language, and access method. Exclude pages that contain personal information unless you have a documented legal and ethical basis for processing them. RTI responses may contain sensitive details and should undergo especially strict review.
Create a source register such as:
source_id | authority | title | url | publication_date | version | language | licence | retrieved_atThis register is essential when a policy changes and you need to identify, regenerate, or remove affected examples.
Extract and normalise the documents
Public documents arrive as HTML pages, PDFs, scanned files, spreadsheets, and poorly encoded regional-language text. Preserve the original file and create a reproducible extraction pipeline rather than editing content manually.
A practical pipeline is:
1. Download files with rate limits and respect robots.txt, terms, and authentication rules.
2. Calculate a file hash so duplicate or changed versions can be detected.
3. Extract native PDF text where available.
4. Apply OCR to scans, retaining page numbers and confidence scores.
5. Detect document structure: headings, tables, footnotes, annexures, forms, and references.
6. Normalise Unicode, whitespace, punctuation, dates, and numerals without destroying language-specific meaning.
7. Remove headers, footers, repeated navigation, and OCR artefacts.
8. Store the cleaned text alongside the raw file and extraction metadata.
OCR errors in names, amounts, dates, and eligibility thresholds can create dangerous answers. Use automated checks for suspicious characters, broken numbers, missing pages, and low-confidence regions. For high-stakes material, have a reviewer compare the extracted text with the original page image.
Convert documents into instruction examples
Build examples from specific passages, not from an entire document pasted into every prompt. Each record should make its evidence boundary clear. A useful JSONL structure is:
{
"instruction": "What documents are required to apply for this scheme?",
"input": "Answer using the eligibility section of the source.",
"output": "The applicant must provide ...",
"source": {
"source_id": "scheme_2026_014",
"pages": [4, 5],
"section": "Required documents",
"retrieved_at": "2026-02-10"
},
"language": "en",
"task": "information_extraction",
"risk": "medium"
}Generate a balanced task mix:
- Direct question answering: answer only from the cited passage.
- Structured extraction: return fields such as deadline, department, eligibility, fee, or contact channel.
- Summarisation: produce short and detailed versions while retaining conditions and exceptions.
- Comparison: distinguish two schemes, notifications, or procedural routes.
- Translation and transliteration: preserve names, figures, legal terms, and official terminology.
- Clarification: ask for missing state, category, date, or document details.
- Refusal and uncertainty: say when the source does not answer the question or is outdated.
- Citation behaviour: identify the document title, section, page, and effective date.
Synthetic questions can increase coverage, but answers should be grounded in source text and reviewed. Do not ask a model to invent answers simply because a document is incomplete. Include adversarial cases involving conflicting versions, ambiguous abbreviations, tables, scanned pages, and misleading user premises.
Handle multilingual and Indian-language data carefully
India-relevant data is not automatically multilingual data. Decide whether you need parallel examples, translated examples, native-language examples, or code-mixed conversations. A machine translation pass can introduce errors in legal, medical, financial, and administrative terminology, so use bilingual reviewers for evaluation and a controlled glossary for names and recurring terms.
Track language at the example level, including script and code-mixing. Measure performance separately for English, Hindi, and each target regional language rather than reporting one combined score. Check whether the model changes eligibility rules, drops honorifics, mistranslates dates, or confuses lakh, crore, and international number formats.
Validate quality, safety, and provenance
Use layered review instead of relying on one annotator. At minimum, check:
- Factual faithfulness: every material claim is supported by the cited source.
- Completeness: conditions, exclusions, deadlines, and exceptions are not omitted.
- Instruction clarity: the prompt tests a real user task and has an unambiguous expected response.
- Language quality: spelling, grammar, terminology, and translation are appropriate.
- Safety: personal data, sensitive attributes, unsafe medical or legal advice, and fabricated certainty are removed.
- Provenance: source, page, version, licence, and reviewer are recorded.
Use a two-pass process: first review source and answer alignment, then review the user experience and risk. Keep disagreement labels and adjudication notes; they reveal which parts of the source are genuinely ambiguous.
For medical datasets, use domain governance and consult ICMR-compliant medical AI data verification in India. For systems that make consequential claims, invest in data veracity infrastructure for high-stakes AI, including source freshness checks and traceable corrections.
Split datasets to test generalisation
Do not randomly split near-duplicate pages across training and test sets. A model may appear accurate simply because it has seen the same scheme or notification in another version. Prefer splits by document, authority, time period, or task family. Keep a challenging holdout set containing:
- Newer documents not used during training
- Unseen departments or states
- Low-quality scans and tables
- Multilingual and code-mixed questions
- Questions requiring clarification or refusal
- Conflicting or superseded policy versions
Evaluate exact extraction fields, citation accuracy, groundedness, refusal quality, translation adequacy, and human usefulness. A fluent answer without a valid source should fail the test.
Govern the dataset after release
Treat the dataset as a maintained product. Version the source register, raw files, transformed examples, annotation guidelines, and evaluation results. Build a removal process for withdrawn documents, personal information, licensing concerns, and incorrect examples. Schedule refreshes for policy-heavy collections and display document dates in the application.
A reliable production pattern is to use retrieval for current facts and instruction tuning for behaviour: tone, formatting, multilingual handling, clarification, and source-aware answers. This reduces the need to retrain whenever an Indian government department changes a deadline or form.
FAQ
Can I train on any government PDF?
No. Check copyright, licence terms, access restrictions, personal-data exposure, and whether redistribution is permitted. Keep evidence of your decision for every source family.
Should I use synthetic instruction examples?
Yes, selectively. Use them to expand task and language coverage, then validate them against the source with human review. Synthetic volume cannot compensate for weak grounding.
How many examples do I need?
There is no universal number. A smaller, diverse, well-reviewed set often outperforms a large set of duplicated or noisy examples. Start with a pilot, measure failure modes, and expand where evaluation shows a gap.
Should instruction tuning replace retrieval?
Usually not for changing public information. Retrieval can supply current documents and citations, while instruction tuning teaches the model how to interpret and present that evidence.