Why this dataset work matters
Urdu models need more than large text collections. They need clear tasks, dependable answers, regional context, and careful handling of script and dialect variation. Public documents from India can provide valuable material on education, welfare, law, health, agriculture, culture, and public administration—but only when converted into well-designed examples rather than copied into a training file.
The goal is a dataset in which each record teaches a model how to respond to an instruction in Urdu. A strong record should identify its source, preserve the meaning of the document, avoid unsupported claims, and make the expected response easy to evaluate. Teams planning broader model adaptation should also review best practices for fine-tuning LLMs on custom data before choosing a training format or model.
1. Define the dataset before collecting documents
Start with a written scope. Decide whether the dataset will train:
- Question answering over Indian public information
- Summarisation of Urdu or multilingual documents
- Translation between Urdu, Hindi, English, and regional languages
- Information extraction, such as dates, schemes, eligibility, and locations
- Classification of public-service requests
- Conversational assistance grounded in cited sources
- Safety behaviours, including refusal of unsafe or unsupported requests
Set target proportions for each task. A dataset made entirely of summaries may produce a model that performs well on summarisation but poorly on factual questions or dialogue. Also define the intended audience: Urdu-medium students, citizens seeking government information, journalists, researchers, or developers.
2. Source Indian documents lawfully and systematically
Prioritise sources with clear provenance and reuse terms. Potential collections include official ministry and state-government portals, parliamentary and legislative material, public-domain archives, educational repositories, court and tribunal publications, census or statistical reports, and openly licensed cultural collections. The National Digital Library of India can help locate material, but its availability does not automatically grant permission to redistribute or train on every item.
For every document, record:
- URL, title, publisher, publication date, and retrieval date
- Language and script, including whether it is Urdu, bilingual, or transliterated
- Licence, copyright status, and redistribution conditions
- Document type, topic, geography, and government level
- Whether personal, medical, legal, or other sensitive information appears
- File hash and version, so later updates can be traced
Do not assume that a government-hosted PDF is unrestricted. Exclude material with unclear rights unless you have permission. Remove personal identifiers and avoid constructing examples that expose private individuals. For medical material, add domain review and provenance controls; ICMR-compliant medical AI data verification in India offers a useful reference point for higher-risk workflows.
3. Extract Urdu text without destroying its structure
Urdu PDFs may contain selectable text, embedded fonts, images, or a mixture of all three. Test a small sample before processing a large archive. For born-digital files, use a PDF parser and check whether extracted text preserves reading order. For scans, use an Urdu-capable OCR pipeline such as Tesseract with the appropriate language model, followed by manual review.
OCR quality problems commonly include:
- Confusion between visually similar Arabic-derived characters
- Incorrect joining or separation of Urdu words
- Missing diacritics and punctuation
- Reversed columns, headers, footnotes, and tables
- Errors in numerals, dates, names, and place names
- Mixing Urdu, Arabic, Persian, Hindi, and English tokens
Keep the original file and page references alongside the extracted text. Never overwrite raw OCR output. A practical structure is raw/, ocr/, clean/, annotated/, and release/, with a manifest connecting each derived file to its source.
4. Normalise carefully, not aggressively
Urdu normalisation should improve consistency without erasing distinctions that matter. Standardise invisible characters, repeated whitespace, line breaks caused by layout, and obvious OCR artefacts. Be cautious with character substitutions involving ی, ے, ہ, ھ, ک, ق, and ؤ. Preserve punctuation where it affects meaning, and retain numerals in the form used by the source unless the task requires conversion.
Remove repeated page furniture such as running headers and page numbers, but retain headings, lists, tables, citations, and section boundaries when they help answer a question. Deduplicate near-identical documents using hashes and text similarity. Flag uncertain passages instead of silently “correcting” them. A fluent but invented correction is worse than a visibly marked OCR uncertainty.
5. Turn documents into instruction-response examples
Create examples from bounded passages, with a source citation attached to every answer. Useful templates include:
- Extractive question: “اس دستاویز کے مطابق درخواست جمع کرنے کی آخری تاریخ کیا ہے؟”
- Grounded summary: “درج ذیل حصے کا اردو میں پانچ نکاتی خلاصہ لکھیں۔”
- Eligibility extraction: “اس اسکیم کے لیے اہلیت کی شرائط درج کریں۔”
- Translation: “اس سرکاری پیراگراف کا مفہوم برقرار رکھتے ہوئے اردو ترجمہ کریں۔”
- Comparison: “دونوں دفعات میں بنیادی فرق واضح کریں۔”
- Unsupported request: “اگر جواب ماخذ میں موجود نہ ہو تو واضح طور پر کہیں کہ معلومات دستیاب نہیں ہے۔”
Write instructions in natural Urdu, but include realistic code-switching where Indian users would use English terms, names, acronyms, or numbers. Do not translate every proper noun mechanically. For high-value tasks, have an Urdu-speaking annotator draft the answer and a second reviewer verify meaning against the source.
Synthetic generation can accelerate production, but it should create drafts—not final truth. Require the generator to quote or point to the relevant passage, then have reviewers check factuality, completeness, grammar, register, and whether the answer exceeds the evidence. For annotation operations, build a rubric before scaling and pilot it on at least a few hundred examples.
6. Validate quality and representation
Use automated checks for empty fields, duplicate prompts, malformed Unicode, unexpectedly long outputs, language mismatch, and source citations that do not resolve. Add Urdu-specific checks for script proportion, abnormal character sequences, and accidental Roman Urdu when Urdu script is required.
Human review should score each example on:
- Fidelity to the source
- Correctness of Urdu grammar and vocabulary
- Clarity of the instruction
- Completeness and appropriate brevity of the answer
- Cultural and regional suitability
- Safety, privacy, and legal sensitivity
- Whether uncertainty is represented honestly
Split data by document or source, not randomly by individual examples. Random splitting can place near-identical passages in training and evaluation sets, producing misleading scores. Maintain separate evaluation sets for public-service QA, long-context understanding, translation, summarisation, and refusal behaviour. For critical deployments, use a small expert-reviewed challenge set rather than relying only on aggregate loss.
7. Package the dataset for reproducibility
A useful record might contain:
{
"id": "urdu_india_000123",
"instruction": "اس حصے کا خلاصہ لکھیں۔",
"input": "...",
"output": "...",
"language": "ur",
"source_url": "https://example.gov.in/document.pdf",
"source_pages": [4, 5],
"licence": "...",
"review_status": "two_pass_verified",
"version": "1.0"
}Publish a dataset card describing sources, licences, OCR models, cleaning rules, annotator guidance, known gaps, sensitive-content handling, and intended uses. Include a removal process for rights-holder or privacy requests. If the project is intended for Indian developers, consider documenting the pipeline alongside Indian open-source AI developer projects and make scripts, manifests, and evaluation prompts version-controlled.
8. Measure usefulness after fine-tuning
Compare a baseline model with the adapted model on held-out Urdu evaluations. Track factual accuracy, citation accuracy, answer completeness, script fidelity, translation adequacy, and harmful hallucination rates. Test both standard Urdu and realistic Indian usage, including mixed-language prompts and names from different regions.
A smaller, carefully verified dataset can outperform a much larger noisy corpus. When performance improves, inspect which task types and source domains caused the gain. When it declines, check for OCR corruption, duplicated passages, licence-driven removals, overfitting to one government department, or answers that reward confident invention.
FAQ
Can any Urdu PDF from an Indian website be used?
No. Check copyright, licence, terms of use, and privacy implications before collection or redistribution.
Should I use machine translation to create Urdu answers?
Use it for drafts or augmentation, then require Urdu-speaking review—especially for legal, health, welfare, and eligibility content.
How much data is enough?
Start with a small, balanced pilot and a strong evaluation set. Quality, task coverage, and source diversity matter more than an arbitrary document count.
Should I remove all English words?
No. Preserve meaningful names, acronyms, technical terms, and natural Indian code-switching; define the expected register for each task.
What is the most important quality control?
Traceability: every answer should be reviewable against a specific source passage, with uncertainty and licensing recorded.