Bengali instruction-tuning data should do more than increase token count. It should teach a model to answer questions, follow constraints, summarise official material, and acknowledge uncertainty in Bengali—without losing the meaning or legal context of the source. Indian public documents are a valuable starting point, but they must be collected, licensed, cleaned, and annotated with care.
This guide presents a practical workflow for builders working on multilingual models, public-service assistants, search systems, and Bengali-language applications in India. It focuses on what turns raw documents into dependable instruction-response examples.
Define the dataset’s job first
Start with the model behaviours you want to improve. A narrow, testable scope produces better data than a large archive with vague objectives. Useful Bengali tasks include:
- Question answering: Answer a question using only the supplied document.
- Plain-language explanation: Explain a government scheme, notice, or regulation for a general reader.
- Summarisation: Produce short or structured summaries while preserving dates, eligibility rules, and exceptions.
- Information extraction: Identify names, departments, deadlines, locations, amounts, and required documents.
- Translation and rewriting: Convert English administrative text into natural Bengali while retaining precise meaning.
- Classification: Assign documents to topics such as education, health, agriculture, welfare, or taxation.
- Refusal and uncertainty: State when the source does not contain enough information instead of inventing an answer.
Write a task specification for each category. Define the expected language variety, answer length, citation behaviour, formatting, and acceptable refusals. For model-development context, review best practices for fine-tuning LLMs on custom data before choosing a training strategy.
Select Indian public sources responsibly
Prioritise authoritative, stable sources over convenience. Potential sources include:
- Central and state government portals, department circulars, gazette notifications, and scheme guidelines.
- District administration websites and municipal notices, where Bengali content may reflect practical local terminology.
- Public-sector education, health, agriculture, and disaster-management material.
- Legislative, regulatory, and court-related documents where redistribution and processing rights permit it.
- Public datasets and repositories with explicit licences or terms of use.
“Publicly accessible” does not automatically mean “free to reuse.” Record the source URL, publisher, publication date, language, licence or terms, access date, and any restrictions. Exclude personal data unless there is a documented, lawful reason to retain it. Remove phone numbers, personal addresses, identification numbers, signatures, and case-specific details when they are not required for the task.
Create a source register before extraction. It should let an auditor trace every training example back to its original document and version. This is a core part of data veracity infrastructure for high-stakes AI, particularly when models may support public services or regulated workflows.
Extract and preserve the original evidence
Download documents through permitted methods and retain the original file in read-only storage. Keep a cryptographic hash, source metadata, and extraction log. For HTML, capture meaningful page content rather than navigation, advertisements, or duplicated menus. For PDFs, determine whether the text layer is usable before applying OCR.
OCR is often necessary for scanned Bengali documents, but Bengali script introduces recurring errors involving যুক্তাক্ষর, vowel signs, punctuation, and line breaks. Use OCR confidence scores and route low-confidence pages for human review. Do not silently “correct” numbers, names, dates, or legal phrases; compare them against the page image.
Preserve two versions:
- Evidence text: The closest faithful transcription of the source.
- Training text: A cleaned version used to construct examples.
Keep page, section, paragraph, and table references wherever possible. Tables should be converted into a clear structure rather than flattened into ambiguous text. If a document mixes Bengali and English, retain important official names and technical terms while documenting the language policy.
Clean Bengali text without erasing meaning
Normalisation should improve consistency, not impose artificial uniformity. Build rules for Unicode normalisation, whitespace, punctuation, repeated headers, broken lines, and boilerplate. Review Bengali-specific issues such as visually similar characters, inconsistent danda usage, numerals, transliterated names, and spacing around punctuation.
Avoid aggressive stemming or stop-word removal for instruction tuning. These operations can damage legal meaning and make responses unnatural. Preserve dates, units, acronyms, citations, section numbers, and eligibility conditions. Deduplicate at the document, paragraph, and semantic-example levels. Near-duplicate government circulars can otherwise dominate the dataset.
Run automated checks for:
- Script and language proportions.
- Missing pages or truncated paragraphs.
- OCR artefacts and unusually high symbol rates.
- Personally identifiable information.
- Conflicting versions of the same policy.
- Excessive repetition or copied answer templates.
Convert documents into instruction-response examples
Do not treat every paragraph as a training example. Construct examples around a verifiable user goal and attach the relevant evidence. A useful record can include:
{
"instruction": "এই নথি অনুযায়ী আবেদনের শেষ তারিখ কী?",
"context": "[relevant Bengali passage]",
"response": "নথি অনুযায়ী আবেদনের শেষ তারিখ ...।",
"source_id": "department_notice_2026_014",
"language": "bn-IN",
"task": "document_qa",
"evidence": "page 2, section 4",
"quality_status": "reviewed"
}Generate several task types from the same source, but avoid producing answers that merely copy long passages. Include concise, detailed, structured, and follow-up variants. Ask annotators to distinguish between explicitly stated, reasonably derived, and not stated information. For the last category, the preferred response should say that the document does not provide the answer.
Synthetic drafts can accelerate production, but they must not become unreviewed ground truth. A Bengali-speaking reviewer should check factual fidelity, grammar, register, ambiguity, and whether the response adds unsupported claims. For sensitive domains, use a subject-matter reviewer as well. In medical datasets, align the process with ICMR-compliant medical AI data verification in India.
Build a quality and evaluation system
Use separate people, or at least separate stages, for drafting and approval. Track annotation disagreements rather than hiding them; disagreement often exposes unclear instructions or source conflicts. Score examples on:
- Source faithfulness.
- Bengali fluency and naturalness.
- Completeness and instruction compliance.
- Correct handling of uncertainty.
- Safety, privacy, and harmful-content risks.
- Correct formatting and evidence references.
Create a held-out evaluation set from documents and publishers not used in training. Test Bengali script, code-mixed Bengali-English prompts, regional administrative vocabulary, long documents, tables, dates, numerals, and adversarial requests. Measure exact extraction accuracy where appropriate, but also use bilingual human evaluation for fluency and factuality. Compare against a base model and a non-tuned baseline to verify that tuning actually helps.
Split data by source document, not random rows. Random splitting can place near-identical paragraphs in both training and test sets, producing misleading results. Version the dataset and publish a datasheet covering sources, licences, languages, demographic limitations, transformations, known OCR errors, and intended uses.
Recommended production workflow
A small India-focused team can work in the following order:
1. Define tasks, risk levels, and acceptance criteria.
2. Register sources and verify reuse permissions.
3. Download, hash, and archive original documents.
4. Extract text, apply OCR where needed, and preserve page references.
5. Clean and deduplicate without removing important meaning.
6. Draft examples with evidence and explicit uncertainty labels.
7. Review with Bengali language and domain specialists.
8. Run privacy, licence, and automated quality checks.
9. Split by document and create a held-out evaluation set.
10. Fine-tune, evaluate, inspect failures, and revise the data.
For teams building reusable infrastructure, Indian open-source AI developer projects offer useful patterns for transparent tooling and community review. Keep the pipeline reproducible so a corrected source or improved OCR model can regenerate the dataset without manual guesswork.
Common mistakes to avoid
- Scraping news or government pages without checking licence terms.
- Treating OCR output as authoritative without page-level review.
- Mixing Bengali, Banglish, and English without labelling the intended use.
- Training on outdated circulars while evaluating against current rules.
- Generating synthetic answers that introduce facts absent from the source.
- Measuring only loss or benchmark scores instead of Bengali factuality.
- Publishing documents that contain unnecessary personal information.
High-quality Bengali instruction tuning data is a governed knowledge asset, not a pile of scraped text. A traceable source register, careful Bengali review, explicit evidence, and document-level evaluation will produce a smaller but more dependable dataset—and a model that is safer to deploy in Indian contexts.