Malayalam AI systems need more than large volumes of text. They need carefully designed examples showing how a model should answer questions, summarise government notices, extract facts, translate terminology, and refuse unsafe or unsupported requests. Indian public documents can provide valuable source material, but only when they are collected lawfully, cleaned carefully, and converted into examples that represent real Malayalam usage.
This guide explains how to create Malayalam instruction-tuning data from Indian public documents in a way that is useful for builders, researchers, and public-interest technology teams. The workflow applies to Malayalam documents from Kerala departments, local bodies, courts, universities, public-sector institutions, and other openly accessible sources.
Define the dataset before collecting documents
Start with the model behaviour you want to improve. A document dump is a language corpus, not an instruction-tuning dataset. Write a task specification before downloading files.
Useful Malayalam tasks include:
- Question answering over public schemes, regulations, and citizen services.
- Short and detailed summarisation of circulars, reports, and notices.
- Information extraction, such as dates, eligibility rules, fees, offices, and contact details.
- Malayalam-to-English and English-to-Malayalam translation of administrative terminology.
- Classification of documents by department, topic, urgency, or audience.
- Rewriting formal government language into plain Malayalam without changing meaning.
- Citation-grounded responses that identify the supporting document and section.
Set target proportions for each task. Include both formal Malayalam and plain-language Malayalam, but do not silently “correct” legitimate regional or institutional usage. If the dataset is intended for public-service assistants, define policies for uncertainty, outdated schemes, personal data, and requests requiring an official source.
For teams planning model adaptation, document the intended base model, context length, output format, and evaluation method. Guidance on best practices for fine-tuning LLMs on custom data is particularly relevant when deciding how much data to create and whether supervised fine-tuning is the right approach.
Select Indian public documents responsibly
Prioritise stable, authoritative sources rather than indiscriminate web collection. Potential sources include:
- Kerala government department and local-government portals.
- Gazette notifications, circulars, tenders, annual reports, and service manuals.
- Public university repositories and government research publications.
- Election, health, education, agriculture, and disaster-management information published by official bodies.
- Public-sector company reports and statutory disclosures.
Publicly accessible does not automatically mean freely reusable. Record the URL, publisher, publication date, retrieval date, licence or terms of use, and document language for every source. Exclude material with unclear rights, paywalled content, private personal information, or restrictions that conflict with model training and redistribution.
Create a source register with a stable document ID. Preserve the original file and a cryptographic hash so later reviewers can identify exactly which version produced each example. For changing webpages, save permitted metadata and an archived reference where possible.
Extract Malayalam text without destroying structure
Use a format-specific pipeline rather than one universal parser:
- HTML: capture the main article or document body, headings, tables, and publication metadata while removing navigation and repeated menus.
- Born-digital PDFs: use layout-aware extraction and retain page numbers, headings, lists, and table boundaries.
- Scanned PDFs: run Malayalam-capable OCR, then manually inspect names, numerals, punctuation, conjunct characters, and government abbreviations.
- Tables: export them separately when row and column relationships matter for answering questions.
OCR errors can create plausible but incorrect Malayalam. Keep both the raw OCR output and the corrected version, and record whether a human verified each page. Do not normalise Unicode blindly: Malayalam characters, vowel signs, zero-width characters, and punctuation require language-aware checks. Preserve numerals when they carry legal, financial, or eligibility meaning.
Useful building blocks include Python, PyMuPDF or pdfplumber, OCR engines with Malayalam language support, and structured storage in JSONL or Parquet. Every extracted chunk should retain document ID, page or section, source URL, language, date, extraction method, and confidence status.
Clean and segment the corpus
Remove headers, footers, page numbers, duplicate pages, navigation text, and OCR artefacts. Keep meaningful legal qualifiers such as “subject to”, “not applicable”, “may”, and “shall”. A summarisation example that loses an exception can teach the model to give harmful advice.
Segment by semantic boundaries—heading, paragraph, list, table, or section—not only by character count. Avoid splitting an eligibility rule from its conditions. For long documents, create overlapping retrieval chunks for evaluation, but use concise source passages when generating supervised examples.
Run deduplication at multiple levels:
- Exact duplicate detection using hashes.
- Near-duplicate detection for repeated circulars and mirrored webpages.
- Template detection for boilerplate disclaimers and recurring forms.
- Cross-source checks for copied press releases.
A small manually reviewed sample should establish the acceptable error rate before processing the full collection. Track Malayalam script, English insertions, transliteration, code-switching, and documents containing other Indian languages.
Convert documents into instruction examples
Each example should make the task, evidence, and expected behaviour explicit. A practical JSONL record can contain:
{"id":"kerala_000184","instruction":"ഈ അറിയിപ്പിന്റെ പ്രധാന യോഗ്യതാ വ്യവസ്ഥകൾ ചുരുക്കുക.","context":"[verified Malayalam passage]","response":"[concise answer]","source":{"document_id":"doc_42","page":3},"task":"summarisation","review":"verified"}Create examples from the source rather than asking a model to invent facts. For each document, produce a mix of straightforward and difficult cases:
- Questions requiring one fact and questions requiring multiple conditions.
- Summaries with specified length and audience.
- Extraction into a fixed schema.
- “Not stated in the document” cases.
- Conflicting or outdated notices where the answer must flag uncertainty.
- Requests that contain personal data or ask for unsupported legal or medical conclusions.
Synthetic drafting can accelerate annotation, but it should never be the final authority. Ask annotators to cite the exact passage supporting each answer. For high-stakes domains, use a second reviewer and a domain specialist. The principles behind data veracity infrastructure for high-stakes AI are useful for building provenance, review trails, and correction workflows.
Annotate Malayalam quality and safety
Use native Malayalam reviewers who understand the document domain. Give them a written rubric covering:
- Factual faithfulness to the cited passage.
- Grammar, spelling, readability, and appropriate register.
- Correct handling of names, dates, amounts, and place names.
- Preservation of exceptions, uncertainty, and scope.
- Whether the answer follows the instruction without adding unsupported claims.
- Privacy, safety, and harmful-request handling.
Measure agreement on a shared sample before scaling annotation. Label unacceptable outputs rather than quietly editing every response; failure labels help diagnose whether the problem is language quality, extraction, task design, or model generation.
Evaluate before fine-tuning
Split data by document, publisher, and time period—not randomly by example. Otherwise, near-identical passages may appear in both training and test sets. Keep a held-out set containing new departments, scanned documents, tables, and plain-language requests.
Evaluate with both automatic and human methods:
- Exact or field-level accuracy for extraction.
- Citation and evidence accuracy for grounded answers.
- Human ratings for Malayalam fluency, faithfulness, and usefulness.
- Adversarial tests for missing information, outdated rules, and contradictory sources.
- Memorisation and privacy checks for names, phone numbers, and sensitive records.
Compare the tuned model with the base model and a retrieval-augmented baseline. Instruction tuning may improve style while reducing factual reliability if the source examples are weak. For many public-document applications, a strong retrieval pipeline plus a smaller, carefully reviewed tuning set is more maintainable than training on every available page.
Operate the dataset as a maintained asset
Version the source register, extraction code, annotation guidelines, and JSONL files together. Publish a dataset card that states coverage, licences, exclusions, known OCR problems, demographic or geographic gaps, and intended uses. Add a takedown and correction process, especially when documents contain personal information or are later withdrawn.
Schedule refreshes for schemes, regulations, and service information. Mark superseded documents instead of deleting history, and ensure evaluation reports identify the data version used. If your project also serves voice interfaces, test Malayalam responses with the target speech system; the practical considerations in voice agent services for Indian businesses illustrate why text quality alone does not guarantee usable interaction.
A practical launch checklist
Before training, confirm that:
- Every document has provenance, retrieval date, and reuse status.
- OCR output has been sampled and critical pages have human verification.
- Examples cite evidence and preserve dates, conditions, and exceptions.
- Train, validation, and test sets are separated by source document.
- Native Malayalam reviewers have checked quality and register.
- Personal data and unsafe instructions have been removed or labelled.
- Baseline, tuned, and retrieval-based systems have been compared.
- The dataset has a version, licence statement, documentation, and correction path.
The goal is not simply to produce more Malayalam tokens. It is to create traceable examples that teach a model when to answer, how to express an answer clearly, and when to defer to an official source. That standard makes Malayalam AI more dependable for citizens, researchers, and Indian builders.