0tokens

Apply for AI Grants India

Financial support for innovators building the future of AI in India.

Apply now

Chat · how to clean data for indic small language models

How to Clean Data for Indic Small Language Models

  1. aigi

    Small language models can perform strongly on Indic tasks when their training data is relevant, consistent, and carefully documented. They can also learn the wrong patterns quickly: duplicated web pages, broken Unicode, machine-translated phrasing, mixed scripts, toxic content, and evaluation leakage can all produce misleading gains.

    This guide explains how to clean data for Indic small language models. It is designed for teams building classifiers, retrieval systems, assistants, speech-text pipelines, and compact generative models for Indian languages. The goal is not to erase linguistic variation. It is to remove technical noise while preserving dialects, code-mixing, register, cultural references, and legitimate minority usage.

    Start with a data specification

    Before writing cleaning code, define what the dataset should contain. Record:

    • Target languages, dialects, scripts, and expected proportions.
    • Intended tasks: next-token prediction, instruction tuning, classification, translation, or retrieval.
    • Accepted sources and licensing conditions.
    • Minimum quality thresholds for text length, readability, and metadata completeness.
    • Sensitive categories that require removal, masking, restricted access, or specialist review.

    For low-resource languages, data strategy matters as much as model architecture. The low-resource Indic natural language processing guide covers sampling, annotation, and evaluation choices that should inform this specification.

    Keep a manifest for every source: URL or repository, collection date, language, script, licence, document type, and processing version. This makes later audits and dataset rollback possible.

    Preserve Indic text before normalising it

    Indic data often arrives through PDFs, OCR, websites, messaging platforms, and transliteration-heavy user content. Do not begin by lowercasing or stripping punctuation. First preserve the original text and create a versioned cleaned copy.

    Unicode and script checks

    Normalise Unicode consistently, usually with NFC, and inspect combining marks rather than deleting them blindly. Common issues include:

    • Devanagari, Bengali, Tamil, Telugu, Malayalam, Kannada, Gujarati, Gurmukhi, and other scripts represented inconsistently.
    • Nukta characters split from their base letters.
    • Invisible zero-width characters copied from web pages.
    • Non-standard spaces, repeated line breaks, and OCR artefacts.
    • Numerals represented in both Latin and Indic forms.

    Use a Unicode-aware library and create diagnostic reports showing which characters were changed. Keep a raw-to-clean mapping for samples so a reviewer can confirm that normalisation did not alter meaning.

    Script and language identification

    Run language identification at the document, paragraph, and sometimes sentence level. A single page may contain Hindi in Devanagari, English product names, Urdu in Perso-Arabic script, and Romanised Hindi. Treat this as structured multilingual data, not automatically as noise.

    Flag uncertain examples for review. For code-mixed data, retain the original and add metadata such as language_sequence, script, and is_code_mixed. A model intended for Indian customer support may need this behaviour; a monolingual benchmark may not.

    Remove structural noise and duplicates

    Strip navigation menus, cookie notices, repeated headers, page numbers, HTML boilerplate, tracking parameters, and OCR fragments. Preserve headings, lists, tables, and paragraph boundaries where they carry meaning. For instruction data, retain the relationship between prompt, context, and answer.

    Deduplicate at several levels:

    • Exact duplicates: identical strings after safe whitespace normalisation.
    • Near duplicates: copied news articles, mirrored government pages, and repeated legal or educational material.
    • Template duplicates: pages where only a name, date, or location changes.
    • Cross-script duplicates: the same sentence in native script and Roman transliteration.

    Use hashes for exact matching and MinHash or locality-sensitive hashing for near-duplicate detection. Deduplicate before splitting into train, validation, and test sets. Otherwise, a memorised passage can appear in training and make evaluation scores meaningless.

    For a deeper treatment of provenance, lineage, and reliability controls, see the guide to data veracity infrastructure for high-stakes AI.

    Clean transliteration, spelling, and formatting carefully

    Romanised Indic text is highly variable. For example, one speaker may write a Hindi word in several valid-looking forms depending on pronunciation, keyboard habits, or regional usage. Do not force all Romanised text into one spelling unless your application specifically requires it.

    Instead:

    • Store native-script and Romanised forms separately when both exist.
    • Record the transliteration system where known.
    • Standardise obvious keyboard and encoding errors, not legitimate variation.
    • Preserve punctuation and emoji when they signal intent, sentiment, or conversational style.
    • Avoid blanket stop-word removal, especially for generative models, translation, dialogue, and sentiment tasks.

    Stemming and lemmatisation also require caution. Aggressive stemming can destroy case, gender, number, politeness, and tense information. For most modern language-model training, use clean text with a tokenizer suited to the target scripts rather than deleting morphology as a preprocessing shortcut.

    Filter quality, safety, and privacy risks

    Set explicit rules for content quality. Remove or quarantine records that are empty, mostly corrupted, unintelligible, or dominated by boilerplate. Do not remove dialectal, informal, or low-literacy text merely because it differs from formal textbook language.

    Create separate filters for:

    • Personal data such as phone numbers, addresses, Aadhaar-like identifiers, emails, and account details.
    • Credentials, private medical records, and confidential business information.
    • Hate speech, sexual exploitation, extremist propaganda, and instructions for serious wrongdoing.
    • Copyright-restricted material outside your permitted use.
    • Synthetic text that could overwhelm naturally occurring language.

    Use automated detection for scale, followed by human review for borderline cases. For medical, legal, or public-service datasets, document the review protocol and escalation path. The ICMR-compliant medical AI data verification guide is relevant when clinical text or health records are involved.

    Balance the dataset without erasing reality

    Measure coverage by language, dialect, script, region, source, domain, speaker or author type, and document length. A dataset can contain millions of tokens yet remain weak for rural speech, informal customer queries, or smaller language communities.

    Avoid balancing only by duplicating minority examples. Oversampling can increase memorisation and amplify annotation errors. Prefer additional high-quality collection, controlled sampling, or loss weighting. Keep source and demographic metadata private where necessary, but retain aggregate distribution reports for auditing.

    Build leakage-free splits and evaluation sets

    Split by document, user, source, and time—not random lines alone. Near-duplicate documents from the same publisher should remain in one split. If the model will serve Indian users across regions, create evaluation slices for:

    • Each target language and script.
    • Code-mixed and Romanised input.
    • Dialect and regional vocabulary.
    • Spelling noise and OCR corruption.
    • Long-context, factual, safety, and instruction-following tasks.

    Have native speakers review a representative sample and score fluency, faithfulness, cultural appropriateness, and harmful bias. For custom training workflows, combine these checks with the best practices for fine-tuning LLMs on custom data.

    A practical cleaning pipeline

    A reproducible pipeline can follow this order:

    1. Ingest raw files and record provenance.
    2. Detect encoding, language, and script.
    3. Extract text while retaining document structure.
    4. Apply Unicode and whitespace normalisation.
    5. Remove boilerplate and corruption.
    6. Detect personal data, unsafe content, and licence violations.
    7. Deduplicate exact, near, and cross-script matches.
    8. Attach quality, language, domain, and source metadata.
    9. Sample records for native-speaker review.
    10. Create leakage-resistant splits and publish a dataset card.

    Version every transformation. Store rejection reasons, counts by language, and before-and-after samples. A small team should be able to reproduce the final dataset from raw inputs and configuration files.

    Tools and checks for builders

    Python with pandas, regex, Unicode utilities, language-identification libraries, and datasketch can support a first pipeline. Use OpenRefine for exploratory metadata cleaning, and a database or object store for manifests and review queues. For larger corpora, process in batches and write immutable intermediate outputs rather than repeatedly editing one large file.

    Track metrics such as duplicate rate, rejection rate, language confidence, average sequence length, proportion of code-mixed records, personally identifiable information detections, and token coverage. Compare these metrics across pipeline versions. Cleaning is successful when model quality improves on trusted evaluation slices—not merely when the dataset becomes smaller.

    Final checklist

    Before training, confirm that:

    • Every record has provenance and licence metadata.
    • Unicode, scripts, and transliteration have been inspected.
    • Duplicates cannot cross dataset splits.
    • Personal and confidential information is handled under a documented policy.
    • Native speakers have reviewed samples from every target language.
    • Dataset balance reflects the product’s users and not only web availability.
    • Evaluation sets are held out, representative, and contamination-checked.
    • Rejected records and transformation logs are retained for audit.

    For Indian builders, disciplined data preparation is often the most affordable performance improvement available. Clean for consistency, but preserve the language behaviour your model must serve.

    Last updated 23 September 2026

AIGI may be inaccurate. Replies seeded from the guide above.