Kannada-English AI systems need more than a generic multilingual checkpoint and a larger GPU. Kannada has rich morphology, multiple writing conventions, and limited high-quality digital data, while everyday communication often mixes Kannada, English, transliteration, numerals, and emojis in the same message. A useful small language model (SLM) is therefore a data and evaluation project first, and a modelling project second.
This guide explains how to build a compact model for tasks such as text completion, classification, retrieval augmentation, translation assistance, customer support, and voice-agent back ends. For broader context on Indic data constraints, see this builder’s guide to low-resource Indic NLP.
Start with a narrow product objective
Do not begin by promising a general Kannada-English chatbot. Define one measurable job:
- Predict the next token in Kannada, English, or code-mixed text.
- Classify support messages, reviews, or public-service queries.
- Rewrite transliterated Kannada into Kannada script.
- Generate short bilingual replies under strict length and safety limits.
- Rerank or extract information from Kannada-English documents.
The objective determines your dataset, vocabulary, context length, and evaluation set. A 100–300 million parameter model can be highly effective for a constrained workflow, particularly when paired with retrieval and rules. It will not match a frontier model on open-ended reasoning, and the architecture should not be marketed as if it will.
Build a legally usable data pipeline
Collect data with provenance, permissions, and intended-use records. Potential sources include openly licensed government material, public-domain literature, opt-in conversations, educational resources, approved websites, and synthetic examples reviewed by native speakers. Do not scrape private groups or assume that publicly visible text is free to train on.
Create a data manifest containing:
- Source, licence, collection date, and permitted use.
- Language label: Kannada, English, code-mixed, transliterated, or unknown.
- Domain and document type.
- Quality score and processing history.
- A hash or identifier for deduplication.
For Indian deployments, include Karnataka-specific names, places, institutions, units, dates, phone-number formats, and public-service terminology. Balance formal Kannada with conversational usage, but keep separate domain slices so you can diagnose performance rather than blending everything into one opaque corpus.
Clean Kannada, English, and code-mixed text carefully
Preprocessing decisions can change model quality more than a modest increase in parameter count. Preserve Kannada characters and meaningful punctuation; avoid blindly lowercasing Kannada or deleting symbols that carry conversational meaning. Normalise Unicode consistently, standardise repeated whitespace, and remove boilerplate, corrupt markup, spam, and near-duplicate documents.
Treat transliteration as its own data type. Users may write Kannada with Latin characters in several inconsistent ways—for example, the same phrase may appear with different vowel spellings. Keep the original form, add a carefully generated normalised form only when useful, and label the transformation. Do not convert every Latin-script sentence into Kannada: much of it is genuinely English.
Segment code-mixed data at the sentence and message level, then label language spans where feasible. Split train, validation, and test data by document, user, or source—not by random lines alone. Otherwise, duplicated news stories and repeated templates can produce misleading scores.
Choose the tokenizer before training
A multilingual tokenizer may waste capacity on Kannada or split Kannada words into too many fragments. Compare a baseline tokenizer with a SentencePiece or Unigram tokenizer trained on your licensed Kannada-English mixture. Measure:
- Average tokens per sentence in each language.
- Fragmentation of common Kannada words and suffixes.
- Coverage of names, URLs, numerals, emojis, and code-mixed terms.
- Sequence length and estimated inference cost.
A vocabulary of roughly 16,000–32,000 tokens is a reasonable starting range for a compact bilingual model, but the right size depends on corpus scale and deployment memory. Include explicit special tokens for padding, unknown text, language tags, and task boundaries where required. Test tokenisation on real user inputs before committing to pretraining.
Select an efficient model strategy
For a new model, a decoder-only Transformer is the practical default for completion and generation. Keep the design modest: fewer layers, a manageable hidden size, grouped-query attention where supported, and a context window matched to your product. For classification or extraction, fine-tuning an existing multilingual encoder can be cheaper and more reliable than pretraining a generator.
There are three sensible paths:
1. Adapt an existing open model with continued pretraining on Kannada-English text.
2. Train a compact model from scratch when licensing, tokenizer control, or domain specificity requires it.
3. Use a hybrid system: a small model for intent and language handling, retrieval for facts, and deterministic code for critical actions.
The hybrid option is often best for Indian startups. It reduces hallucination risk and lets the SLM run on a modest CPU or a single affordable GPU. It also aligns with the practical deployment concerns covered in building AI apps for the next billion users in India.
Train in stages and track experiments
Start with a small pilot so you can detect data and tokenizer failures quickly. A typical workflow is:
- Prepare and deduplicate the corpus.
- Train or select the tokenizer.
- Run a short continued-pretraining experiment.
- Inspect generated Kannada, English, and mixed outputs manually.
- Scale only after loss curves and samples look credible.
Use PyTorch with Hugging Face Transformers, Datasets, and Accelerate, or an equivalent stack. Record seed, code revision, dataset version, tokenizer hash, batch size, learning rate, sequence length, precision, checkpoint, and hardware. Use gradient accumulation, mixed precision, gradient clipping, and checkpoint resumption to control cost.
Avoid training too long on a small corpus. Monitor validation loss by language and domain, not only the aggregate. If Kannada loss improves while code-mixed quality worsens, adjust sampling or add targeted data instead of simply increasing model size. Hold out a contamination-resistant test set that the training pipeline never sees.
Evaluate real Kannada-English behaviour
Perplexity is useful for tracking training but does not prove product usefulness. Build a human-reviewed test suite with native Kannada speakers and bilingual evaluators. Include formal Kannada, colloquial Kannada, transliterated input, English-only input, mixed sentences, spelling variation, names, numbers, and ambiguous words.
Measure task-specific quality:
- Generation: factuality, relevance, repetition, grammaticality, and language preservation.
- Classification: macro-F1 across Kannada, English, and code-mixed slices.
- Translation or rewriting: chrF or COMET alongside human adequacy and fluency ratings; BLEU alone is insufficient.
- Safety: abusive content, privacy leakage, stereotyping, and unsafe instructions.
- Operations: latency, time to first token, memory, throughput, and failure rate.
Ask reviewers to mark whether an output is understandable, natural, culturally appropriate, and faithful—not merely whether it is grammatically correct. Maintain a regression set of failures and rerun it after every data, tokenizer, or checkpoint change.
Compress and deploy for the target device
Quantisation can make a small model practical on CPU, mobile, or low-cost cloud infrastructure. Test 8-bit and 4-bit variants against the full-precision checkpoint; Kannada tokenisation and code-mixed sequences may expose quality losses that English-only benchmarks miss. Export to a runtime suited to your target, such as ONNX Runtime, llama.cpp-compatible formats, or a mobile inference stack.
Set explicit limits for context length, output tokens, concurrency, and fallback behaviour. Cache repeated system prompts, stream responses where appropriate, and log latency without storing sensitive user text by default. If the application handles voice, keep speech recognition, language identification, the SLM, and text-to-speech as separately measurable components. A related voice-agent architecture guide can help you reason about those boundaries.
Common mistakes to avoid
- Treating Kannada as English with a different script.
- Mixing licences and ignoring source-level restrictions.
- Evaluating only on random held-out text.
- Translating all code-mixed text into one language before modelling.
- Using synthetic data without native-speaker review.
- Reporting one average score that hides Kannada or transliteration failures.
- Deploying a generative model where retrieval or deterministic logic is safer.
A practical first release
For a first production experiment, target one domain, use a clean and traceable corpus, compare a continued-pretrained checkpoint with a small from-scratch baseline, and publish slice-level evaluation. Start with retrieval-backed responses and constrained output formats. Collect opt-in feedback from Kannada speakers, remove personal data, and turn verified failures into a versioned evaluation set.
The winning Kannada-English SLM will usually be the one that is measurable, affordable, and dependable on a specific Indian workflow—not the largest model on a leaderboard. Build the data and testing discipline first; scale parameters only when evidence shows that capacity, rather than data quality or system design, is the bottleneck.