What code-mixed speech ASR means
Code-mixed speech ASR converts spoken language into text when a speaker switches between languages within a sentence, phrase, or conversation. In India, a user may say, “Kal meeting reschedule kar do,” or combine Hindi, English, Tamil, Telugu, Bengali, Marathi, and other languages in the same interaction. The switch may happen at a word boundary, inside a technical term, or across a speaker turn.
This is different from ordinary multilingual ASR, where the system identifies one language for an utterance and transcribes it. A production system must detect language transitions, recognise speech across accents and dialects, preserve named entities, and produce text that downstream search, analytics, translation, or customer-support systems can use.
The opportunity is practical: voice interfaces, contact centres, field-service applications, education tools, healthcare access, and public-service platforms often fail when they assume users will speak one standard language. Code-mixed ASR can make these products more natural and more accessible, but only if it is built around real Indian speech rather than clean laboratory prompts.
Why Indian deployments are difficult
India’s code-mixing patterns are not uniform. A Hindi-English speaker in Delhi may switch differently from a speaker in Bengaluru, Hyderabad, Mumbai, or Guwahati. Users also mix languages with English product names, abbreviations, numbers, addresses, and local pronunciations.
Key sources of difficulty include:
- Language identification at short time scales: The system may need to identify a switch within a few hundred milliseconds without delaying transcription.
- Accent and dialect variation: The same word can have different pronunciations across regions, age groups, and levels of English exposure.
- Transliteration ambiguity: “Mujhe kal call karna” could be represented in Devanagari, Latin script, or a mixture of both. Romanised Indian language is not standardised.
- Sparse labelled audio: High-quality, naturally occurring code-mixed recordings are harder to collect and annotate than monolingual speech.
- Noise and conversational speech: Calls, shops, vehicles, homes, and outdoor environments introduce overlap, compression, echo, and incomplete sentences.
- Named entities and domain vocabulary: Names, PIN codes, medicine brands, local places, and technical terms are frequent sources of errors.
Teams working with limited data should treat this as a broader low-resource language problem. The guidance in Low-Resource Indic Natural Language Processing: A Builder’s Guide is useful for planning annotation, transfer learning, and evaluation across Indian languages.
Data is the main product advantage
Model architecture matters, but representative data usually determines whether a code-mixed ASR system works outside a demo. Start by defining the target user and environment: language pairs, region, device type, average utterance length, background noise, and expected output script.
A useful dataset should capture:
- Natural code-switch points rather than sentences created by inserting random English words.
- Multiple regional accents, genders, age groups, and speaking speeds.
- Conversational repairs, hesitations, repetitions, laughter, and overlapping speech.
- Realistic domain vocabulary, including names, addresses, numbers, and product terms.
- Both audio and carefully specified transcripts, including how to represent punctuation, numerals, English words, and transliteration.
- Consent, licensing, demographic documentation, and clear rules for deleting sensitive content.
Do not rely only on synthetic mixing. Artificially combining monolingual recordings can help pretraining and stress tests, but it rarely reproduces natural switching, pronunciation changes, or conversational context. Synthetic data should supplement human speech, not replace it.
For teams building their own pipeline, Low-Resource Language Datasets for AI Training in India offers a useful framework for sourcing, documenting, and governing scarce language data. Preprocessing can also be automated, but scripts must be auditable; practical patterns are covered in Python Scripts for Automating Data Preprocessing.
A practical modelling stack
A modern baseline usually combines a multilingual or language-capable speech encoder with a decoder trained on the target domains. Depending on latency, licensing, and infrastructure, teams may evaluate wav2vec-style encoders, Conformer models, or encoder-decoder systems with multilingual pretraining.
A sensible development path is:
1. Establish a monolingual baseline for each target language and for English. This reveals whether errors come from the base recogniser or from switching.
2. Add language and script metadata where available. Language-ID signals can help, but hard language gating may make rapid switches worse.
3. Fine-tune on naturally code-mixed audio with balanced sampling across language pairs and environments.
4. Add domain adaptation for names, locations, numbers, and product terminology. Use contextual biasing carefully so it does not force likely but incorrect words.
5. Optimise for the deployment setting. Streaming systems need chunking, endpointing, partial-result stability, and predictable latency; batch transcription can prioritise accuracy and larger context.
6. Keep alternatives where ambiguity matters. Contact centres and public services may benefit from confidence scores, n-best hypotheses, or human review rather than a single overconfident transcript.
Large language models can improve punctuation, normalisation, and post-transcription correction, but they should not silently invent content. Keep the acoustic transcript separate from downstream rewriting, record provenance, and evaluate both layers independently. Local or private deployment may be appropriate for sensitive calls; see How to Deploy Large Language Models Locally for infrastructure considerations.
Evaluate the failures that matter
A single overall word error rate hides the behaviour users experience. Report results by language pair, switch type, speaker group, noise condition, and use case. Track at least:
- Word error rate and character error rate, with tokenisation rules documented.
- Language identification accuracy at utterance and switch-point level.
- Switch-point error rate, especially substitutions near language boundaries.
- Entity accuracy for names, numbers, addresses, and domain terms.
- Script and transliteration consistency for Indian-language output.
- Streaming latency and partial-transcript stability.
- Abstention and confidence calibration, particularly for high-impact workflows.
Create a fixed, access-controlled test set that reflects production conditions and a separate challenge set for rare accents, noisy audio, and underrepresented language pairs. Human review remains essential: annotators should distinguish recognition errors from genuine ambiguity and assess whether the transcript preserves meaning.
Product and deployment decisions
A strong recogniser can still fail as a product. Decide early whether users need verbatim transcription, searchable text, subtitles, translated output, or structured fields. Each goal requires different normalisation and post-processing rules.
For India-focused deployments, prioritise:
- On-device or regional inference when connectivity and privacy are constraints.
- Graceful fallback to batch processing when streaming confidence drops.
- User correction loops that feed approved examples back into evaluation, not automatic retraining.
- Clear display of uncertain words instead of presenting guesses as facts.
- Support for mixed scripts and user preferences without forcing one “correct” writing style.
- Monitoring for drift as vocabulary, devices, accents, and usage regions change.
Privacy is especially important for call recordings, health conversations, financial information, and government services. Obtain informed consent, minimise retention, encrypt audio and transcripts, restrict annotation access, and document whether data is used for model improvement.
What to build first
For a new team, the highest-value first release is usually a narrow, measurable workflow: one or two language pairs, one domain, and a defined acoustic environment. Collect a representative pilot set, establish an auditable baseline, and measure errors that affect the workflow—not just benchmark scores.
By 2026, the winning systems will be those that combine strong multilingual pretraining with disciplined Indian data work, transparent evaluation, and deployment-aware design. Code-mixed speech ASR is not simply a language-switching feature; it is a complete data, modelling, product, and governance problem.
Teams building this capability in India can explore support through AI Grants India, including funding pathways and ecosystem resources for applied AI projects.