India’s speech AI market cannot be evaluated with a single leaderboard score. A model that performs well on clean, urban Hindi may fail for a Telugu-speaking farmer in a noisy call centre, a Marathi user speaking with a regional accent, or a Hinglish request transcribed into the wrong script. A useful multilingual speech model evaluation framework in India must measure linguistic quality, user success, fairness, latency, and operational reliability together.
This guide lays out a production-oriented approach for evaluating automatic speech recognition (ASR), text-to-speech (TTS), and speech-to-intent systems across Indian languages and usage conditions.
Start with a clear evaluation contract
Before selecting metrics, define what the system must do. An ASR engine for searchable call recordings has different requirements from a voice agent handling payments or healthcare queries. Write an evaluation contract covering:
- Languages and variants: Include target languages, dialects, accents, and expected code-switching patterns.
- Use cases: Separate transcription, captions, call analytics, voice search, customer support, and speech-to-intent.
- Output policy: Specify whether the system returns native script, Romanized text, normalized text, or multiple representations.
- Risk level: Set stricter thresholds for medical, financial, identity, and government workflows.
- Operational limits: Define latency, throughput, endpointing, streaming stability, and cost per audio minute.
For teams building an agent, evaluation should extend beyond transcription. A restaurant voice system, for example, must correctly capture menu items, quantities, addresses, and modifications; benchmarks from multilingual voice agents for restaurants in India illustrate why domain success matters more than a generic WER score.
Build a representative Indian speech test set
A benchmark is only as credible as its sampling plan. Do not create one pooled “Indian English and languages” test set and report an average. Maintain language-level and slice-level results so regressions remain visible.
Your test set should stratify by:
- Language and dialect: Include both high-resource languages and priority low-resource languages. Record region rather than assuming language equals accent.
- Speaker profile: Capture age, gender, education, first language, urban or rural background, and speech or hearing differences where ethically and operationally appropriate.
- Audio conditions: Test quiet rooms, traffic, markets, television bleed, reverberation, telephone codecs, far-field microphones, and overlapping speakers.
- Speaking style: Include read speech, spontaneous conversation, commands, hesitations, disfluencies, numbers, names, and rapid speech.
- Domain vocabulary: Add agriculture, finance, healthcare, education, logistics, government services, and local place names.
- Script and normalization: Preserve native-script references while separately testing Romanized output, punctuation, numerals, abbreviations, and borrowed English terms.
Store consent, licence, speaker metadata, recording conditions, transcription policy, and version information with every sample. Keep evaluation audio separate from training data, and use speaker-disjoint splits to prevent memorization from inflating results.
Use the right metrics for ASR
Word Error Rate (WER) remains useful for comparable reporting, but it must not be the only metric. Tokenization and word boundaries vary considerably across Indian scripts and annotation conventions. Report:
- WER: Useful for application-level comparison after documenting normalization rules.
- Character Error Rate (CER): Helpful for script-level transcription quality, but sensitive to orthographic conventions.
- Word Information Lost or semantic error rate: Measures whether important meaning-bearing content was lost.
- Entity and number accuracy: Track names, dates, amounts, phone numbers, account identifiers, and locations separately.
- Code-switching WER: Report errors on native-language spans, English spans, and language boundaries.
- Endpointing and streaming metrics: Measure time to first partial, finalization delay, interruption handling, and dropped audio.
Publish both raw and normalized scores. Normalization should address punctuation, Unicode variants, whitespace, numerals, and common spelling conventions without silently correcting substantive model errors. For agglutinative or morphologically rich languages, supplement WER with character, morpheme, or slot-level measures rather than declaring one metric universally superior.
Evaluate code-switching as a first-class capability
Hinglish, Tanglish, and other mixed forms are normal user behaviour, not edge cases. Annotate each utterance for language spans, borrowed terms, transliteration, named entities, and switches within a word or phrase where relevant.
A useful code-switching evaluation report includes:
- Language identification accuracy at utterance and segment level.
- WER for each language span.
- Boundary error rate at language transitions.
- Accuracy for English technical terms embedded in Indian-language syntax.
- Transliteration consistency across native and Roman scripts.
- Intent and slot accuracy when a switch occurs near a critical field.
Test realistic phrases rather than artificial alternation. “Mera recharge failed hai, refund kab milega?” is valuable because the business meaning depends on both the mixed vocabulary and the extracted issue and timeframe.
Measure TTS beyond MOS
For TTS, Mean Opinion Score (MOS) is useful but expensive and difficult to compare across panels. Pair it with objective and task-specific tests:
- Naturalness: Human ratings for fluency, rhythm, pronunciation, and prosody.
- Speaker similarity: Verify consistency when voice identity is a product requirement.
- Pronunciation accuracy: Create lexicons for names, places, acronyms, code-mixed terms, and numerals.
- Intelligibility: Use transcription-based or comprehension tests with native listeners.
- Prosodic adequacy: Evaluate questions, lists, emphasis, amounts, warnings, and turn-taking cues.
- Safety and appropriateness: Test unwanted additions, incorrect numbers, offensive pronunciation, and voice cloning controls.
Use native-speaker panels with a structured rubric. Ask raters to flag the exact span that sounds wrong, not only assign a global score. For mobile products, combine quality with real-device latency, audio size, battery use, and offline behaviour; AI model optimization for mobile devices is relevant when deployment constraints shape the evaluation target.
Add speech-to-intent and agent success metrics
A low ASR error rate does not guarantee a successful conversation. Evaluate the complete path from audio to action:
- Intent accuracy and macro-F1 across languages.
- Slot precision, recall, and exact-match rate.
- Critical-field accuracy for amounts, dates, addresses, and identifiers.
- Task completion rate and escalation rate.
- Clarification quality when confidence is low.
- Hallucination, refusal, and unsafe-action rates.
- Turns per task, time to resolution, and user abandonment.
For agentic systems, log whether the model asked for confirmation before irreversible actions. Evaluate language routing, tool calls, fallback behaviour, and recovery from partial or incorrect recognition. Framework choices should be tested against the same suite; teams comparing orchestration options can also review this AI agent framework for developers in India.
Create a continuous evaluation pipeline
Treat evaluation as a release gate, not a launch presentation. A practical pipeline has four layers:
1. Static golden set: A versioned, speaker-disjoint set covering every language, domain, and risk slice.
2. Adversarial set: Numbers, names, accents, noise, overlap, code-switching, disfluencies, and rare vocabulary.
3. Production replay set: Privacy-reviewed samples from real failures, with sensitive information removed or securely governed.
4. Human review: Native evaluators inspect statistically significant changes and classify error causes.
Run regression tests for every model, prompt, decoder, lexicon, and normalization change. Set minimum thresholds per language instead of allowing a strong Hindi result to hide a severe decline in Kannada or Assamese. Track confidence calibration: a system that knows when it is uncertain can route risky requests to confirmation or human support.
Report results transparently
A credible report should include model version, data cut-off, decoding settings, normalization rules, confidence thresholds, hardware, latency measurement method, and confidence intervals. Show per-language results and slices for gender, region, noise, device, and code-switching. Disclose missing data and excluded samples.
As of 2026, teams should also document whether evaluation data may have appeared in pretraining, how synthetic speech was used, and whether the model was tested for memorization or privacy leakage. Open-source language-model work, including open-source vision-language models for Indian languages, can provide useful precedent for documenting language coverage and data limitations even when the target system is speech-first.
A practical release checklist
Before production release, confirm that:
- Every supported language has a minimum test-set size and a named owner.
- Speaker, region, device, noise, and domain slices are represented.
- WER or CER is paired with entity, intent, slot, and task metrics.
- Code-switching and Romanized input are tested explicitly.
- TTS pronunciation and intelligibility have native-speaker sign-off.
- High-risk actions require calibrated confirmation or fallback.
- Regression dashboards preserve historical language-level results.
- Data consent, retention, licensing, and access controls are documented.
A strong multilingual speech model evaluation framework in India is not a single score or a list of datasets. It is a repeatable measurement system that reflects how Indians actually speak, what products must accomplish, and where errors carry real consequences. Teams that build this discipline early can improve models faster while making performance claims that users, customers, and regulators can trust.