ASR and TTS models are the two core layers of a voice interface. Automatic Speech Recognition (ASR) turns speech into text; Text-to-Speech (TTS) turns text into spoken audio. Used together, they power call-centre automation, voice search, accessibility tools, classroom assistants, dictation, and conversational agents.
For Indian products, the hard problem is not simply selecting a model with a strong English benchmark. Real deployments must handle code-switching, regional accents, noisy mobile recordings, names and addresses, multiple scripts, and languages with limited labelled data. This guide explains how the systems work, how to evaluate them, and what to plan before putting them into production.
How ASR and TTS fit into a voice system
A typical voice application contains more components than an ASR model and a TTS model:
- Audio capture: A phone, browser, kiosk, or embedded device records speech.
- Pre-processing: Voice activity detection, noise reduction, resampling, and channel handling prepare the audio.
- ASR: The recogniser produces text, timestamps, confidence scores, and sometimes speaker or language information.
- Language and task layer: An application, search system, database, or language model interprets the transcript.
- Response generation: The system creates a reply, action, summary, or structured output.
- TTS: The synthesiser converts the response into audio, often streaming it as it is generated.
- Monitoring: Logs and evaluation pipelines track latency, errors, user corrections, and safety incidents.
This separation makes systems easier to improve. You can replace ASR without rebuilding the dialogue layer, or add a new TTS voice while retaining the same business logic. For language-heavy applications, pair speech components with language resources such as open-source small language models for Hindi rather than assuming one model will solve every part of the product.
ASR models: architectures and selection criteria
Traditional ASR used hidden Markov models, pronunciation dictionaries, and statistical language models. Modern systems are generally neural and may be hybrid, end-to-end, or large multilingual speech models.
- Hybrid systems combine an acoustic model, pronunciation or lexicon resources, and a language model. They remain useful when teams need explicit control over vocabulary, decoding, and domain adaptation.
- Connectionist temporal classification (CTC) models align audio and text without requiring frame-level labels. They are relatively straightforward to train and can be efficient for streaming.
- Sequence-to-sequence models predict text directly from audio and can capture broader context, though decoding and latency need careful management.
- Transducer models are designed for low-latency streaming and are common in interactive applications.
- Multilingual foundation models can transfer knowledge across languages and accents. They are valuable for prototyping, but their average benchmark score may conceal poor performance on a particular Indian language or use case.
When comparing ASR models, measure more than word error rate (WER). Track character error rate (CER) for Indic scripts, latency to first partial transcript, finalisation delay, real-time factor, memory use, and performance under realistic noise. Create separate test sets for Hindi-English code-switching, names, numbers, addresses, dates, domain terms, and regional accents. For Telugu, Sanskrit, Marathi, and other language-specific work, a benchmarking approach such as evaluating NLP models for Telugu and Sanskrit is a useful model for disciplined comparison.
TTS models: naturalness, control, and safety
TTS systems generally contain text normalisation, linguistic analysis, acoustic prediction, and waveform generation. Older concatenative and parametric approaches remain relevant for constrained devices and predictable output, but neural TTS dominates new deployments because it produces more natural rhythm and pronunciation.
Important TTS capabilities include:
- Voice quality: Naturalness, clarity, prosody, and absence of robotic artefacts.
- Pronunciation control: Correct handling of Indian names, abbreviations, currency, dates, numbers, and mixed scripts.
- Style and prosody: Control over pauses, emphasis, speed, pitch, and speaking style.
- Streaming: Low time to first audio for conversational applications.
- Voice coverage: Support for the required languages, genders, age profiles, and regional characteristics.
- Voice protection: Consent, provenance, watermarking, access controls, and safeguards against impersonation.
Do not evaluate TTS only through a listening panel using isolated sentences. Test complete user journeys: a payment amount, a medicine instruction, a delivery address, or a government-service explanation. Ask native speakers to rate intelligibility and pronunciation separately from naturalness. If the system must support translation or multilingual generation, review fine-tuning large language models for Sanskrit translation for the broader data and evaluation issues involved in Indic-language systems.
Building for Indian languages
India’s voice landscape is multilingual, code-switched, and highly variable. A production plan should address:
- Language identification: Detect the user’s language without forcing a language choice too early.
- Code-switching: Preserve English product names and technical terms inside an Indian-language utterance.
- Script handling: Decide whether output should use Latin transliteration, an Indic script, or both.
- Dialect coverage: Record and evaluate regional varieties instead of treating a language as uniform.
- Low-resource adaptation: Use consented speech, transcripts, pronunciation dictionaries, and synthetic augmentation carefully.
- Domain vocabulary: Maintain custom terms for names, villages, medicines, schemes, and local institutions.
Data governance matters as much as model quality. Obtain clear consent for recorded speech, minimise retention, redact personal information, and define who can access raw audio. For sensitive sectors such as healthcare and finance, consider on-device or private-cloud inference and keep a human escalation route.
Deployment choices and production checklist
Cloud APIs can accelerate prototyping, while open-weight models provide more control over data, costs, and customisation. Local deployment may be preferable when connectivity is unreliable or audio cannot leave the device. The right choice depends on traffic, latency, language coverage, compliance, and the team’s operational capacity.
Before launch, validate:
- Latency: Measure end-to-end response time, not only model inference time.
- Reliability: Test poor networks, dropped calls, device variation, and service degradation.
- Cost: Estimate audio minutes, GPU or CPU requirements, storage, egress, and human review.
- Fallbacks: Provide text input, keypad navigation, repetition, and agent handoff.
- Observability: Store privacy-safe metrics for confidence, corrections, failure types, and language mix.
- Security: Protect API keys, audio files, transcripts, and generated voices.
- Scaling: Load-test concurrent streams and peak call volumes before a public rollout.
For teams deploying models in India, how to deploy ML models on AWS Lambda in India offers relevant infrastructure considerations, although sustained streaming workloads may require containers, specialised inference servers, or edge hardware instead of short-lived functions. If the wider application includes a local language model, compare the trade-offs in how to deploy large language models locally.
Common failure modes
The most frequent mistake is optimising a public benchmark while ignoring the target environment. A model may perform well on clean, read speech but fail on call-centre audio, overlapping speakers, or informal conversation. Other recurring issues include incorrect numbers, hallucinated words during silence, poor punctuation, unstable language detection, and TTS mispronunciation of names.
Mitigate these risks with confidence thresholds, confirmation prompts for critical values, constrained grammars where appropriate, human review, and targeted retraining. Never let an uncertain transcript trigger an irreversible payment, medical action, or identity decision without confirmation.
FAQ
Are ASR and TTS the same model?
No. ASR recognises speech and produces text; TTS synthesises speech from text. A voice assistant normally uses both, along with language understanding and application logic.
Which metric should I use?
Use WER or CER for recognition, but also measure latency, domain-specific accuracy, code-switching, and task completion. For TTS, evaluate intelligibility, pronunciation, naturalness, and streaming delay with native listeners.
Should a startup use an API or an open-source model?
Use an API to validate demand quickly. Consider an open model or private deployment when data control, predictable unit economics, offline operation, or language customisation becomes strategic.
What should Indian founders prioritise first?
Start with one high-value workflow, collect representative and consented data, define failure thresholds, and test with native speakers across accents and real operating conditions.
Apply for AI Grants India
If you are building an Indian-language voice product, accessibility tool, speech dataset, or deployment infrastructure, apply to AI Grants India for funding and support. A strong application should state the target users, language coverage, evaluation plan, privacy safeguards, and the measurable problem the system will solve.