Speech-to-text benchmarks are useful only when they predict what users will experience in production. A single WER score on clean English audio can hide failures caused by Indian accents, code-switching, noisy phone recordings, names, addresses, numbers, or domain terminology. A credible evaluation therefore measures accuracy, speed, robustness, and business-critical errors on data that resembles your product.
This guide explains how to design that benchmark in 2026, compare vendors or models fairly, and turn transcription errors into an engineering roadmap.
Start with the production decision
Define what the benchmark must help you decide before collecting audio. The right test for a call summariser is different from the right test for a voice-payment flow or a medical dictation tool.
Specify:
- Supported languages and scripts, including whether users mix English with Hindi, Tamil, Telugu, or another language.
- Audio conditions: smartphone microphone, Bluetooth headset, call recording, far-field microphone, background noise, and network variability.
- Response mode: batch transcription, streaming captions, voice commands, or post-call analytics.
- Critical fields: names, amounts, dates, account numbers, addresses, product codes, and negations such as “not”.
- Operational limits: maximum latency, cost per audio minute, data residency, retention, and concurrency.
For products serving several Indian languages, pair this work with a review of AI speech recognition for Indian regional languages. It helps distinguish a genuinely multilingual system from one that performs well only on a narrow English test set.
Build a representative test set
Use a held-out evaluation set that the model or vendor has never seen. Keep training, tuning, and test speakers separate; otherwise, speaker familiarity can inflate results.
A practical dataset should be stratified by:
- Language and locale: Include the languages your product promises, with meaningful sample sizes rather than token examples.
- Code-switching: Capture realistic utterances such as “ticket book kar do”, not artificially alternating sentences.
- Speaker diversity: Vary region, age, gender, speaking rate, pronunciation, and device experience.
- Acoustic conditions: Include traffic, fans, television, office conversations, reverberation, clipping, and low signal-to-noise recordings.
- Content type: Test commands, conversations, names, numbers, dates, currency, acronyms, and domain vocabulary.
- Audio duration: Include short commands, overlapping turns, and long-form speech.
Document sampling quotas and publish a dataset card internally. A benchmark that overrepresents one city, one call-centre microphone, or one highly scripted prompt will not generalise. For language-specific evaluation, compare your design with the principles used in benchmarking NLP models for Telugu and Sanskrit.
Create trustworthy ground truth
Ground truth is a human-verified transcript, not an automatically generated reference. Give annotators clear rules for punctuation, casing, numerals, abbreviations, hesitation sounds, disfluencies, speaker turns, and code-switched words.
For high-stakes applications, use two independent transcribers and adjudicate disagreements. Track agreement rates and flag clips where the audio is genuinely unintelligible. Do not silently delete difficult samples: report them as a separate quality tier.
Preserve at least two reference forms where useful:
- Verbatim reference: Records what was spoken, including disfluencies and repetitions.
- Canonical reference: Normalises acceptable variants such as “ten rupees” and “₹10”, when the application treats them as equivalent.
This distinction prevents formatting conventions from overwhelming the linguistic signal. Store speaker, language, environment, device, and annotation metadata alongside every clip, while applying consent, access control, and retention policies appropriate for personal audio.
Measure WER, CER, and task-critical accuracy
The standard metric is Word Error Rate (WER):
WER = (Substitutions + Deletions + Insertions) / Reference words
WER is useful for comparing like-for-like systems, but it is not a percentage of words that are correct and can exceed 100% when insertions are high. Always report substitutions, deletions, and insertions separately. A high deletion rate may indicate missed speech; excessive insertions may signal hallucinated words or unstable decoding.
Add complementary measures:
- Character Error Rate (CER): Helpful for Indian scripts, spelling-sensitive terms, and cases where tokenisation rules vary.
- Entity accuracy: Exact or normalised accuracy for names, locations, dates, amounts, IDs, and other fields that drive workflows.
- Command or intent accuracy: Whether the downstream action is correct, not merely whether every word matches. This connects naturally to intent extraction in short text.
- Speaker and language identification accuracy: Important in multilingual calls and meeting audio.
- Real-Time Factor (RTF): Processing time divided by audio duration. For streaming systems, also measure time to first partial, endpointing delay, and final-transcript delay.
- Cost and failure rate: Include API cost, timeouts, unsupported formats, and rate-limit behaviour.
Do not publish one blended score without showing slices. Report results by language, accent or region, noise level, device, utterance length, and content type. Include confidence intervals or bootstrap intervals so small differences are not mistaken for meaningful improvements.
Normalise consistently, not conveniently
Run the same documented normalisation policy on references and hypotheses. Typical rules include Unicode normalisation, case handling, punctuation, whitespace, numerals, currency symbols, and approved spelling variants. Keep raw transcripts for auditability; never overwrite them with scored text.
Be careful with Indian-language scripts and transliteration. “Namaste” and its native-script equivalent may be semantically identical but should not be treated as interchangeable unless your product accepts both. Similarly, “UPI”, “U P I”, and a spoken expansion may need a domain-specific policy.
Use established tooling such as JiWER, Hugging Face Evaluate, or SCTK, but inspect alignments rather than trusting a command-line score. A reproducible scoring script should pin package versions, save configuration, and produce per-utterance error records.
Analyse errors like a product team
After scoring, review the largest and most consequential error groups. Build an error taxonomy covering:
- Language switches and transliteration.
- Proper nouns and regional place names.
- Numbers, dates, currency, and alphanumeric identifiers.
- Negation and instruction verbs.
- Overlapping speakers and turn boundaries.
- Noise, clipping, reverberation, and silence detection.
- Unsupported vocabulary and pronunciation variants.
Use confusion matrices and example clips to decide whether the remedy is better prompting, vocabulary injection, endpointing, audio preprocessing, a domain language model, or a different provider. For a voice product, latency is part of usability; review low-latency audio-to-text processing for Indian startups alongside accuracy results.
Compare systems fairly and monitor drift
Run every candidate on identical audio, decoding settings, normalisation, and hardware or API conditions. Warm up services before latency tests, repeat runs when systems are nondeterministic, and separate transcription quality from post-processing such as punctuation or summarisation.
Keep a locked regression set for releases and a fresh production-like set for drift. Re-test after changes to microphones, codecs, language coverage, vocabulary, or vendor models. Track performance over time by cohort; an overall score can remain stable while one regional language deteriorates.
A practical benchmark report
Your final report should include:
- Dataset size, sampling plan, consent status, and metadata coverage.
- Ground-truth policy and normalisation rules.
- WER, CER, substitutions, deletions, and insertions.
- Entity, command, language-ID, and speaker-turn accuracy where relevant.
- Streaming latency, RTF, cost, timeout, and failure rates.
- Confidence intervals and results by important slices.
- Representative error examples and recommended fixes.
- A go/no-go threshold tied to user and business impact.
For multilingual products, combine transcription testing with the broader evaluation discipline described in multilingual LLM benchmarking in India, especially when transcripts feed search, assistants, or automated decisions.
Benchmarking checklist
- [ ] Define production languages, users, environments, and critical entities.
- [ ] Create a speaker-independent, representative test set.
- [ ] Produce human-verified verbatim and canonical references where needed.
- [ ] Freeze normalisation and scoring code.
- [ ] Report WER, CER, error types, latency, cost, and failures.
- [ ] Slice results by language, noise, device, speaker, and domain.
- [ ] Audit code-switching, names, numbers, and high-impact errors manually.
- [ ] Lock a regression set and monitor production drift.
The best benchmark is not the one that produces the lowest headline WER. It is the one that exposes where a speech system will fail Indian users, quantifies the cost of those failures, and gives builders a clear path to improve it.