Why Indic ASR benchmarking needs care
Benchmarking an automatic speech recognition (ASR) model is more than running inference and reporting one Word Error Rate (WER) number. Indian speech data spans multiple scripts, code-switching patterns, accents, dialects, recording conditions, and levels of literacy. A model that performs well on clean Hindi audio may struggle with Marathi-English code-switching, Tamil names, noisy call-centre recordings, or regional pronunciation.
A useful benchmark therefore answers a practical question: which model works best for a defined Indian user, language, audio environment, and latency budget? This guide presents a reproducible Hugging Face workflow for answering it.
If your project works with other low-resource language technologies, pair this process with the principles in Low-Resource Indic Natural Language Processing: A Builder’s Guide. ASR evaluation is one part of a broader language pipeline.
Define the evaluation before choosing a model
Write down the intended use case first. Your benchmark design will differ depending on whether you are building:
- A voice agent for customer support
- Meeting or classroom transcription
- Search over spoken content
- Voice notes and messaging
- Accessibility tools
- On-device or edge inference
Specify the target language, script, dialect coverage, expected audio duration, microphone quality, background noise, code-switching tolerance, and acceptable response time. Also decide whether punctuation, capitalisation, numerals, and named entities matter to the product.
For example, a voice assistant may prioritise conversational accuracy and streaming latency, while a legal transcription workflow may need stronger handling of names, numbers, and domain terminology. Teams building customer-facing systems can also review top-rated voice agent services for Indian businesses to understand where ASR quality affects the wider application.
Select models from the Hugging Face Hub
Search the Hugging Face Model Hub by language, task, architecture, and license. Consider multilingual models such as Whisper variants, Indic-focused checkpoints, and wav2vec 2.0 or XLS-R models fine-tuned for particular languages.
For every candidate, record:
- Model name, revision, and architecture
- Supported languages and scripts
- Sampling-rate expectations
- License and commercial-use restrictions
- Training-data description, where available
- Whether timestamps, language detection, or streaming are supported
- Parameter count, memory footprint, and quantised versions
Do not assume that a model labelled “multilingual” is equally capable across all Indian languages. A smaller language-specific checkpoint may outperform a much larger general model on a carefully matched domain. Pin the exact model revision in your experiment so that future results remain reproducible.
Build a representative test set
Public datasets are useful starting points, but a product benchmark should include audio that resembles real use. Common sources include Mozilla Common Voice, AI4Bharat resources, language-specific corpora, and licensed internal recordings. Check each dataset’s licence, consent terms, speaker restrictions, and redistribution conditions before using it.
Create a test set with speaker-level separation. The same speaker must not appear in both development and test data, since memorised voice characteristics can inflate results. Stratify the test set by:
- Language, dialect, and region
- Gender and approximate age group, where ethically and legally appropriate
- Clean, noisy, reverberant, and mobile-recorded audio
- Short commands, conversational turns, long-form speech, and numbers
- Native-language and code-switched utterances
- Names, places, acronyms, and domain-specific vocabulary
Keep a locked test set that is never used for model selection. Store audio paths, language labels, speaker IDs, duration, sampling rate, transcript, and data provenance in a consistent schema.
Standardise audio and transcripts
Before inference, convert audio to the model’s expected format. Most pipelines require mono PCM audio at a particular sampling rate, often 16 kHz. Do not silently resample or trim files without logging the operation.
pip install -U transformers datasets evaluate accelerate soundfile librosa jiwer pandasNormalise transcripts separately from predictions. Decide in advance how to handle:
- Unicode normalisation and Indic combining marks
- Punctuation and quotation marks
- Whitespace and repeated spaces
- Numerals versus number words
- English words in mixed-language speech
- Abbreviations and spelling variants
Report both strict and normalised scores when possible. Over-aggressive normalisation can hide meaningful errors, especially in names, numbers, and code-switched text.
Run a reproducible Hugging Face evaluation
For supported datasets and models, the pipeline API provides a quick baseline:
from datasets import load_dataset
from transformers import pipeline
model_id = "openai/whisper-small"
asr = pipeline(
task="automatic-speech-recognition",
model=model_id,
chunk_length_s=30,
device=0,
)
test = load_dataset("mozilla-foundation/common_voice_17_0", "hi", split="test")
predictions, references = [], []
for row in test:
audio = {"raw": row["audio"]["array"],
"sampling_rate": row["audio"]["sampling_rate"]}
predictions.append(asr(audio)["text"])
references.append(row["sentence"])For serious comparisons, use batched inference, fixed decoding settings, and a saved prediction file. Record hardware, batch size, precision, model revision, library versions, audio preprocessing, beam settings, and failed or skipped examples. Run each model on the identical utterance list.
Long recordings require special care. Compare the model’s native chunking and your production segmentation separately. Poor chunk boundaries can create deletions and duplicated words that are not representative of the acoustic model itself.
Measure accuracy, speed, and cost
At minimum, calculate:
- WER: word substitutions, deletions, and insertions divided by reference words
- CER: useful for languages or applications where character-level differences matter
- Sentence accuracy: the percentage of utterances transcribed without an error
- Real-time factor (RTF): processing time divided by audio duration
- First-token or first-result latency: important for streaming experiences
- Memory and compute use: relevant to cloud cost and edge deployment
import evaluate
wer = evaluate.load("wer")
cer = evaluate.load("cer")
print("WER:", wer.compute(predictions=predictions, references=references))
print("CER:", cer.compute(predictions=predictions, references=references))Never publish only an aggregate score. Include sample counts, confidence intervals where feasible, per-language results, and breakdowns by noise, duration, and speaker group. A macro-average across languages prevents a high-resource language from dominating the headline result.
Perform error analysis, not just leaderboard comparison
Export aligned reference-prediction pairs and classify the largest failures. Useful categories include phonetic confusions, script errors, code-switching, named entities, numerals, hallucinated text during silence, repeated phrases, and missed words at chunk boundaries.
Look for systematic patterns. High WER on one district’s speakers may indicate accent coverage rather than general model weakness. Strong overall results with poor number recognition may still be unacceptable for finance or healthcare. Review a stratified sample manually and calculate substitution, deletion, and insertion rates separately.
For downstream systems, evaluate task success as well. A transcript with a minor spelling variation may still yield the correct intent, while one wrong medication name can be critical. If ASR feeds an intent classifier, compare the complete pipeline and document how transcription errors affect decisions. Guidance on intent extraction from short text is useful when connecting ASR to command or support workflows.
Turn results into a model decision
Create a scorecard rather than selecting the lowest WER automatically. Weight the metrics according to your product:
- Accuracy by target language and use case
- Robustness to real recordings
- Streaming or batch latency
- GPU, CPU, and memory requirements
- Licence and data-governance fit
- Ease of fine-tuning and deployment
- Monitoring and rollback options
A two-stage process works well: shortlist models using public data, then validate finalists on a private, consented test set. If no checkpoint meets the requirement, fine-tune only after diagnosing the failure mode and confirming that additional data covers it.
Common mistakes to avoid
- Comparing models on different test sets or transcript conventions
- Mixing speakers across development and test splits
- Reporting WER without defining text normalisation
- Ignoring code-switching and named entities
- Measuring GPU throughput but not end-to-end latency
- Treating a public dataset score as production performance
- Failing to pin model revisions and software versions
- Publishing audio or transcripts without checking consent and licences
A practical 2026 benchmark checklist
Before sharing results, confirm that you have:
- A locked, speaker-independent test set
- Documented language, script, and audio coverage
- Reproducible preprocessing and decoding settings
- WER, CER, sentence accuracy, RTF, and resource measurements
- Per-language and per-condition breakdowns
- Manual review of representative errors
- Model licence and dataset-governance checks
- Saved predictions, metadata, and experiment configuration
Reliable benchmarking gives Indian AI builders evidence they can use—not just a leaderboard number. It exposes where a model works, where it fails, and what data or engineering change is most likely to improve the product.