Real-time speech recognition is a systems problem, not simply a model-selection exercise. For a voice agent, call-analytics product, accessibility tool, or multilingual application, users experience the combined delay of audio capture, network transport, inference, endpointing, text cleanup, and the next response. A highly accurate transcript that arrives too late is still a poor product.
For Indian startups, the challenge is sharper: mobile networks vary widely, users code-switch between English and Indic languages, names and local businesses are difficult to recognise, and GPU budgets are limited. This guide explains how to design low latency audio to text processing for startups without overbuilding the platform.
Define latency before choosing a model
Use a measurable latency budget rather than calling a system “real time.” Track at least:
- Time to first partial: audio capture to the first interim transcript.
- Partial-result cadence: how frequently new hypotheses are emitted.
- Finalisation latency: end of speech to the stable transcript.
- End-to-end turn latency: end of speech to the start of the application response.
- Real-time factor (RTF): processing time divided by audio duration. An RTF below 1 means the system is faster than playback, but concurrency and queueing still matter.
A sensible early target is first partial text within 200–400 ms, final text within roughly 500–900 ms after a clear turn, and a complete voice-agent response beginning within about one second. Measure p50, p95, and p99; averages hide the network and queueing failures that users notice most.
Instrument every stage with a trace ID. Record capture timestamps, packet arrival, VAD decisions, model start and finish, decoder time, post-processing, and downstream response time. This turns “the bot feels slow” into an actionable budget.
Use streaming transport and bounded audio chunks
Batch transcription is unsuitable for interactive conversations because the server waits for an entire file. A streaming pipeline sends small audio frames continuously and returns interim hypotheses while the speaker is still talking.
Choose transport according to the product:
- WebSockets are straightforward for browser and mobile clients that need persistent, bidirectional messaging.
- gRPC over HTTP/2 works well for service-to-service streaming and typed contracts.
- WebRTC is useful when you need low-jitter media transport, echo handling, and interruption support.
Send linear PCM or a carefully tested compressed format at a consistent sample rate. Frames of 20–40 ms are common; larger chunks reduce protocol overhead but increase waiting time. Add sequence numbers, timestamps, heartbeat messages, reconnect logic, and bounded server-side buffers. Never allow a slow consumer to create unbounded memory growth.
If you are building a complete voice product, pair this transcription layer with the design principles in low-latency conversational AI for businesses in India. Audio-to-text is only one part of the perceived response time.
Select an ASR architecture for the workload
The best model is the one that meets accuracy, latency, concurrency, and cost requirements together.
- Streaming-first ASR provides partial results with controlled look-ahead and is usually the safest choice for live agents.
- Optimised Whisper variants can work well for prototypes and asynchronous workloads, especially with faster inference runtimes and smaller checkpoints. Standard Whisper should not be treated as a native streaming solution without testing its wrapper, buffering, and stability behaviour.
- Cloud speech APIs can shorten time to market, but evaluate regional availability, data handling, language coverage, custom vocabulary support, quotas, and predictable p95 latency.
- On-device models reduce transport delay and improve privacy, but device CPU, memory, battery, and thermal limits constrain model size.
Benchmark with representative Indian audio rather than public English samples alone. Include accents, code-switching, noisy roads, call-centre microphones, overlapping speakers, and domain terms. Report word error rate by language and scenario, not just one blended number. For Indic deployments, review the broader issues covered in the low-resource Indic natural language processing guide.
Treat VAD and endpointing as product logic
Voice Activity Detection decides whether speech is present; endpointing decides when a turn is complete. Both strongly affect perceived speed.
A long silence threshold makes the system feel unresponsive. A threshold that is too short cuts off thoughtful speakers and creates fragmented transcripts. Start with a 400–700 ms silence window, then tune it by language, channel, and use case. Telephone audio, noisy environments, and conversational agents often need different settings.
Use three transcript states: interim, stable, and final. Downstream systems should avoid triggering expensive actions on every interim revision. For voice agents, support barge-in: detect new speech, stop or pause TTS immediately, and preserve the correct conversational state. Add a maximum utterance duration and fallback behaviour for continuous noise or silent connections.
Optimise inference and concurrency
Latency under one user says little about production performance. Benchmark concurrent streams, warm and cold starts, queue depth, GPU memory, and degradation during traffic spikes.
Practical optimisations include:
- Keep model weights resident in memory; do not load a checkpoint per request.
- Use FP16 or INT8 where accuracy remains acceptable, validating WER after every quantisation change.
- Test inference servers and runtimes such as TensorRT, ONNX Runtime, or vendor-supported serving stacks.
- Separate audio ingestion from inference workers so reconnects and backpressure are controlled.
- Batch only when the added waiting time is bounded; dynamic batching can improve throughput but harm interactive latency.
- Pin compatible GPU drivers and maintain a reproducible benchmark environment.
A single T4, L4, or equivalent may be sufficient for an initial deployment, but capacity planning must use streams per GPU at the target p95 latency. CPUs remain viable for VAD, resampling, routing, and small on-device models; they are not automatically inadequate, but large-model concurrent inference can become expensive quickly. For a wider platform view, use this low-latency AI model deployment guide.
Improve Indian-language accuracy without slowing every request
Indian speech products should plan for Hindi-English and other code-switched combinations from the beginning. A language-identification decision made too late can add delay and cause the recogniser to restart. Prefer multilingual models that maintain a single streaming session when users switch languages.
Build a domain phrase list for names, products, locations, acronyms, and financial terms. Test whether phrase biasing improves recognition without increasing false positives. Post-processing can normalise terms such as UPI, GST, KYC, or product SKUs, but keep the raw transcript for auditing and debugging.
Do not assume one punctuation, numeral, or transliteration policy fits every application. A call summary, a search box, and a voice-command parser need different output formats. If the transcript feeds an intent classifier, keep the pipeline lightweight and evaluate it alongside recognition; the intent extraction guide for short text offers a useful downstream framing.
Decide between cloud, edge, and hybrid processing
Cloud processing is usually easiest to update and scale. On-device processing removes round-trip delay and can keep sensitive audio local. A hybrid architecture often gives Indian startups the best trade-off:
- Run VAD, denoising, resampling, and possibly a small recogniser on the device.
- Stream only active speech to the nearest practical region.
- Fall back to cloud inference when confidence is low or the device is constrained.
- Cache model assets and reconnect gracefully on unstable networks.
Treat privacy as an architectural requirement. Document retention, encryption, access controls, consent, deletion, and whether audio is used for training. Minimise raw-audio retention when transcripts are sufficient, while preserving carefully governed samples for quality improvement.
Control cost and operate the system
Measure cost per audio minute, GPU utilisation, concurrent streams per worker, and failed or retried sessions. Autoscale on active streams and queue age rather than CPU alone. Keep a warm pool for interactive traffic, use cheaper capacity for asynchronous jobs, and avoid spot capacity for sessions that cannot tolerate interruption.
Create dashboards for first-partial latency, finalisation latency, WER by language, endpointing errors, disconnect rate, GPU saturation, and transcript correction rates. Run replay tests from anonymised production samples before changing models or VAD thresholds.
A practical launch sequence
1. Define latency and accuracy targets by use case.
2. Build a streaming prototype with full stage-level tracing.
3. Benchmark representative Indian audio at realistic concurrency.
4. Tune VAD, endpointing, phrase biasing, and partial-result policy.
5. Quantise and serve a warm model; verify p95 latency after each change.
6. Add retries, backpressure, privacy controls, and graceful degradation.
7. Launch with a narrow language and domain scope, then expand using measured error data.
The strongest systems do not chase the smallest model or the fastest benchmark in isolation. They control the entire path from microphone to application action, make trade-offs explicit, and optimise for the conditions Indian users actually face.