A voice agent is a real-time system, not a chatbot with audio added at the end. It must capture speech, recognise words, decide what to say, speak quickly, and stop naturally when the user interrupts. For Indian products, the engineering challenge is amplified by code-switching, uneven connectivity, regional languages, background noise, and diverse accents.
This guide explains how to build a production-oriented voice agent with Whisper for speech-to-text (STT) and ElevenLabs for text-to-speech (TTS). It also covers model choices, streaming architecture, latency budgets, privacy, evaluation, and the decisions that matter before you launch.
If you need the fundamentals first, start with what a voice agent is. If you are comparing vendors rather than building in-house, review voice agent software for small businesses.
Reference architecture
A reliable pipeline has six layers:
1. Audio capture: A web or mobile client records microphone input through WebRTC or a native audio API.
2. Voice activity detection: VAD detects when speech starts and ends, while noise suppression improves input quality.
3. Speech recognition: Whisper converts audio into text and returns language, timestamps, or confidence signals where available.
4. Agent orchestration: An LLM, business rules, tools, and retrieval determine the response.
5. Speech synthesis: ElevenLabs converts response text into streamed audio.
6. Session control: The system handles interruptions, retries, authentication, logging, and handoff to a human.
Use WebSockets or WebRTC for a live session. A conventional request-response API can work for voicemail and asynchronous workflows, but it usually feels slow for conversation.
Choose the right Whisper deployment
You can use Whisper through a hosted API or run an implementation such as faster-whisper yourself.
- Hosted Whisper: Fastest to integrate, with no GPU operations. It is suitable for pilots and workloads where audio can leave your infrastructure.
- Self-hosted faster-whisper: Gives you more control over data residency, throughput, and cost at scale, but requires GPU capacity, monitoring, and model optimisation.
- Whisper.cpp or smaller models: Useful for edge and offline scenarios, though accuracy and language coverage may vary.
Model selection is a measured trade-off. Larger models can improve difficult audio and multilingual recognition, but they demand more memory and increase processing time. Benchmark with recordings from your actual users rather than relying on English-language test clips.
For Hinglish, do not force a language unless you have a strong reason. Allow automatic language detection during exploration, then test whether explicit routing—such as Hindi, English, or a regional language based on the user’s first utterance—improves accuracy. Keep a domain vocabulary list for names, product terms, PIN codes, locations, and abbreviations.
A hosted transcription call can look like this:
from openai import OpenAI
client = OpenAI()
def transcribe(path: str) -> str:
with open(path, "rb") as audio:
result = client.audio.transcriptions.create(
model="whisper-1",
file=audio,
language="en",
prompt="Product names: ...; locations: ..."
)
return result.textTreat the code as an integration starting point. Production systems need file validation, time limits, retries, authentication, and protection against oversized uploads.
Stream the agent instead of waiting for a full answer
The slowest design waits for complete transcription, a complete LLM response, and complete audio synthesis before playing anything. A better design pipelines these stages:
- Send audio frames continuously or in short segments.
- Begin transcription after a stable speech boundary.
- Start the LLM as soon as the user’s turn is sufficiently clear.
- Split the response into short, meaningful clauses or sentences.
- Send safe text chunks to ElevenLabs while the LLM is still generating.
- Play audio immediately and buffer only a small amount ahead.
Use punctuation-aware chunking rather than sending every token. Very small chunks sound unnatural and increase request overhead; very large chunks increase first-audio latency. Avoid splitting numbers, URLs, names, or sentences in the middle.
ElevenLabs supports streaming patterns through its SDK and API. The exact model names and availability can change, so select a current low-latency model from the provider’s documentation and benchmark it with your target voice, language, and output format.
Design the conversation layer for speech
Voice responses must be shorter and clearer than screen-based responses. Set explicit rules for the agent:
- Lead with the answer or next action.
- Keep most turns under two sentences.
- Ask one question at a time.
- Read numbers, dates, and amounts unambiguously.
- Confirm high-risk actions before execution.
- Avoid markdown, tables, long lists, and visual references.
A finance assistant, for example, should say, “Your payment of ₹2,450 is scheduled for Friday. Would you like me to change the date?” It should not read a dense transaction table aloud.
Separate conversational memory from operational state. The LLM can remember the user’s recent intent, while a deterministic state machine tracks authentication, consent, payment status, and tool results. This limits hallucinations and makes audits easier.
Latency targets and optimisation
Measure each stage separately instead of reporting one vague response-time number. Track:
- Time from speech end to transcript availability
- LLM time to first token
- Time to first audio byte
- Total turn duration
- Interruption detection time
- Percentage of turns requiring a repeat
For a natural interaction, aim to make first audio arrive quickly, even if the complete response continues streaming. Regional deployment in an AWS Mumbai or similar location can reduce orchestration round trips, but it cannot remove provider-network latency. Keep connections warm, reuse HTTP clients, compress only where it helps, and avoid unnecessary database calls in the critical path.
Use VAD and client-side noise suppression to reduce dead air and irrelevant audio. Cache stable phrases such as greetings and compliance disclosures, but do not cache personalised or sensitive speech without a clear retention policy.
Interruption handling is essential
Users will speak while the agent is talking. When incoming speech crosses your VAD threshold, immediately stop playback, cancel pending synthesis where possible, and preserve the new utterance. This is called barge-in.
A practical control loop is:
1. Detect new speech on the input stream.
2. Stop the audio player locally so the user hears silence immediately.
3. Cancel or discard queued TTS chunks.
4. Mark the previous response as interrupted.
5. Transcribe the new turn and continue from the latest confirmed state.
Do not let an interrupted tool call continue silently if it could create a side effect. For payments, bookings, or account changes, require explicit confirmation and make tool execution idempotent.
Indian-language and deployment considerations
Test with real audio from the languages and devices you support. Include women’s voices, older speakers, rural and urban accents, low-end phones, call-quality recordings, fan noise, traffic, and code-switching. Accuracy in a quiet English demo says little about performance in a multilingual contact-centre environment.
Provide a clear fallback: keypad input, text chat, callback scheduling, or human transfer. For healthcare workflows, review the additional operational and compliance requirements in patient follow-up with a voice agent. For food-delivery workflows, an order-status flow may need different confirmations and escalation rules, as shown in this Zomato and Swiggy order automation guide.
Privacy, consent, and safety
Audio and transcripts can contain personal, financial, and health information. Before launch:
- Tell users they are speaking with an AI system where required by your policy or applicable law.
- Obtain consent for recording and voice cloning.
- Encrypt data in transit and at rest.
- Minimise retention of raw audio and redact sensitive transcript fields.
- Restrict staff access and maintain audit logs.
- Define deletion, correction, and human-escalation processes.
- Review provider data-use, residency, and subprocessors terms.
Never clone a person’s voice without documented permission. Add authentication before exposing account details, and treat caller-provided instructions as untrusted input.
Cost and production checklist
Your cost model includes transcription, LLM tokens, TTS characters, telephony or WebRTC infrastructure, GPUs if self-hosting, storage, observability, and human handoffs. Calculate cost per completed task—not just cost per minute—and monitor repeat turns, failed recognitions, and abandoned calls. The voice agent pricing guide provides a useful framework for comparing these components.
Before launch, verify that you have:
- A latency dashboard with stage-level metrics
- Golden test recordings in every supported language
- Barge-in and retry tests
- Tool permissions and confirmation gates
- Redaction and retention controls
- Human escalation with transcript context
- Load tests for concurrent sessions
- A rollback path for prompts, models, and voices
Whisper and ElevenLabs can provide strong building blocks, but the product quality comes from orchestration. Start with one narrow workflow, measure recognition and task completion with Indian users, then expand languages, tools, and voice styles only when the evidence supports it.