0tokens

Apply for AI Grants India

Financial support for innovators building the future of AI in India.

Apply now

Chat · building a voice agent with Whisper and ElevenLabs

Building a Voice Agent with Whisper and ElevenLabs

  1. aigi

    Whisper and ElevenLabs are a practical foundation for voice agents that need accurate transcription and natural speech. Whisper handles speech-to-text (STT), an LLM manages reasoning and dialogue, and ElevenLabs converts the response back into audio. The difficult work is not simply connecting three APIs: it is designing turn-taking, interruption handling, streaming, observability, privacy, and fallbacks so the conversation remains useful when networks or models behave imperfectly.

    This guide focuses on an implementation that can serve Indian users, including Hindi, Hinglish, Indian English, and noisy mobile or call-centre audio. If you are still defining the product, first understand what a voice agent is and how voice AI works in 2026. The architecture below applies to customer support, appointment booking, collections, education, and internal operations.

    Reference architecture

    A production voice agent has six layers:

    1. Audio capture: A browser, mobile app, SIP provider, or telephony gateway receives microphone or phone audio.
    2. Voice activity detection: WebRTC VAD or Silero VAD identifies speech, pauses, and probable end-of-turn events.
    3. Speech recognition: Whisper transcribes audio, ideally with language detection and timestamps.
    4. Dialogue orchestration: An LLM, business rules, tools, and conversation state determine the next response.
    5. Speech synthesis: ElevenLabs generates audio from response fragments as they become available.
    6. Playback and control: The client plays audio, detects barge-in, stops playback when needed, and reports telemetry.

    Keep these concerns separate. A telephony adapter should not contain business logic, and the LLM should not be responsible for enforcing payment limits, refund rules, or identity checks. Use a session object containing a conversation ID, caller or user ID, detected language, consent status, tool permissions, and recent turns.

    Choosing and configuring Whisper

    You have two main deployment paths:

    • Managed transcription: Send completed or short audio segments to a hosted Whisper endpoint. This is quick to launch and reduces infrastructure work.
    • Self-hosted transcription: Run Faster-Whisper or another CTranslate2-based implementation on your own GPU or inference service. This can improve control over data handling and unit economics at higher volume, but requires capacity planning and model operations.

    For live conversations, do not wait for an entire recording if your stack supports partial transcripts. Use VAD to segment speech, but retain enough trailing audio to avoid cutting off the final syllable. In Indian deployments, test microphones and codecs rather than relying only on clean studio samples. Background traffic, fans, code-switching, names, local place names, and low-bandwidth phone audio can materially change accuracy.

    Useful transcription practices include:

    • Pass a domain glossary or post-process common names, product terms, and locations.
    • Store confidence scores and timestamps where available.
    • Detect language per turn, but avoid switching voices or prompts on a single uncertain token.
    • Treat transcripts as untrusted input; never execute a tool solely because a caller says “ignore the rules.”
    • Evaluate Hindi, Hinglish, Indian English, and the languages your actual users speak—not just English benchmarks.

    Streaming ElevenLabs output

    ElevenLabs works best when the agent sends complete, speakable phrases rather than individual LLM tokens. Buffer streamed output until you reach punctuation, a natural clause boundary, or a short character threshold. Sending every token creates choppy prosody and excess requests; waiting for the full answer increases perceived latency.

    A practical sequence is:

    1. Start the LLM response in streaming mode.
    2. Buffer text until a sentence or phrase is ready.
    3. Send that segment through the ElevenLabs streaming interface.
    4. Begin playback as soon as the first audio chunk arrives.
    5. Continue synthesising later segments while the first segment plays.

    Choose a voice that is intelligible at telephone bitrate, not merely impressive in a high-quality demo. Test pronunciation of Indian names, acronyms, numbers, dates, rupee amounts, addresses, and mixed Hindi-English phrases. Use SSML or text normalisation where supported, but keep prompts and transformations deterministic for sensitive information.

    Voice cloning requires explicit permission and strong controls. Maintain a voice inventory with ownership, consent records, permitted use, and a revocation process. Do not clone a public figure or an employee’s voice without documented authorisation.

    Turn-taking, interruption, and latency

    A natural agent must know when a user has finished speaking and must yield immediately when the user interrupts. Measure the complete path rather than only model speed:

    • audio capture to end-of-turn detection;
    • end of turn to first transcript;
    • transcript to first LLM token;
    • first LLM token to first ElevenLabs audio byte; and
    • audio arrival to playback.

    Stream each stage, keep services geographically close to users where possible, reuse connections, and avoid unnecessary serial API calls. For Indian users, benchmark your actual cloud region, telecom route, and device mix. A fast model in a distant region may feel slower than a slightly larger model nearby.

    For barge-in, continue listening while audio is playing. When VAD detects a sufficiently long user utterance, stop playback immediately, cancel pending synthesis, and mark the abandoned assistant turn in state. Do not send the entire interrupted answer as if it was completed. A short acknowledgement such as “Go ahead” may help, but avoid adding filler to every interruption.

    Orchestration and tool safety

    The LLM should decide how to phrase an answer, while deterministic code decides whether an action is allowed. For example, a restaurant booking agent may collect date, time, party size, and contact details, then call a booking API that validates availability. A fintech collections agent should authenticate the customer and apply approved scripts before discussing account information.

    Use structured tool schemas, timeouts, retries with idempotency keys, and explicit confirmation before consequential actions. Provide a human handoff when confidence is low, the user requests an agent, authentication fails, or the workflow reaches a regulated or emotionally sensitive boundary. Teams comparing deployment approaches can also review top-rated voice agent services for Indian businesses before building every component internally.

    India-specific product decisions

    Design for the realities of Indian voice traffic:

    • Support code-switching without forcing users into a language menu.
    • Confirm rupee amounts, dates, phone numbers, and addresses digit by digit when errors are costly.
    • Offer DTMF or keypad fallback for noisy calls and users who do not want to speak.
    • Minimise collection of Aadhaar, bank, health, and other sensitive data; redact it from logs.
    • Provide a clear disclosure that the caller is interacting with an AI system.
    • Make escalation routes visible and preserve a concise transcript for the human agent.

    For restaurants, a focused booking flow is easier to evaluate than a general-purpose assistant; compare it with this restaurant table booking voice agent guide for India. For business planning, estimate minutes, retries, telephony, LLM tokens, transcription, synthesis, storage, monitoring, and human escalation—not only the headline API rates. A voice agent pricing and ROI framework can help structure that calculation.

    Evaluation and production checklist

    Before launch, create a test set from real, consented conversations and score both technical and business outcomes:

    • word and entity accuracy for names, numbers, and addresses;
    • time to first audio and full-turn latency;
    • interruption recovery rate;
    • task completion and transfer rate;
    • hallucinated actions or unsupported claims;
    • call drop, timeout, and retry rates;
    • user satisfaction by language, device, and network type; and
    • cost per successful resolution.

    Run adversarial tests for prompt injection, replayed audio, impersonation, abusive callers, silent calls, overlapping speakers, and tool failures. Encrypt audio and transcripts, set retention limits, restrict operator access, and document vendor data-processing terms. Keep a fallback response when Whisper, the LLM, or ElevenLabs is unavailable; a short transfer message is better than silence.

    Recommended build sequence

    Start with one narrow workflow and one channel. Build a text-only orchestration layer first, then add Whisper transcription, ElevenLabs streaming, VAD, and barge-in. Pilot with internal users, inspect failed turns manually, and expand language coverage only after the core workflow is reliable. If the team lacks real-time audio or telephony experience, assess how to hire voice agent developers before committing to a complex in-house build.

    The strongest Whisper-and-ElevenLabs agents are not the ones with the most human-like voice. They are the ones that understand the user, respond quickly, disclose their limits, protect sensitive data, and complete a useful task consistently.

    Last updated 23 September 2026

AIGI may be inaccurate. Replies seeded from the guide above.