The ElevenLabs Worldwide Hackathon in London on 11 December 2025 was a useful snapshot of where voice agents are heading: away from scripted demos and towards systems that can listen, respond, interrupt, retrieve information, and act in real time. The strongest ideas were not defined by voice quality alone. They combined responsive interaction design with clear product utility.
For Indian builders, the relevance is immediate. Voice remains a natural interface for users who are more comfortable speaking than typing, while multilingual and code-switched conversations are common across support, commerce, education, healthcare, and field operations. This recap extracts the most reusable patterns from the event and turns them into an implementation guide for teams building in 2026.
What made the strongest projects stand out
A conventional voice assistant follows a simple chain: speech-to-text, language-model response, then text-to-speech. That architecture is easy to prototype but often feels slow and brittle. Users notice pauses, repeated confirmations, incorrect turn-taking, and audio that continues after they have started speaking.
The more convincing projects treated conversation as a real-time control system, not a sequence of API calls. Their priorities were:
- Start speaking as soon as a safe response is available.
- Stop speaking immediately when the user interrupts.
- Separate fast interaction decisions from slower reasoning.
- Preserve context without sending an entire conversation to every model.
- Make failure visible and recoverable instead of pretending the agent understood.
Builders new to voice can begin with this Whisper and ElevenLabs voice-agent walkthrough, then add the production patterns below one at a time.
Pattern 1: Stream every stage of the conversation
The best agents reduced perceived latency by streaming wherever possible. The system did not wait for a complete transcript, a complete model answer, or a complete audio file before moving to the next step.
A practical pipeline looks like this:
1. Capture microphone audio in short frames.
2. Detect speech activity and partial transcription.
3. Send a stable utterance or intent signal to the reasoning layer.
4. Begin generating a response in incremental text chunks.
5. Convert approved chunks into audio and play them immediately.
6. Continue reasoning in the background only when necessary.
Measure latency by stages rather than relying on one end-to-end number. Useful metrics include time to speech detection, final-transcript delay, time to first model token, time to first audio byte, and time from user barge-in to playback stop. A system that starts talking in 600 milliseconds but takes two seconds to stop is still frustrating.
For Indian deployments, test latency on ordinary mobile networks and lower-cost Android hardware. A design that works only on a fast laptop connection is not ready for a Tier 2 or Tier 3 user base.
Pattern 2: Make interruption a first-class feature
Natural conversation is turn-taking, not push-to-talk. Users interrupt to correct a detail, ask a follow-up, or signal that they already know the next step. High-quality agents therefore need three separate controls:
- Voice activity detection: identify when the user has started speaking.
- Playback cancellation: stop queued and currently buffered audio.
- Conversation reconciliation: decide whether the unfinished response should be discarded, summarised, or retained as internal context.
Use a small event-driven state machine rather than scattered conditionals. States such as listening, thinking, speaking, interrupted, and recovering make it easier to debug race conditions. Log every transition with timestamps; without this, teams often blame the language model for problems caused by buffering or WebSocket management.
Backchanneling also needs restraint. Short acknowledgements such as “okay” or “right” can make an agent feel attentive, but they should not overlap with important user speech or add cost to every turn. Treat them as a product decision, not a cosmetic voice setting.
Pattern 3: Put a fast orchestrator in front of the main model
A recurring lesson was to avoid using a large model for every event. A lightweight classifier, rules engine, or small language model can handle routine decisions such as greetings, silence, confirmation, interruption, and known commands. The main model can then focus on reasoning, retrieval, and tool use.
A useful tiered design is:
- Fast path: speech activity, intent routing, basic acknowledgements, authentication checks, and safe exits.
- Reasoning path: complex questions, retrieval-augmented generation, planning, and tool calls.
- Voice path: pronunciation, pacing, audio streaming, and interruption handling.
This structure improves both responsiveness and cost control. It also supports graceful degradation: if the reasoning service is slow, the agent can acknowledge the request, provide a status update, or offer a callback rather than remaining silent.
For Indian products, route high-frequency intents locally where practical. Appointment status, order tracking, loan-document checklists, and simple FAQs rarely need an expensive model. Keep escalation available for ambiguous language, mixed Hindi-English input, sensitive domains, and requests requiring human review.
Pattern 4: Add multimodal context only when it solves a real problem
Voice becomes substantially more useful when the agent can inspect an image, screen, document, or device state. A support agent might analyse a machine panel; an education product might discuss a photographed worksheet; a field-sales assistant might extract information from an invoice.
The winning design principle is ground the response in observable evidence. Ask the vision system for structured facts—such as error codes, labels, or visible components—then pass only the relevant facts to the conversational model. Avoid sending raw images repeatedly when a compact representation will do.
Multimodal systems introduce privacy and reliability risks. Tell users when a camera or document is being processed, minimise retention, redact sensitive fields, and require confirmation before taking consequential actions. In India, this matters particularly for healthcare, finance, education records, and identity documents.
Pattern 5: Treat prosody as product logic, not decoration
Voice agents need more than a pleasant default voice. The system should adapt pacing, pauses, verbosity, and confirmation style to the task and the user’s state. Frustration detection can trigger shorter responses and a human handoff; a complex instruction may require slower delivery and explicit checkpoints.
Do not infer emotion too confidently from speech alone. Accent, language, microphone quality, disability, and cultural communication styles can distort sentiment signals. Use emotional classification as a weak signal, combine it with explicit user feedback, and let users correct the agent.
For multilingual Indian products, prioritise comprehensibility over imitation. Test Hindi-English, Tamil-English, Bengali-English, and other code-switched flows with native speakers. Check names, numbers, addresses, dates, and domain vocabulary separately. A technically fluent sentence can still fail if the agent pronounces a locality or rupee amount incorrectly.
A practical stack for a 2026 prototype
A lean build can use a browser or mobile audio client, WebRTC or WebSocket transport, streaming speech recognition, an orchestration service, a retrieval layer, and ElevenLabs for speech generation. LiveKit or Daily can help with real-time media, while a conventional database can often handle session state before a vector database is justified.
Keep these boundaries explicit:
- Session memory: what was said in the current call.
- User memory: durable preferences retained with consent.
- Knowledge base: source documents that can be updated independently.
- Action layer: authenticated tools such as booking, payment, or CRM updates.
Build evaluation before adding features. Create a test set covering interruptions, silence, background noise, accents, code-switching, ambiguous requests, unsafe requests, and tool failures. Score task completion, factual accuracy, first-response latency, interruption recovery, escalation quality, and cost per completed interaction.
Teams exploring the ecosystem can also review voice AI themes from London in 2026 and use the AI hackathons and grants guide for beginners to find suitable build opportunities.
What Indian founders should build next
The strongest opportunity is not a generic voice chatbot. It is a focused agent with a measurable workflow: verifying a delivery, guiding a technician, helping a student practise, collecting a field report, or assisting a customer through a regional-language process.
Start with one job, one user group, and a narrow set of authorised actions. Design for unreliable networks, noisy environments, shared devices, accents, and easy human escalation. Then publish evidence from real calls rather than relying on a polished demo.
The London hackathon’s durable lesson is straightforward: the winning voice experience is engineered around timing, recovery, and trust. Indian teams that combine those fundamentals with local language knowledge can build agents that are not merely impressive to watch, but useful at scale.