The Latency Budget: Engineering Voice Agents That Feel Human
Ask people what makes a voice assistant feel robotic and most will point to the voice itself, the flat intonation, the slightly wrong emphasis. They are half right. But the thing that breaks the illusion fastest is not how the agent sounds. It is how long it waits before it starts talking.
In natural human conversation, the gap between one person finishing and the next beginning is remarkably short, often around 200 milliseconds, and sometimes we even start before the other person is done. That rhythm is deeply wired into us. When a voice agent takes a second and a half to respond, the delay does not read as "thinking." It reads as "broken." Building Vani, our human-grade voice agent, taught us that latency is not a performance metric you optimise at the end. It is the whole game.
Why the pause matters more than the voice
A caller forgives a lot. They will forgive a slightly synthetic timbre, an occasional odd phrasing, even a wrong answer they can correct. What they will not forgive is dead air. Silence on a phone line is ambiguous in the worst way, the caller does not know if the agent is thinking, has misheard, or has hung up. They start talking again, the agent starts talking over them, and the conversation collapses into the awkward dance everyone has had with an automated system.
Users do not experience your model’s accuracy. They experience the silence before it answers.
Where the milliseconds go
A single conversational turn is not one operation, it is a pipeline, and every stage spends from the same budget. If you want the agent to begin replying within a human-feeling window, you have to account for all of it:
- Turn detection, deciding the caller has actually finished speaking, not just paused mid-sentence. Wait too long and the agent feels slow; too little and it interrupts.
- Speech-to-text, transcribing what was said, ideally streaming the transcript as the words arrive rather than waiting for the end.
- Reasoning, the language model deciding what to do and generating a response, often while calling tools or looking up data mid-turn.
- Text-to-speech, turning the response back into audio, which should start playing from the first words rather than after the whole reply is synthesised.
- Network, every hop between caller, telephony, and model adds round-trips that are invisible in a demo on your laptop and very visible on a real phone call.
Engineering under a budget
The core discipline is to treat the end-to-end response time as a fixed budget and refuse to let any single stage overspend. In practice that means a few things we have learned to insist on:
- Stream everything. Transcribe as the caller speaks and begin speaking as the first tokens generate. The goal is to overlap stages, not run them in sequence.
- Start speaking before you have the whole answer. A short, natural opening ("Sure, let me check that for you") buys real compute time while keeping the line alive, exactly as a human would.
- Right-size the model. The fastest reliable model that clears the quality bar beats the most capable model that misses the latency bar. Reasoning you cannot deliver in time is reasoning the caller never hears.
- Handle interruptions as a first-class event. Humans talk over each other and recover gracefully. An agent that cannot be interrupted feels like a recording; one that stops instantly when the caller speaks feels alive.
None of this is about a single clever trick. It is about relentlessly protecting the budget at every stage, because the caller only ever experiences the sum. A voice agent that sounds perfect but answers a beat too late will always feel like a machine. One that answers in the rhythm of real conversation earns something rarer than accuracy, it earns the caller forgetting, for a moment, that they are talking to software at all.
Want to go deeper?
Talk to the team building this. We'd love to hear about the problems you're trying to solve.
Get in touch →