ai contact center · stub
Latency and turn-taking in voice AI agents
Verified 2026-10-02 · 56 sources · tier 1–4
Also known as voice agent turn detection, voice bot response latency.
Stub. This topic has 56 sources and no published article. The sources below are everything recorded so far.
So far, no summary has been generated for this topic. The sources below are everything recorded so far.
See also
Related to
- Troubleshooting audio delay and talk-over on voice calls — Network and PSTN delay add to bot response latency; network troubleshooting is covered there not here
- Virtual agents and conversational IVR design — Turn detection and barge-in settings are configuration surfaces of the virtual agents covered there
Referenced by
- Real-time audio streaming APIs for contact center AIstub — Streaming transport and codec choices feed the latency budget of voice AI agents; no vendor in this packet publishes a streaming latency figure
- Troubleshooting audio delay and talk-over on voice calls — Network delay adds to bot response latency; not covered here
Sources
- 1Azure Speech Speech_SegmentationSilenceTimeoutMs, the in-phrase silence before a phrase is considered done, defaults to 500 ms with an allowed range of 100 to 5000 ms; higher values delay results.How to recognize speech · Change how silence is handled, segmentation silence timeout · Checked 2026-10-02
- 2For Amazon Lex (Classic) bots used from Amazon Connect, barge-in is disabled globally by default and must be enabled with x-amz-lex:barge-in-enabled; that attribute does not control DTMF barge-in.Flow block in Connect Customer: Get customer input · Barge-in configuration and usage for Amazon Lex, Amazon Lex (Classic) tab · Checked 2026-10-02
- 3In an Amazon Connect Get customer input block calling an Amazon Lex V2 bot, the End Silence Threshold session attribute x-amz-lex:audio:end-timeout-ms defaults to 600 ms of silence before the utterance is treated as finished.Flow block in Connect Customer: Get customer input · Configurable time-outs for voice input, Amazon Lex tab, End Silence Threshold · Checked 2026-10-02
- 4In Amazon Connect with Lex V2, Start Silence Threshold (x-amz-lex:audio:start-timeout-ms) defaults to 3000 ms and Max Speech Duration (x-amz-lex:audio:max-length-ms) defaults to 12000 ms with a 55000 ms maximum.Flow block in Connect Customer: Get customer input · Configurable time-outs for voice input, Amazon Lex tab · Checked 2026-10-02
- 5For Amazon Lex V2 bots used from Amazon Connect, barge-in is enabled globally by default and can be changed in the Lex V2 console or with the x-amz-lex:allow-interrupt session attribute.Flow block in Connect Customer: Get customer input · Barge-in configuration and usage for Amazon Lex, Amazon Lex tab · Checked 2026-10-02
- 6Deepgram streaming endpointing is enabled by default with a 10 millisecond silence window and can be changed with the endpointing parameter in milliseconds or disabled with endpointing=false.Endpointing · Endpointing overview; Enable feature · Checked 2026-10-02
- 7Deepgram's Flux documentation states approximately 260 ms end-of-turn detection latency without describing the measurement method on that page.Flux quickstart · Introduction · Checked 2026-10-02
- 8Deepgram strongly recommends sending Flux audio in 80 ms chunks, and Flux accepts mulaw and alaw encodings and an 8000 Hz sample rate, and uses only the /v2/listen endpoint.Flux quickstart · Audio specifications; API requirements · Checked 2026-10-02
- 9Deepgram Flux can emit EagerEndOfTurn so an application starts drafting an LLM response early and TurnResumed if the user keeps talking so the draft is cancelled; this requires setting eager_eot_threshold (0.3 to 0.9, off by default).Flux quickstart · Turn events; Configuration parameters table · Checked 2026-10-02
- 10Deepgram Flux eot_threshold, the confidence needed to emit EndOfTurn, ranges 0.5 to 1.0 with default 0.7, and eot_timeout_ms forces EndOfTurn after 500 to 60000 ms of silence with default 5000.Flux quickstart · Configuration parameters table · Checked 2026-10-02
- 11Deepgram interim results, which stream preliminary transcripts that may be revised before is_final is true, are disabled by default and enabled with interim_results=true.Interim Results · Interim Results overview; Enable feature · Checked 2026-10-02
- 12Deepgram documents that background noise such as music or a ringing phone can keep its VAD triggered so audio-based endpointing never sees silence, and offers UtteranceEnd, based on gaps in word timings, for that case.Understanding end of speech detection while live streaming · UtteranceEnd section · Checked 2026-10-02
- 13When Deepgram endpointing detects an endpoint it finalizes the transcript and sets speech_final to true, which is distinct from is_final.Endpointing · Results · Checked 2026-10-02
- 14Deepgram advises setting utterance_end_ms to 1000 ms or higher.Understanding end of speech detection while live streaming · UtteranceEnd section, configuration · Checked 2026-10-02
- 15Using Deepgram utterance_end_ms requires interim_results=true, and Deepgram recommends acting on the first of speech_final=true or an UtteranceEnd message not preceded by speech_final.Understanding end of speech detection while live streaming · Using endpointing and UtteranceEnd together · Checked 2026-10-02
- 16Enabling Dialogflow CX barge-in lets callers interrupt response audio but increases billable duration, because input and output seconds are billed simultaneously while the agent listens during its own speech.Advanced speech settings · Barge-in · Checked 2026-10-02
- 17Dialogflow CX End of speech sensitivity ranges from 0 (less likely to end speech) to 100 (more likely to end speech), and an advanced timeout-based option turns that value into a relative silence timeout.Advanced speech settings · End of speech sensitivity; Advanced timeout-based end of speech sensitivity · Checked 2026-10-02
- 18Dialogflow CX No speech timeout, after which it stops waiting for caller audio and raises a no-input event, defaults to 5 seconds with a 60 second maximum.Advanced speech settings · No speech timeout · Checked 2026-10-02
- 19ElevenLabs states its Flash TTS models deliver about 75 ms inference and that this figure is model inference time only, with end-to-end latency varying by location and endpoint type.Latency optimization · Use Flash models · Checked 2026-10-02
- 20Vendor-published latency figures are not directly comparable because they measure different spans, for example ElevenLabs' 75 ms is model inference only while Twilio's turn gap targets run from end of user speech to audio at the caller's ear.inferredLatency optimization · Use Flash models (caveat sentence) · Checked 2026-10-02
- 21A PSTN-fed voice agent receives 8 kHz G.711 narrowband audio, so any recognizer or realtime model that does not accept 8 kHz G.711 natively needs transcoding and upsampling in the media path, which cannot restore bandwidth lost on the call leg.inferredMedia Streams - WebSocket Messages · Start message, mediaFormat · Checked 2026-10-02
- 22Silence-based end-of-turn timers set a floor on bot response time: with Amazon Connect's default 600 ms Lex end silence, the bot cannot begin replying until at least that silence has elapsed, before STT, LLM and TTS time is added.inferredFlow block in Connect Customer: Get customer input · Configurable time-outs for voice input, End Silence Threshold definition · Checked 2026-10-02
- 23ITU-T G.114 (05/2003, in force) gives one-way transmission time planning guidance in which about 150 ms is broadly acceptable and 400 ms is the general upper planning limit.G.114 : One-way transmission time · Clause 4 (per summarised fetch; not confirmed against extracted text) · Checked 2026-10-02
- 24An Amazon Lex V2 bot on a bidirectional audio stream treats user input that arrives before the application sends PlaybackCompletion as an interruption and sends a PlaybackInterruptionEvent; interruptibility can be turned off per slot prompt.Enabling your Amazon Lex V2 bot to be interrupted by the user · Introduction paragraphs; console procedure step 9 · Checked 2026-10-02
- 25LiveKit Agents defaults to a turn detector model that predicts end of turn from speech meaning and acoustics on top of VAD, and suggests VAD-only detection when minimal latency or an unsupported language matters.Turn detection and interruptions · Turn detection modes · Checked 2026-10-02
- 26OpenAI semantic_vad has an eagerness setting of low, medium, high or auto, where auto is the default and equals medium; low lets the user take more time and high chunks audio as soon as possible.Voice activity detection (VAD) · Semantic VAD, eagerness · Checked 2026-10-02
- 27The OpenAI Realtime API accepts G.711 audio directly as audio/pcmu or audio/pcma at 8 kHz, alongside 16-bit PCM at 24 kHz (the default) or 16 kHz.Realtime API with WebSocket · Audio formats · Checked 2026-10-02
- 28In the OpenAI Realtime API, when the server emits input_audio_buffer.speech_started while a response is in progress, the server automatically cancels the in-progress response and emits response.cancelled.Realtime conversations · Interruption and truncation · Checked 2026-10-02
- 29OpenAI semantic_vad decides end of turn from what the user has said rather than from silence duration alone, so it can wait longer when an utterance sounds unfinished.Voice activity detection (VAD) · Semantic VAD · Checked 2026-10-02
- 30server_vad is the default turn detection mode for OpenAI Realtime sessions that support turn detection.Voice activity detection (VAD) · Overview · Checked 2026-10-02
- 31On a WebSocket Realtime connection the client manages playback, so after an interruption it must stop playback and send conversation.item.truncate with audio_end_ms so the model does not assume the caller heard audio that was never played.Realtime conversations · Interruption and truncation, WebSocket handling · Checked 2026-10-02
- 32The OpenAI VAD guide's server_vad configuration example uses threshold 0.5, prefix_padding_ms 300 and silence_duration_ms 500, where silence_duration_ms is the silence needed to detect speech stop and prefix_padding_ms is audio kept from before detected speech.Voice activity detection (VAD) · Server VAD, example session configuration and parameter list · Checked 2026-10-02
- 33The OpenAI Realtime API offers two automatic turn detection modes: server_vad, which chunks audio on periods of silence, and semantic_vad, which uses a semantic classifier on the spoken content to judge when the user has finished.Voice activity detection (VAD) · Overview; Server VAD; Semantic VAD · Checked 2026-10-02
- 34On a WebRTC Realtime connection the server keeps an output audio buffer, automatically truncates unplayed audio on user interruption, and accepts output_audio_buffer.clear to discard unplayed audio.Realtime conversations · Interruption and truncation, WebRTC handling · Checked 2026-10-02
- 35Pipecat Smart Turn v3.2 is a BSD 2-Clause open audio turn detection model that takes up to 8 seconds of 16 kHz mono PCM and runs alongside a lightweight VAD such as Silero, only during silences.pipecat-ai/smart-turn README · README, overview and how it works · Checked 2026-10-02
- 36The Twilio latency guide recommends co-locating STT, TTS and orchestration near the media edge because each cross-region hop adds delay that compounds across the pipeline.Core Latency in AI Voice Agents · Service architecture section · Checked 2026-10-02
- 37The Twilio latency guide budgets LLM time to first token at 375 ms target and 750 ms upper limit.Core Latency in AI Voice Agents · Latency targets table, LLM TTFT row · Checked 2026-10-02
- 38A Twilio engineering guide sets a mouth-to-ear turn gap target of 1115 ms with a 1400 ms upper limit, and a platform turn gap excluding internet and PSTN transit of 885 ms target and 1100 ms limit.Core Latency in AI Voice Agents · Latency targets table · Checked 2026-10-02
- 39The Twilio latency guide budgets speech-to-text at 350 ms target and 500 ms upper limit.Core Latency in AI Voice Agents · Latency targets table, Speech-to-Text row · Checked 2026-10-02
- 40The Twilio latency guide budgets text-to-speech time to first byte at 100 ms target and 250 ms upper limit.Core Latency in AI Voice Agents · Latency targets table, Text-to-Speech TTFB row · Checked 2026-10-02
- 41Twilio ConversationRelay's default transcriptionProvider is Deepgram, with Google as the alternative.TwiML Voice: <ConversationRelay> · Attributes table, transcriptionProvider · Checked 2026-10-02
- 42Twilio ConversationRelay's default ttsProvider is ElevenLabs, with Google and Amazon as alternatives.TwiML Voice: <ConversationRelay> · Attributes table, ttsProvider · Checked 2026-10-02
- 43Twilio ConversationRelay eotThreshold, the confidence required to finish a turn, accepts 0.5 to 0.9 and defaults to 0.8.TwiML Voice: <ConversationRelay> · Attributes table, eotThreshold · Checked 2026-10-02
- 44Twilio ConversationRelay ignoreBackchannel, which stops short feedback such as 'uh-huh' from interrupting the agent, defaults to false.TwiML Voice: <ConversationRelay> · Attributes table, ignoreBackchannel · Checked 2026-10-02
- 45Twilio ConversationRelay interruptible accepts none, dtmf, speech or any and defaults to any, so by default caller speech or DTMF stops TTS playback.TwiML Voice: <ConversationRelay> · Attributes table, interruptible · Checked 2026-10-02
- 46Twilio ConversationRelay interruptSensitivity (high, medium or low) controls how easily caller speech interrupts TTS based on recognition confidence and input length, and defaults to high, the most easily triggered setting.TwiML Voice: <ConversationRelay> · Attributes table, interruptSensitivity · Checked 2026-10-02
- 47The Twilio dashboard breaks virtual agent response time into network round trip to the developer WebSocket, speech-to-text finalization, application time to first token, and text-to-speech time to first audio.Voice Insights Conversation Relay Insights Dashboard · Virtual agent response time · Checked 2026-10-02
- 48Twilio's Conversation Relay Insights dashboard defines Time to First Audio as the average time for the virtual agent to begin speaking after the customer finishes talking.Voice Insights Conversation Relay Insights Dashboard · Latency metrics, Time to First Audio · Checked 2026-10-02
- 49Twilio states that high rates of callers interrupting the virtual agent may indicate high conversation latency or an overly verbose agent.Voice Insights Conversation Relay Insights Dashboard · Interruption metrics · Checked 2026-10-02
- 50To stop bot audio on barge-in over Twilio Media Streams, the application sends a clear message, which empties all buffered audio and causes pending mark messages to be returned.Media Streams - WebSocket Messages · Send a clear message · Checked 2026-10-02
- 51Twilio Media Streams always delivers audio as audio/x-mulaw at 8000 Hz on one channel.Media Streams - WebSocket Messages · Start message, mediaFormat · Checked 2026-10-02
- 52Azure Voice Live turn detection defaults to server_vad with silence_duration_ms 500 and threshold 0.5, and also offers azure_semantic_vad and azure_semantic_vad_multilingual for all models.How to use the Voice Live API · Conversational enhancements, Turn Detection Parameters table (type, threshold, silence_duration_ms rows) · Checked 2026-10-02
- 53Voice Live server echo cancellation with the default server reference assumes the client plays response audio immediately, and quality degrades if playback is delayed more than two seconds.How to use the Voice Live API · Noise suppression and echo cancellation, Note · Checked 2026-10-02
- 54Voice Live remove_filler_words (default false) ignores English fillers such as um, uh and hmm during an ongoing response to reduce false barge-in, and assumes the client plays response audio as soon as it arrives.How to use the Voice Live API · Turn Detection Parameters table, remove_filler_words row · Checked 2026-10-02
- 55Voice Live input_audio_sampling_rate supports 16000 and 24000 with a default of 24000.How to use the Voice Live API · Input audio properties table, input_audio_sampling_rate · Checked 2026-10-02
- 56From Voice Live API version 2026-04-10 the prefix_padding_ms default is 400 for server_vad and 420 for the Azure semantic VAD types; earlier API versions default to 300 for all types.How to use the Voice Live API · Turn Detection Parameters table, prefix_padding_ms row · Checked 2026-10-02
Documents
tier 1 standards and regulators
G.114 : One-way transmission time
tier 2 current vendor documentation
Advanced speech settings
tier 2 current vendor documentation
Enabling your Amazon Lex V2 bot to be interrupted by the user
tier 2 current vendor documentation
Endpointing
tier 2 current vendor documentation
Flow block in Connect Customer: Get customer input
tier 2 current vendor documentation
Flux quickstart
tier 2 current vendor documentation
How to recognize speech
tier 2 current vendor documentation
How to use the Voice Live API
tier 2 current vendor documentation
Interim Results
tier 2 current vendor documentation
Latency optimization
tier 2 current vendor documentation
Media Streams - WebSocket Messages
tier 2 current vendor documentation
Realtime API with WebSocket
tier 2 current vendor documentation
Realtime conversations
tier 2 current vendor documentation
Turn detection and interruptions
tier 2 current vendor documentation
TwiML Voice: <ConversationRelay>
tier 2 current vendor documentation
Understanding end of speech detection while live streaming
tier 2 current vendor documentation
Voice activity detection (VAD)
tier 2 current vendor documentation
Voice Insights Conversation Relay Insights Dashboard
tier 3 vendor-maintained repositories
pipecat-ai/smart-turn README
tier 4 archived vendor documentation
Core Latency in AI Voice Agents
Cite this page
APA
WarmTransfer. (2026, October 2). Latency and turn-taking in voice AI agents. WarmTransfer. https://warmtransfer.net/knowledge/voice-ai-agent-latency
BibTeX
@misc{warmtransfer-voice-ai-agent-latency,
title = {Latency and turn-taking in voice AI agents},
author = {{WarmTransfer}},
year = {2026},
url = {https://warmtransfer.net/knowledge/voice-ai-agent-latency},
note = {Verified 2026-10-02}
}