ai contact center · stub

Latency and turn-taking in voice AI agents

Verified 2026-10-02 · 56 sources · tier 1–4

Also known as voice agent turn detection, voice bot response latency.

Stub. This topic has 56 sources and no published article. The sources below are everything recorded so far.

So far, no summary has been generated for this topic. The sources below are everything recorded so far.

See also

Related to

Referenced by

Sources

  1. 1
    Azure Speech Speech_SegmentationSilenceTimeoutMs, the in-phrase silence before a phrase is considered done, defaults to 500 ms with an allowed range of 100 to 5000 ms; higher values delay results.
    How to recognize speech · Change how silence is handled, segmentation silence timeout · Checked 2026-10-02
  2. 2
    For Amazon Lex (Classic) bots used from Amazon Connect, barge-in is disabled globally by default and must be enabled with x-amz-lex:barge-in-enabled; that attribute does not control DTMF barge-in.
    Flow block in Connect Customer: Get customer input · Barge-in configuration and usage for Amazon Lex, Amazon Lex (Classic) tab · Checked 2026-10-02
  3. 3
    In an Amazon Connect Get customer input block calling an Amazon Lex V2 bot, the End Silence Threshold session attribute x-amz-lex:audio:end-timeout-ms defaults to 600 ms of silence before the utterance is treated as finished.
    Flow block in Connect Customer: Get customer input · Configurable time-outs for voice input, Amazon Lex tab, End Silence Threshold · Checked 2026-10-02
  4. 4
    In Amazon Connect with Lex V2, Start Silence Threshold (x-amz-lex:audio:start-timeout-ms) defaults to 3000 ms and Max Speech Duration (x-amz-lex:audio:max-length-ms) defaults to 12000 ms with a 55000 ms maximum.
    Flow block in Connect Customer: Get customer input · Configurable time-outs for voice input, Amazon Lex tab · Checked 2026-10-02
  5. 5
    For Amazon Lex V2 bots used from Amazon Connect, barge-in is enabled globally by default and can be changed in the Lex V2 console or with the x-amz-lex:allow-interrupt session attribute.
    Flow block in Connect Customer: Get customer input · Barge-in configuration and usage for Amazon Lex, Amazon Lex tab · Checked 2026-10-02
  6. 6
    Deepgram streaming endpointing is enabled by default with a 10 millisecond silence window and can be changed with the endpointing parameter in milliseconds or disabled with endpointing=false.
    Endpointing · Endpointing overview; Enable feature · Checked 2026-10-02
  7. 7
    Deepgram's Flux documentation states approximately 260 ms end-of-turn detection latency without describing the measurement method on that page.
    Flux quickstart · Introduction · Checked 2026-10-02
  8. 8
    Deepgram strongly recommends sending Flux audio in 80 ms chunks, and Flux accepts mulaw and alaw encodings and an 8000 Hz sample rate, and uses only the /v2/listen endpoint.
    Flux quickstart · Audio specifications; API requirements · Checked 2026-10-02
  9. 9
    Deepgram Flux can emit EagerEndOfTurn so an application starts drafting an LLM response early and TurnResumed if the user keeps talking so the draft is cancelled; this requires setting eager_eot_threshold (0.3 to 0.9, off by default).
    Flux quickstart · Turn events; Configuration parameters table · Checked 2026-10-02
  10. 10
    Deepgram Flux eot_threshold, the confidence needed to emit EndOfTurn, ranges 0.5 to 1.0 with default 0.7, and eot_timeout_ms forces EndOfTurn after 500 to 60000 ms of silence with default 5000.
    Flux quickstart · Configuration parameters table · Checked 2026-10-02
  11. 11
    Deepgram interim results, which stream preliminary transcripts that may be revised before is_final is true, are disabled by default and enabled with interim_results=true.
    Interim Results · Interim Results overview; Enable feature · Checked 2026-10-02
  12. 12
    Deepgram documents that background noise such as music or a ringing phone can keep its VAD triggered so audio-based endpointing never sees silence, and offers UtteranceEnd, based on gaps in word timings, for that case.
    Understanding end of speech detection while live streaming · UtteranceEnd section · Checked 2026-10-02
  13. 13
    When Deepgram endpointing detects an endpoint it finalizes the transcript and sets speech_final to true, which is distinct from is_final.
    Endpointing · Results · Checked 2026-10-02
  14. 14
    Deepgram advises setting utterance_end_ms to 1000 ms or higher.
    Understanding end of speech detection while live streaming · UtteranceEnd section, configuration · Checked 2026-10-02
  15. 15
    Using Deepgram utterance_end_ms requires interim_results=true, and Deepgram recommends acting on the first of speech_final=true or an UtteranceEnd message not preceded by speech_final.
    Understanding end of speech detection while live streaming · Using endpointing and UtteranceEnd together · Checked 2026-10-02
  16. 16
    Enabling Dialogflow CX barge-in lets callers interrupt response audio but increases billable duration, because input and output seconds are billed simultaneously while the agent listens during its own speech.
    Advanced speech settings · Barge-in · Checked 2026-10-02
  17. 17
    Dialogflow CX End of speech sensitivity ranges from 0 (less likely to end speech) to 100 (more likely to end speech), and an advanced timeout-based option turns that value into a relative silence timeout.
    Advanced speech settings · End of speech sensitivity; Advanced timeout-based end of speech sensitivity · Checked 2026-10-02
  18. 18
    Dialogflow CX No speech timeout, after which it stops waiting for caller audio and raises a no-input event, defaults to 5 seconds with a 60 second maximum.
    Advanced speech settings · No speech timeout · Checked 2026-10-02
  19. 19
    ElevenLabs states its Flash TTS models deliver about 75 ms inference and that this figure is model inference time only, with end-to-end latency varying by location and endpoint type.
    Latency optimization · Use Flash models · Checked 2026-10-02
  20. 20
    Vendor-published latency figures are not directly comparable because they measure different spans, for example ElevenLabs' 75 ms is model inference only while Twilio's turn gap targets run from end of user speech to audio at the caller's ear.inferred
    Latency optimization · Use Flash models (caveat sentence) · Checked 2026-10-02
  21. 21
    A PSTN-fed voice agent receives 8 kHz G.711 narrowband audio, so any recognizer or realtime model that does not accept 8 kHz G.711 natively needs transcoding and upsampling in the media path, which cannot restore bandwidth lost on the call leg.inferred
    Media Streams - WebSocket Messages · Start message, mediaFormat · Checked 2026-10-02
  22. 22
    Silence-based end-of-turn timers set a floor on bot response time: with Amazon Connect's default 600 ms Lex end silence, the bot cannot begin replying until at least that silence has elapsed, before STT, LLM and TTS time is added.inferred
    Flow block in Connect Customer: Get customer input · Configurable time-outs for voice input, End Silence Threshold definition · Checked 2026-10-02
  23. 23
    ITU-T G.114 (05/2003, in force) gives one-way transmission time planning guidance in which about 150 ms is broadly acceptable and 400 ms is the general upper planning limit.
    G.114 : One-way transmission time · Clause 4 (per summarised fetch; not confirmed against extracted text) · Checked 2026-10-02
  24. 24
    An Amazon Lex V2 bot on a bidirectional audio stream treats user input that arrives before the application sends PlaybackCompletion as an interruption and sends a PlaybackInterruptionEvent; interruptibility can be turned off per slot prompt.
    Enabling your Amazon Lex V2 bot to be interrupted by the user · Introduction paragraphs; console procedure step 9 · Checked 2026-10-02
  25. 25
    LiveKit Agents defaults to a turn detector model that predicts end of turn from speech meaning and acoustics on top of VAD, and suggests VAD-only detection when minimal latency or an unsupported language matters.
    Turn detection and interruptions · Turn detection modes · Checked 2026-10-02
  26. 26
    OpenAI semantic_vad has an eagerness setting of low, medium, high or auto, where auto is the default and equals medium; low lets the user take more time and high chunks audio as soon as possible.
    Voice activity detection (VAD) · Semantic VAD, eagerness · Checked 2026-10-02
  27. 27
    The OpenAI Realtime API accepts G.711 audio directly as audio/pcmu or audio/pcma at 8 kHz, alongside 16-bit PCM at 24 kHz (the default) or 16 kHz.
    Realtime API with WebSocket · Audio formats · Checked 2026-10-02
  28. 28
    In the OpenAI Realtime API, when the server emits input_audio_buffer.speech_started while a response is in progress, the server automatically cancels the in-progress response and emits response.cancelled.
    Realtime conversations · Interruption and truncation · Checked 2026-10-02
  29. 29
    OpenAI semantic_vad decides end of turn from what the user has said rather than from silence duration alone, so it can wait longer when an utterance sounds unfinished.
    Voice activity detection (VAD) · Semantic VAD · Checked 2026-10-02
  30. 30
    server_vad is the default turn detection mode for OpenAI Realtime sessions that support turn detection.
    Voice activity detection (VAD) · Overview · Checked 2026-10-02
  31. 31
    On a WebSocket Realtime connection the client manages playback, so after an interruption it must stop playback and send conversation.item.truncate with audio_end_ms so the model does not assume the caller heard audio that was never played.
    Realtime conversations · Interruption and truncation, WebSocket handling · Checked 2026-10-02
  32. 32
    The OpenAI VAD guide's server_vad configuration example uses threshold 0.5, prefix_padding_ms 300 and silence_duration_ms 500, where silence_duration_ms is the silence needed to detect speech stop and prefix_padding_ms is audio kept from before detected speech.
    Voice activity detection (VAD) · Server VAD, example session configuration and parameter list · Checked 2026-10-02
  33. 33
    The OpenAI Realtime API offers two automatic turn detection modes: server_vad, which chunks audio on periods of silence, and semantic_vad, which uses a semantic classifier on the spoken content to judge when the user has finished.
    Voice activity detection (VAD) · Overview; Server VAD; Semantic VAD · Checked 2026-10-02
  34. 34
    On a WebRTC Realtime connection the server keeps an output audio buffer, automatically truncates unplayed audio on user interruption, and accepts output_audio_buffer.clear to discard unplayed audio.
    Realtime conversations · Interruption and truncation, WebRTC handling · Checked 2026-10-02
  35. 35
    Pipecat Smart Turn v3.2 is a BSD 2-Clause open audio turn detection model that takes up to 8 seconds of 16 kHz mono PCM and runs alongside a lightweight VAD such as Silero, only during silences.
    pipecat-ai/smart-turn README · README, overview and how it works · Checked 2026-10-02
  36. 36
    The Twilio latency guide recommends co-locating STT, TTS and orchestration near the media edge because each cross-region hop adds delay that compounds across the pipeline.
    Core Latency in AI Voice Agents · Service architecture section · Checked 2026-10-02
  37. 37
    The Twilio latency guide budgets LLM time to first token at 375 ms target and 750 ms upper limit.
    Core Latency in AI Voice Agents · Latency targets table, LLM TTFT row · Checked 2026-10-02
  38. 38
    A Twilio engineering guide sets a mouth-to-ear turn gap target of 1115 ms with a 1400 ms upper limit, and a platform turn gap excluding internet and PSTN transit of 885 ms target and 1100 ms limit.
    Core Latency in AI Voice Agents · Latency targets table · Checked 2026-10-02
  39. 39
    The Twilio latency guide budgets speech-to-text at 350 ms target and 500 ms upper limit.
    Core Latency in AI Voice Agents · Latency targets table, Speech-to-Text row · Checked 2026-10-02
  40. 40
    The Twilio latency guide budgets text-to-speech time to first byte at 100 ms target and 250 ms upper limit.
    Core Latency in AI Voice Agents · Latency targets table, Text-to-Speech TTFB row · Checked 2026-10-02
  41. 41
    Twilio ConversationRelay's default transcriptionProvider is Deepgram, with Google as the alternative.
    TwiML Voice: <ConversationRelay> · Attributes table, transcriptionProvider · Checked 2026-10-02
  42. 42
    Twilio ConversationRelay's default ttsProvider is ElevenLabs, with Google and Amazon as alternatives.
    TwiML Voice: <ConversationRelay> · Attributes table, ttsProvider · Checked 2026-10-02
  43. 43
    Twilio ConversationRelay eotThreshold, the confidence required to finish a turn, accepts 0.5 to 0.9 and defaults to 0.8.
    TwiML Voice: <ConversationRelay> · Attributes table, eotThreshold · Checked 2026-10-02
  44. 44
    Twilio ConversationRelay ignoreBackchannel, which stops short feedback such as 'uh-huh' from interrupting the agent, defaults to false.
    TwiML Voice: <ConversationRelay> · Attributes table, ignoreBackchannel · Checked 2026-10-02
  45. 45
    Twilio ConversationRelay interruptible accepts none, dtmf, speech or any and defaults to any, so by default caller speech or DTMF stops TTS playback.
    TwiML Voice: <ConversationRelay> · Attributes table, interruptible · Checked 2026-10-02
  46. 46
    Twilio ConversationRelay interruptSensitivity (high, medium or low) controls how easily caller speech interrupts TTS based on recognition confidence and input length, and defaults to high, the most easily triggered setting.
    TwiML Voice: <ConversationRelay> · Attributes table, interruptSensitivity · Checked 2026-10-02
  47. 47
    The Twilio dashboard breaks virtual agent response time into network round trip to the developer WebSocket, speech-to-text finalization, application time to first token, and text-to-speech time to first audio.
    Voice Insights Conversation Relay Insights Dashboard · Virtual agent response time · Checked 2026-10-02
  48. 48
    Twilio's Conversation Relay Insights dashboard defines Time to First Audio as the average time for the virtual agent to begin speaking after the customer finishes talking.
    Voice Insights Conversation Relay Insights Dashboard · Latency metrics, Time to First Audio · Checked 2026-10-02
  49. 49
    Twilio states that high rates of callers interrupting the virtual agent may indicate high conversation latency or an overly verbose agent.
    Voice Insights Conversation Relay Insights Dashboard · Interruption metrics · Checked 2026-10-02
  50. 50
    To stop bot audio on barge-in over Twilio Media Streams, the application sends a clear message, which empties all buffered audio and causes pending mark messages to be returned.
    Media Streams - WebSocket Messages · Send a clear message · Checked 2026-10-02
  51. 51
    Twilio Media Streams always delivers audio as audio/x-mulaw at 8000 Hz on one channel.
    Media Streams - WebSocket Messages · Start message, mediaFormat · Checked 2026-10-02
  52. 52
    Azure Voice Live turn detection defaults to server_vad with silence_duration_ms 500 and threshold 0.5, and also offers azure_semantic_vad and azure_semantic_vad_multilingual for all models.
    How to use the Voice Live API · Conversational enhancements, Turn Detection Parameters table (type, threshold, silence_duration_ms rows) · Checked 2026-10-02
  53. 53
    Voice Live server echo cancellation with the default server reference assumes the client plays response audio immediately, and quality degrades if playback is delayed more than two seconds.
    How to use the Voice Live API · Noise suppression and echo cancellation, Note · Checked 2026-10-02
  54. 54
    Voice Live remove_filler_words (default false) ignores English fillers such as um, uh and hmm during an ongoing response to reduce false barge-in, and assumes the client plays response audio as soon as it arrives.
    How to use the Voice Live API · Turn Detection Parameters table, remove_filler_words row · Checked 2026-10-02
  55. 55
    Voice Live input_audio_sampling_rate supports 16000 and 24000 with a default of 24000.
    How to use the Voice Live API · Input audio properties table, input_audio_sampling_rate · Checked 2026-10-02
  56. 56
    From Voice Live API version 2026-04-10 the prefix_padding_ms default is 400 for server_vad and 420 for the Azure semantic VAD types; earlier API versions default to 300 for all types.
    How to use the Voice Live API · Turn Detection Parameters table, prefix_padding_ms row · Checked 2026-10-02

Documents

tier 1 standards and regulators

G.114 : One-way transmission time

International Telecommunication Union (ITU-T) · 2003-05-07 · accessed 2026-10-02

tier 2 current vendor documentation

Advanced speech settings

Google Cloud · 2026-09-30 · accessed 2026-10-02

tier 2 current vendor documentation

Enabling your Amazon Lex V2 bot to be interrupted by the user

Amazon Web Services · accessed 2026-09-25

tier 2 current vendor documentation

Endpointing

Deepgram · accessed 2026-10-02

tier 2 current vendor documentation

Flow block in Connect Customer: Get customer input

Amazon Web Services · accessed 2026-09-25

tier 2 current vendor documentation

Flux quickstart

Deepgram · accessed 2026-10-02

tier 2 current vendor documentation

How to recognize speech

Microsoft · accessed 2026-10-02

tier 2 current vendor documentation

How to use the Voice Live API

Microsoft · 2026-09-24 · accessed 2026-10-02

tier 2 current vendor documentation

Interim Results

Deepgram · accessed 2026-10-02

tier 2 current vendor documentation

Latency optimization

ElevenLabs · accessed 2026-10-02

tier 2 current vendor documentation

Media Streams - WebSocket Messages

Twilio · accessed 2026-10-02

tier 2 current vendor documentation

Realtime API with WebSocket

OpenAI · accessed 2026-10-02

tier 2 current vendor documentation

Realtime conversations

OpenAI · accessed 2026-10-02

tier 2 current vendor documentation

Turn detection and interruptions

LiveKit · accessed 2026-10-02

tier 2 current vendor documentation

TwiML Voice: <ConversationRelay>

Twilio · accessed 2026-10-02

tier 2 current vendor documentation

Understanding end of speech detection while live streaming

Deepgram · accessed 2026-10-02

tier 2 current vendor documentation

Voice activity detection (VAD)

OpenAI · accessed 2026-10-02

tier 2 current vendor documentation

Voice Insights Conversation Relay Insights Dashboard

Twilio · 2026-08-24 · accessed 2026-10-02

tier 3 vendor-maintained repositories

pipecat-ai/smart-turn README

Pipecat (Daily) · accessed 2026-10-02

tier 4 archived vendor documentation

Core Latency in AI Voice Agents

Twilio · 2025-11-17 · accessed 2026-10-02

Cite this page

APA

WarmTransfer. (2026, October 2). Latency and turn-taking in voice AI agents. WarmTransfer. https://warmtransfer.net/knowledge/voice-ai-agent-latency

BibTeX

@misc{warmtransfer-voice-ai-agent-latency,
  title  = {Latency and turn-taking in voice AI agents},
  author = {{WarmTransfer}},
  year   = {2026},
  url    = {https://warmtransfer.net/knowledge/voice-ai-agent-latency},
  note   = {Verified 2026-10-02}
}