All field notes

Voice AI Testing

·

Testing turn detection in voice AI agents

TTS and LLM latency keep falling, so endpointing now sets the ceiling on voice-agent responsiveness. Here is how to test turn detection before you ship.

Daniel Gauci Mizzi

Daniel Gauci Mizzi

Co-founder & CTO @ Roark

9 min read
Testing turn detection in voice AI agents

The latency budget inside a voice agent has been collapsing from the edges inward. Microsoft announced MAI-Voice-2.1-Flash on October 1 at roughly 45ms inference latency. Inception's Mercury Voice does 45 seconds of audio with about 150ms end-to-end latency. Deepgram Flux cut end-of-turn detection to about 260ms P50. The LLM step, especially with preemptive generation, keeps shrinking too.

What has not collapsed is the thing that decides when the model gets to speak. End-of-turn detection, phrase endpointing, the "has the caller actually finished?" question. In 2024 that was a 500ms silence timer. In 2026 it is a stack of VAD, a semantic turn model, STT-side endpointing hints, and a configurable commit delay, and whether it feels right still depends almost entirely on how it behaves with real callers who hesitate, trail off, read phone numbers aloud, or say "uhh" while they think. If TTS is now 45ms, every bad turn-detection decision is visible to the user as either a cut-off or an awkward hang.

This post is about how to test that layer before it ships. The core question is not "which turn detector is best"; it is "does ours handle my callers?" and the only honest way to answer it is to run thousands of calls that resemble the ones you will actually get.

What "turn detection" actually means in 2026

A turn-detection stack usually has three parts that interact:

  1. VAD segments audio into speech and non-speech. Silero VAD is the near-universal default: it processes 30ms chunks in under 1ms on a single CPU thread. It tells you speech is happening, nothing more.
  2. An end-of-utterance model decides whether the current silence is a pause or a terminal boundary. The current wave of these are audio-native: LiveKit Turn Detector v1 "listens to speech directly instead of waiting on a transcript, fusing semantic and acoustic understanding", and Pipecat Smart Turn v3 classifies raw PCM with a Whisper-Tiny backbone at 12ms on CPU. Some STTs now fuse this into recognition: Deepgram Flux delivers turn-complete transcripts with native end-of-turn detection, AssemblyAI's Universal-3 Pro exposes min_turn_silence and max_turn_silence, and OpenAI's Realtime API offers semantic_vad with an eagerness setting alongside the legacy server_vad.
  3. An endpointing delay policy sits on top. LiveKit documents this as min_delay, max_delay, and an optional dynamic mode that adapts within the range to pause patterns. Pipecat exposes the equivalent via its user-turn strategies. These knobs are where most teams actually set the latency vs cutoff tradeoff.
End-of-turn decisions pass through VAD, a semantic model, and a commit-delay policy before the agent can speak
End-of-turn decisions pass through VAD, a semantic model, and a commit-delay policy before the agent can speak

The important observation is that none of these components is deterministic in a way a unit test can cover. A turn detector trained on English conversational data has to behave sensibly on a caller reading back a 17-digit policy number, slowly, with umms. Semantic VAD can score low end-of-turn probability when audio trails off with "uhhm", which is exactly what you want for one caller and exactly wrong for another. The only validation that generalizes is behavioral.

Why the old tests do not catch turn-detection bugs

Three failure modes dominate production incidents I see teams hit, and all three slip past typical pre-launch QA.

False cutoffs. The agent decides the caller is done when they are mid-thought, interrupts, and the caller has to repeat themselves. Scripted eval sets rarely catch this because scripts speak in grammatically complete bursts. Real callers do not.

Dead air. The agent waits too long after the caller actually finishes. On short-utterance patterns like "yes" or "okay," Smart Turn v3.2 specifically targeted a 40% accuracy improvement because short confirmations were a known weakness. If your test suite does not contain short acknowledgements against a backdrop of longer turns, you will not see this.

False barge-in. The agent treats the caller's "mhm" or a background cough as a new turn and either stops talking or starts generating a response. LiveKit added adaptive interruption handling and resume_false_interruption exactly because this failure is frequent enough to need a dedicated remediation path. You cannot assess it by reading transcripts; you need the actual audio and the actual timing.

A text-based eval harness, even a sophisticated one, cannot score any of these because they live in the audio and the timing, not in the words.

A pre-launch test plan that actually finds the regressions

The following is the scaffolding I use when auditing a team's turn-detection setup. Each scenario is a dialed call, generated against a persona, with a pass criterion that lives in the audio.

1. Trailing-off and filled pauses

Caller trails off mid-sentence with "ummm," "so I think it's... yeah, it's," or a long inhale. The pass criterion is simple: the agent does not take the floor until the caller has stopped speaking for the full configured commit delay after the final content word. This single scenario reveals whether your semantic turn detector is actually behaving semantically or has silently fallen back to raw silence.

2. Digit and alphanumeric readback with pauses

"My phone number is five, five, five... two, one, four... one..." The spacing between groups will trip pure-silence endpointing every time. If you are on a fixed endpointing policy with min_delay=500ms, you will cut the caller off between groups. If you are on dynamic endpointing, you will see whether the per-session pause-statistics adaptation actually catches up in time.

3. Short confirmations

"Yes." "Sure." "Correct." A persona that answers in one-word confirmations against an agent built for longer turns will expose the opposite failure: the agent sits silent for 1.5 seconds after a 200ms "yes." Measure the lag.

4. Caller interrupts the agent

Have the persona start speaking while the agent is in the middle of a disclosure. The pass criterion has two parts: the agent must cut TTS within the configured interrupt threshold, and it must not treat a short "mhm" backchannel as an interrupt. This is where min_duration and resume_false_interruption get stress-tested.

5. Background noise

Run the same scripts in a quiet environment, in traffic noise, and with a TV in the background. Smart Turn v3.2 explicitly focused on noisy environments because production is nothing like a clean mic. Expect your false-cutoff rate to double in noise and know how much is acceptable.

6. Code-switching and accented speech

Flux Multilingual claims end-of-turn decisions under 400ms across languages via model-based turn detection rather than silence thresholds. Smart Turn v3 covers 23 languages. The LiveKit multilingual turn detector covers 14. The way to verify is to have personas speak with the accents and language mixes your actual callers will use, and measure the cutoff rate per language separately.

7. Mid-turn tool-call latency

If the agent kicks off a tool call speculatively during the caller's turn (preemptive generation), make sure the turn detector does not get confused when the caller continues speaking after the "eager" end-of-turn fires. Flux exposes an eager-end-of-turn event specifically to support speculative response generation, and the "turn-resumed" case needs explicit coverage.

8. The DTMF / backchannel edge

Many agents handle DTMF or short acknowledgements as side signals. Make sure they do not force a false turn boundary. This is the most common "worked in dev, broke in production" case.

An illustrative pre-launch turn-detection suite with per-scenario pass/fail
An illustrative pre-launch turn-detection suite with per-scenario pass/fail

Metrics that make the tradeoff visible

Testing turn detection is not pass/fail on a single number; it is a Pareto curve between responsiveness and cutoffs. LiveKit's open benchmark explicitly sweeps policy settings and measures "the tradeoff between response latency and false cutoffs." Your internal suite should do the same.

The minimum metric set I would track per release:

  • False cutoff rate: percentage of turns where the agent spoke while the caller was still mid-utterance. Scored from the audio, not the transcript.
  • End-of-turn response latency P50 / P95: time from the caller's last content word to the first agent audio. Separate the P95 because that is where cutoffs and dead air both live.
  • False interrupt rate: percentage of agent turns truncated by a backchannel or noise rather than a real caller interruption.
  • Dead-air rate: percentage of turns with >1.5s of silence after a clear end-of-utterance.
  • Per-persona breakdown: the same five metrics bucketed by language, accent, pace, and background noise. A single-number average will hide the only regressions that matter.

The one thing worth adding over LiveKit's benchmark approach is per-scenario thresholds. "Trailing-off" turns can tolerate a slightly higher P95 latency; "short confirmation" turns cannot. Encode that in your pass criteria.

Where Roark fits

Validating any of the above requires one thing a text eval cannot provide: a lot of calls, over real telephony, from callers who sound different enough from one another to shake out the long-tail cases.

That is what Roark is built for. Simulations dial your agent over real phone calls (PSTN) and WebRTC, driven by personas that specify voice, language, accent, speech pace, emotional register, and background-noise environment. For turn detection specifically, this matters: the agent experiences the simulated caller the same way it will experience a real one, through the same audio pipeline, with the same codec and jitter, so the turn detector sees audio that resembles production rather than clean studio TTS.

Scoring is audio-native. Roark's models score the sound of the call, including pace and pauses and interruptions, so false cutoffs and dead air are measurable without hand-labeling. Every call ends up with scores against the metric suite, and failures are filed as issues automatically.

An illustrative scored call where early endpointing cuts off the caller mid-utterance
An illustrative scored call where early endpointing cuts off the caller mid-utterance

The piece that closes the loop is production replay. When a real call shows a false cutoff, Roark can capture it and replay it against updated agent logic, turning the exact audio that broke you into a repeatable regression test. For turn detection this is especially valuable because the failure is almost always something you did not think to script: a specific pause pattern, a particular accent, an unusual filler word. Catching one in production and converting it into a permanent test is the only way to stop seeing it twice.

The integrations line up with wherever your agent lives: Roark has one-click integrations for Vapi, Retell, LiveKit, Pipecat, Bland, and ElevenLabs. Pre-launch simulation suites can be gated on CI over HTTP, so a change to your endpointing delays, VAD threshold, or STT provider gets validated against the full turn-detection suite before it reaches a caller.

The short version

Model latency has fallen fast enough that endpointing is now the thing people notice. The current generation of turn detectors (LiveKit v1, Smart Turn v3, Flux, Realtime semantic_vad) are genuinely better than silence-timer baselines, but none of them removes the need to test behavior against audio that resembles your actual callers. The practical test plan is maybe eight scenarios, a handful of audio-native metrics, and a mechanism to turn production failures into regressions. Build that, and the "sometimes it cuts people off" bug stops recurring.

Daniel Gauci Mizzi

Written by

Daniel Gauci Mizzi · Co-founder & CTO @ Roark

Building Roark — the quality platform that simulates, monitors, and auto-improves voice and chat agents.

Bring a recording.
We’ll score it live.

See your own agent measured on the audio it actually produced, in the demo, in real time. Stop guessing whether your voice AI works.

Or start free with $50 in credit · read the docs · support@roark.ai