A week ago Decagon shipped Voice 3 with a new in-house speech model and a <a href="https://decagon.ai/blog/voice-3">duplex architecture that lets the agent listen, speak, and act at the same time</a>. The pitch is the same one Deepgram's Flux, LiveKit's preemptive generation, and a growing stack of arxiv papers have been building toward for a year: stop waiting until the caller is clearly finished before you start thinking. Draft a response while they might still be talking, cancel it if they keep going, commit it the moment they stop.
The latency win is real. The problem is that "draft and maybe commit" is a brand new class of state in your voice pipeline, and most teams have no tests for it. The agent that sounds snappy on your demo call is the same agent that will, in production, speak half a sentence from a draft the caller invalidated two hundred milliseconds ago. If you turn on eager_eot_threshold and don't change how you test, you ship a race condition with better p50.
What preemptive generation actually does
Deepgram's Flux API exposes the pattern directly. You set an eager_eot_threshold somewhere in the <a href="https://developers.deepgram.com/docs/flux/configuration">0.3 to 0.9 range</a>, and when the model crosses that confidence it fires an EagerEndOfTurn event with a candidate transcript. Your agent uses that transcript to start generating a response. If the caller resumes, Flux fires TurnResumed and <a href="https://developers.deepgram.com/docs/flux/configuration">you cancel the draft</a>. If they really did finish, EndOfTurn arrives and the transcript matches exactly, so you commit.
LiveKit's agents stack wires the same idea end to end. In the Deepgram plugin's <a href="https://docs.livekit.io/agents/models/stt/plugins/deepgram/">STTv2 configuration</a> you pass eager_eot_threshold=0.4 and set turn_detection="stt":
session = AgentSession(
turn_detection="stt",
stt=deepgram.STTv2(
model="flux-general-en",
eager_eot_threshold=0.4,
),
vad=silero.VAD.load(),
# ... llm, tts, etc.
)LiveKit also <a href="https://livekit.com/blog/turn-detection-and-interruption-handling">documents preemptive LLM generation and optional preemptive speech synthesis</a>, which pushes the draft further down the pipeline into TTS. Decagon's duplex design generalizes this: a <a href="https://www.chatpicture.com/en/blog/your-ai-agent-is-now-calling-customer-support-and-businesses-have-to-answer">low-latency model handles listening and speaking, while a heavier model manages reasoning, tool calling, and guardrails behind the conversation</a>. The academic version is Voice-Light and the growing speculative-interaction-agent literature, where an <a href="https://arxiv.org/pdf/2509.01920">approximation agent generates tentative planning steps while a stronger target agent verifies them asynchronously</a>.
All of these share the same shape: a fast path commits optimistically, a slow path is the source of truth, and something has to reconcile them when they disagree.

The failure modes you actually need to test for
If you only test "does the agent answer the question", preemptive generation looks like free latency. The failures live in the gaps between draft, cancel, and commit.
Stale drafts resurfacing. This is the signature bug. The eager event fires, the LLM starts generating, the caller adds a word, you fire TurnResumed, but the cancellation doesn't reach every component in time. A pending LLM generation, a queued synthesis job, or an unexecuted function call can <a href="https://theagenticstack.substack.com/p/when-a-voice-agent-answers-too-soon">resurface as a response to a question the user already abandoned</a>. The symptom is a clean-sounding agent answering the wrong, half-formed version of a question.
Draft and target divergence. Decagon's duplex and the speculative-planning literature both assume the fast model's answer is close enough to the slow model's that commits are safe. When they diverge, the fast layer has already said something the slow layer wouldn't have. The <a href="https://arxiv.org/pdf/2509.01920">latency of speculative pipelines depends critically on the alignment between draft and target components</a>; when alignment slips, you get a voice agent that announces "I'll transfer you to billing" before the authoritative reasoner decides the right action is to collect a card number.
Incorrect eager commits. The eager threshold is a probability, not a promise. Deepgram's docs are explicit that <a href="https://developers.deepgram.com/docs/flux/configuration">lower values give earlier triggers but more false starts</a>. In test, you need coverage of the specific speech patterns that fool the model: trailing conjunctions ("and..."), ambient breath, a sharp inhale that reads as end-of-utterance, a short "yeah" interjection in the middle of a long thought.
Interruption races with speculative TTS. When the draft advances all the way to synthesized audio before being cancelled, the race moves to whether the audio buffer is drained before any frames are rendered to the caller. If a single hundred-millisecond frame escapes, you ship a tiny, decontextualized phoneme into the call, which the caller will remember as "it talks over itself."
Rising p99 under adversarial callers. Preemptive generation improves p50 by making the common case fast. It can make the tail worse. A caller who hesitates mid-sentence pays the cost of a cancelled draft plus a fresh generation plus whatever queue the TTS now has to drain, and your p99 end-of-turn latency moves the wrong direction even as p50 looks great.
Discard rate as a leading indicator. As the pattern note on preemptive generation puts it, <a href="https://theagenticstack.substack.com/p/when-a-voice-agent-answers-too-soon">measure how often prepared responses survive, along with the extra model and speech usage</a>. A high discard rate is a cost signal and a UX signal at the same time: you are paying for LLM and TTS tokens that never reach the caller, and your threshold is set too aggressively.
A test plan that covers the draft/commit boundary
A reasonable testing strategy for preemptive generation has three layers.
1. Deterministic unit tests on the state machine
Stand up a fake STT that emits a scripted sequence of EagerEndOfTurn, TurnResumed, and EndOfTurn events, then assert that your agent code:
- Cancels the pending generation on
TurnResumedand the cancellation is observable in metrics within one event loop tick. - Tags every speculative generation with an attempt id, and discards results arriving on an older id.
- Never writes to the TTS output buffer from a cancelled attempt, even if the LLM completes after cancellation.
- Correctly treats the
EndOfTurntranscript as authoritative when it differs from the earlier eager transcript, including the case where the eager transcript was committed to any downstream tool call.
These are the invariants Flux's own documentation hints at when it notes the eager transcript "will exactly match the transcript in the subsequent EndOfTurn event (if no TurnResumed occurs)." Your job is to prove the "if no TurnResumed occurs" case is handled.
2. Audio-level scenarios over real calls
Unit tests catch logic bugs. They do not catch the physics of a real codec, jitter buffer, and VAD under load. For that you need to dial the agent over a real phone call with scenario audio engineered to stress the eager path. The scenarios that find bugs:
- Mid-sentence breath. A caller says "my account number is one two three" (breath) "four five six seven eight". The breath often reads above threshold.
- Trailing conjunction. "I'd like to cancel my policy and..." then two seconds of silence, then "and also move the renewal date."
- Backchannel from the caller. The agent is partway through a sentence; the caller says "mhm" without intending to take the turn.
- Short confirmation inside a long turn. "Yes, that one, the twenty-four month plan from last March, I think."
- Code-switching. A sentence that switches language mid-utterance, which semantic turn detectors handle unevenly.
- Fast repeated corrections. "Tuesday, no Wednesday, no actually Thursday."
Each of these should be run as a persona with distinct voices, pacing, and background noise, because eager thresholds behave very differently against clean studio audio than against a kitchen with a dishwasher. This is the kind of coverage simulation testing is built for: Roark's simulations <a href="https://roark.ai">dial your agent over real telephony with personas defined by voice, language, pace, emotional register, and background noise</a>, and record the full audio for every run.

3. Production replay of eager-boundary failures
The scenarios you design in advance are a floor. The ones you catch in production are the ceiling. Every call where the agent spoke, got cut off within 400 ms, and then switched topic is a candidate for a preemptive-generation failure. The fix is a loop: capture that call, replay it against the next build with the same audio, and keep it as a regression test forever. Roark's production replay is designed exactly for this, taking real captured calls and <a href="https://roark.ai">replaying them against updated agent logic so real failures become repeatable regression tests</a>.
Metrics worth tracking
Scoring these calls with transcript-only evaluators misses the point. A stale draft and a correct answer look nearly identical in text. The signal is in the audio and the timing.
The metric set that actually catches preemptive-generation regressions:
| Metric | What it tells you |
|---|---|
| Final-VAD-to-first-audio p50 and p99 | Whether speculation is paying off, and whether the tail is quietly getting worse |
| Draft discard rate | Share of eager generations never committed; too high and you are wasting spend, too low and you are not actually speculating |
| Stale-utterance rate | Share of calls where any audio from a cancelled draft reached the caller |
| Eager/commit transcript divergence | Share of eager transcripts that did not match the eventual EndOfTurn transcript |
| Interruption-after-eager rate | Share of eager-committed turns where the caller interrupted within 500 ms of first audio, which usually means a bad commit |
| Turn-boundary dead air | Pauses longer than a threshold immediately after a committed eager turn, which can mean the slow model rejected the draft |
The last four of these are only visible if your scoring layer is audio-native. Roark's models <a href="https://roark.ai">score the sound of the call, not just the transcript: pronunciation, emotion, vocal stress, pace and pauses, and interruptions</a>, which is what you need to distinguish "the agent answered correctly" from "the agent answered the right question with a hundred-millisecond ghost of the wrong one."

A note on thresholds
The right eager_eot_threshold is empirical and workload-dependent. Deepgram's own guidance is that lower values (0.3 to 0.5) give earlier triggers and more false starts, higher values (0.6 to 0.8) are more conservative with fewer cancellations. The practical approach is to pick a starting value, run a representative scenario suite plus a replay set of captured production calls, and compare three numbers across threshold sweeps: p50 latency, discard rate, stale-utterance rate. The curve is almost never monotonic in a way the docs can predict for your traffic.
Treat the threshold as a config you own, not a vendor default you accept. The same discipline applies to LiveKit's preemptive speech synthesis toggle, Decagon's duplex layer whenever you tune its behavior through the integration, and any custom speculative tool-calling layer you build on top. Every one of them is a knob that improves the median by trading tail risk you can only see with audio-native scoring and replay.
What to do this week
If you have shipped a voice agent and have not touched the eager-generation knobs yet, you are leaving the latency on the table. If you have turned them on and your tests still assert only on final transcripts, you are shipping ghosts.
Three concrete steps:
- Add a scenario suite that specifically targets eager-boundary failures: breaths, trailing conjunctions, backchannels, code-switching, rapid corrections. Run it as audio over real telephony, not transcript playback.
- Measure discard rate, stale-utterance rate, and eager-commit transcript divergence as first-class metrics alongside your latency numbers.
- Capture every production call with a sub-second interruption after the agent spoke, and replay the audio against each new build. These are the calls where preemptive generation most often misbehaves.
The underlying lesson is older than voice AI: optimistic concurrency is cheap on the happy path and expensive at the edges, and the only way to know your edges are safe is to point tests at them. The duplex and eager-EOT architectures that make 2026 voice agents feel human are worth the complexity. They are also worth testing like the concurrency systems they are.

