Full-duplex voice models are showing up in production. On August 3, NVIDIA released NemotronLabs VoiceChat 11B, an open-weight speech-to-speech model that listens and speaks at the same time, and it is the first open full-duplex model to support tool calling without pausing the conversation. Independent measurements put its smooth turn-taking latency at 448 ms on Full-Duplex-Bench 1.0 with a 1.00 takeover rate at 480 ms. Proprietary duplex APIs from OpenAI and Google are already in general availability, and Tencent, Alibaba, and NVIDIA have released research variants over the last quarter.
If you are still testing these agents the way you tested your STT → LLM → TTS stack, you will miss most of what breaks. The failure surface is different. The metrics your platform reports are different. And a paper posted to arXiv on September 17 just gave the field a fresh reason to worry about content-driven behavior in duplex models. This post is a practical guide to what breaks in full-duplex agents, which behaviors you need to test, and how to actually simulate them over the phone.
The pipeline collapsed. So did your test harness.
A cascaded voice agent is easy to reason about because each stage has an observable interface. Endpointing produces a final transcript. The LLM produces tokens. TTS produces audio. You can log timestamps, replay each stage in isolation, and inject text into the LLM to test policy without ever speaking. This is how most testing tooling built up over the last two years works.
A full-duplex model does not expose those seams. NVIDIA describes it plainly: one unified architecture jointly performs streaming speech understanding and speech generation, eliminating the multi-model orchestration and API handoffs a cascaded stack requires. There is no "final transcript" moment. There is no clean LLM input to fuzz. The only interface that reliably exists is audio in, audio out, over real time.
That has three practical consequences for testing:
- Text-only regression tests do not exercise the model. The model's behavior is a function of the audio timeline (overlap, silence, pitch, backchannels), not a token sequence. A prompt-level test suite tells you almost nothing.
- Stage-level latency budgets are meaningless. There is no time-to-first-token to measure. You have to measure floor-management behavior end to end, at the audio timeline.
- Turn-taking is now a model property, not an orchestration knob. As LiveKit put it in their turn-detection guide, some realtime multimodal models handle turn-taking natively, which reduces moving parts but also reduces your direct control over the behavior. If the model gets it wrong, you cannot fix it by tuning a VAD threshold. You can only fix it upstream (system prompt, fine-tune) or downstream (guardrail model listening on top).

Four failure modes you need to test for
Here are the ones I have seen bite teams the moment they moved a workflow onto a duplex model.
1. Content-driven intervention (the "why won't it push back?" bug)
The Sept 17 paper, Full-Duplex Speech Models Take the Floor When Asked, Not When Needed, is worth reading in full. The authors constructed context-matched monologues that varied only in whether the trigger utterance was neutral, contained a false fact, or contained a hazard. Across five model families they found that being directly addressed and pauses in the user's speech are far more reliable triggers for the model to speak than false facts or hazards. Even given the floor, the proportion of non-empty replies that challenge a false claim is only 14 to 15 percent, and the proportion of hazard replies that warn of danger is 4 to 7 percent.
Read that again with a production hat on. If your agent is a health triage bot, a claims intake agent, or a payment collector, "the model rarely challenges a false claim it heard" is a P0 failure mode. And it does not show up in any cascaded-stack test suite you already have, because in a cascaded stack the LLM sees the final transcript and can be prompted to correct. In a full-duplex model the decision to speak is not text-conditioned in the same way.
You need explicit tests where the simulated caller states something false or dangerous and the agent has clear grounds to intervene, then you score whether the agent takes the floor and whether the content of its response is a correction or a warning. Not one of these tests. Dozens, across your specific domain vocabulary.
2. Barge-in accuracy and false barge-in rate
Full-duplex marketing headlines almost always emphasize "the user can interrupt." What they rarely emphasize is that the model can also mis-interrupt itself when a caller says "uh-huh" or "right" as a backchannel, or when there is a bark or a horn in the background.
The FireRedChat authors define two metrics that are the right shape here: barge-in success rate at fixed offsets from user-speech onset (0 ms, 50 ms, 100 ms, and so on, with a summary metric of the minimum latency to reach 90 percent, called T90), and false barge-in rate, which measures erroneous interruptions triggered when the primary speaker is silent. A 2022 Alibaba study of rule-based barge-in detection found that under a simple ASR-confidence rule, 89 percent of interruption events were false interruptions. Neural full-duplex models do better, but the failure modes have the same shape: user greetings, backchannels, ambient noise, and non-primary speakers.
You need to simulate:
- A caller barging in mid-agent-utterance with a real interruption ("wait, no, actually...").
- A caller producing backchannels ("mm-hmm," "right," "yeah") without wanting the floor. Full-Duplex-Bench v1.5 treats these as separate capabilities for a reason.
- A caller talking to someone else in the room, so the agent should not respond at all.
- Background noise events (a door, a dog, a car horn).
For each, you need to score both whether the agent yielded when it should have and whether it kept talking when it should have.
3. Post-interruption recovery
Yielding the floor is the easy half. What the agent says next is the harder half. The IHBench paper formalizes this as post-interruption recovery within a structured workflow: after the interruption is handled, does the agent resume the checklist, the collection form, the diagnostic tree, at the right step?
A related paper on self-listening in duplex models argues that a full-duplex spoken language model can only know where it actually got to in its own speech if it receives its own already-played audio as an input stream. The authors call this anchoring: when a user interrupts a structured spoken response, the model must answer with respect to the last item the user actually heard, not the item the text generator has already planned. If the model gets this wrong, you get a very specific bug: the agent skips items, references items the caller never heard, or restarts sections. Traditional test suites do not catch it because the interruption is what makes it happen.
Test it directly. Have your simulated caller interrupt at controlled offsets (250 ms in, 1 s in, mid-list-item) and score whether the agent's next utterance is correctly anchored to what actually played.
4. Tool calling without silence
The classic cascaded pattern for a tool call is: the LLM emits a tool_call token, the orchestrator pauses TTS or emits a filler, the tool runs, the LLM resumes with the result. Full-duplex models like NemotronLabs VoiceChat handle this with a separate output channel for tool-call scripts plus operator-defined "on-hold" messages that fill the gap while an API runs. Neat, but it introduces its own failure surface.
Independent evaluation of the same model reports 82.5 percent tool-selection F1 on Full-Duplex-Bench 3.0 but only 42.2 percent tool argument accuracy. Selection is not execution. And the on-hold pattern only helps if the on-hold message actually plays, actually times out if the tool never returns, and does not get barged in on by the same audio the model is producing.
Test it directly. Trigger tool calls under interruption, under overlapping speech, and under slow tool response, and score whether the on-hold message played, whether the returned value was integrated into the following utterance, and whether the caller ever heard silence longer than your policy allows. We wrote about the slow-tool-call case in a cascaded context already; in a full-duplex model, the failure modes multiply.
The metrics that actually mean something
If you are moving a workflow onto a full-duplex model, throw out most of your legacy latency dashboards and adopt something closer to this list:
| Metric | What it measures | Fail signal |
|---|---|---|
| Barge-in T90 | Minimum latency at which 90% of user interruptions are yielded to | Model keeps talking over the caller |
| False barge-in rate | Fraction of agent turns interrupted by non-floor-taking audio | Model yields to "mm-hmm," coughs, background |
| Turn-taking latency (P50, P95) | Delay from user speech end to agent speech start | Long tail feels unreliable, P95 over 1.5 s is the standard warning line |
| Content-intervention rate | Fraction of false-fact or hazard turns where the agent challenges/warns | Silent agreement to wrong or dangerous content |
| Post-interruption anchor accuracy | After a barge-in, does the next utterance reference what actually played | Restarts, skips, hallucinated context |
| Tool argument accuracy | Fraction of tool calls with correct arguments end-to-end | Wrong lookups, wrong actions taken |
| On-hold coverage | Fraction of tool executions with a spoken on-hold message | Dead air during tool runs |
Notice what is missing: time-to-first-token, LLM latency, TTS latency. None of those exist as separable measurements in a duplex model, and pretending they do just moves the problem sideways.
How to actually run these tests
You need three things a text-based eval harness cannot give you:
- Real audio, over real telephony. The model's behavior on a WebSocket loopback is not its behavior on a PSTN call with jitter, packet loss, and codec compression. If you plan to serve the agent over the phone, test it over the phone.
- Precise event timing. You need to inject interruptions at controlled offsets from the model's own speech onset. That means driving both sides of the conversation from a harness with a shared clock, not just replaying a WAV file.
- Audio-native scoring. Transcript scoring cannot detect a false barge-in or a T90 latency, cannot measure whether the on-hold message actually played, and cannot score pace or overlap. You need scoring that listens to the audio itself.
This is what Roark is built for. Roark's simulation testing dials your agent over real phone calls, using personas that define the caller's voice, language, accent, pace, emotional register, and background-noise environment. You script the scenarios (mid-utterance barge-in at 800 ms, backchannel-only response, false-fact injection, hazard injection, tool call under overlap), and Roark runs them, scoring each call on audio-native metrics that pick up pace, pauses, interruptions, and vocal stress rather than just the transcript.

Every production call is scored against the same metric suite, so when NemotronLabs VoiceChat ships a new checkpoint, or you swap from OpenAI's Realtime API to Google's Gemini Live, or your fine-tune lands, you can turn every real failure into a replayable regression test. The failures the Sept 17 paper surfaces at the model level, the failures the FireRedChat authors describe at the barge-in level, and the failures the IHBench authors surface at the recovery level all become concrete scored scenarios you run before each release.

The bottom line
Full-duplex is a genuine architectural shift, and it moves failure into places your current test suite cannot see. The good news is the failure modes are well characterized in the literature over the last few months. The bad news is that if you do not adopt tests for them before your next model swap, you will find them in production, one call at a time, from an angry customer.
If you are evaluating Nemotron 3 VoiceChat, OpenAI Realtime, Gemini Live, or any of the open duplex models for a workflow you already run on a cascaded stack, treat the migration as a new agent, not a drop-in upgrade. Run barge-in T90, false barge-in, content-intervention, post-interruption anchoring, and tool-argument accuracy as first-class release gates. If any of them regress, do not ship.
That is the whole job.

