All field notes

Voice AI Testing

·

How to test voice agent barge-in and interruption handling

Barge-in is where voice agents feel human or robotic. Here is how to test interruption handling systematically instead of hoping it works in production.

Daniel Gauci Mizzi

Daniel Gauci Mizzi

Co-founder & CTO @ Roark

7 min read
How to test voice agent barge-in and interruption handling

Barge-in, a caller interrupting the agent mid-sentence, is the single feature that decides whether your voice agent feels human or infuriating. Get it right and conversations flow. Get it wrong and the agent talks over people, ignores "stop", or cuts itself off at every "um". Yet almost nobody tests it, because it is invisible in a transcript and only shows up in the audio.

This is a walkthrough of how to test interruption handling deliberately, and how Roark turns those tests into a suite you run before every deploy.

Why barge-in breaks in production

Interruption handling sits at the seam of three subsystems that each have their own timing, and the bugs live in how they disagree:

  • Endpointing, deciding the caller has stopped speaking. Tune it too eager and the agent barges in on a mid-sentence pause; too patient and it feels laggy. Providers like Deepgram expose this directly.
  • VAD and turn detection, whether inbound speech during agent playback counts as an interruption or as background noise. LiveKit's turn detection is a good reference model.
  • TTS cancellation, how fast the agent actually stops talking once it decides it was interrupted. Buffered audio that keeps playing for 400ms after "stop" is what makes an agent feel like it is not listening.
How an interruption travels through the stack
How an interruption travels through the stack

None of this is visible in a text transcript. Two agents with identical transcripts can feel completely different depending on when each word was spoken and how fast the agent yielded the floor.

What a text log will never show you

A response that starts 3.8 seconds late reads as a perfectly good transcript. In the audio it is a caller sitting in silence, wondering if the line dropped.

A scored production call in Roark
A scored production call in Roark

Roark scores what happened in the audio, so dead air, talk-over, and slow yields become numbers instead of vibes.

Testing it by hand does not scale

You can catch the worst cases by dialing the agent and rudely interrupting it a few times. That finds the obvious "it never stops" bug. It does not find the regression three weeks later when someone bumps the endpointing threshold by 100ms to fix a different complaint and quietly reintroduces talk-over on every third call.

Ad hoc testing vs. a repeatable suite
Ad hoc testing vs. a repeatable suite

Manual testing also cannot answer the question that actually matters: did this change make interruption handling better or worse, on average, across a representative set of callers? That is a measurement problem, and measurement needs to be automated.

Turning barge-in into a repeatable suite with Roark

Roark simulates real callers against your agent, including the interruption patterns above, and scores what happens in the audio. A barge-in suite runs the hard-interrupt, backchannel, mid-pause and noise cases on every candidate build:

A pre-launch barge-in suite
A pre-launch barge-in suite

Because it is a suite, you run it in CI on every prompt change, model swap or TTS update. A regression in interruption latency fails the build the same way a broken unit test would, before it reaches a caller, not after a week of bad calls.

Where to start

Pick your worst offender first. Most teams find their false barge-in rate is the embarrassing one: the agent stops talking every time the caller breathes. Write that case, get a baseline score, tune endpointing, and watch the number move. Then add the hard-interrupt case as your guardrail so you never trade one failure for the other.

The point is not to hit a perfect score. It is to make interruption handling a number you can see and defend, instead of a vibe you hope survived the last deploy. That is the whole idea behind Roark: if you can measure it, you can improve it, and you can prove you did.

Daniel Gauci Mizzi

Written by

Daniel Gauci Mizzi · Co-founder & CTO @ Roark

Building Roark — the quality platform that simulates, monitors, and auto-improves voice and chat agents.

Bring a recording.
We’ll score it live.

See your own agent measured on the audio it actually produced — in the demo, in real time. Stop guessing whether your voice AI works.

support@roark.ai · we reply fast