All field notes

Voice AI Testing

·

Testing speech-to-speech voice agents before launch

Speech-to-speech voice agents don't produce a transcript to grade. Here's what to measure before launch, and the test suite PMs should ship with.

James Zammit

James Zammit

Co-founder & CEO @ Roark

8 min read
Testing speech-to-speech voice agents before launch

On October 1, Inworld acquired Ultravox to accelerate its speech-to-speech work. The same week, Decagon shipped Voice 3 built around Chord, its own in-house speech model. OpenAI's gpt-realtime line keeps folding reasoning into the audio model rather than bolting it on. Microsoft, meanwhile, has introduced MAI-Transcribe-2-Streaming aimed at ultra-realistic voice agents. The direction is unambiguous: the STT then LLM then TTS pipeline is being replaced, call by call, with single models that take audio in and emit audio out.

That shift breaks most voice-AI test suites. If your pre-launch QA depends on reading transcripts and asserting against text, you are grading an artifact the production model does not actually produce. The sound of the call is the output now, not a byproduct of it. If you cannot score sound, you cannot sign off on the agent.

What speech-to-speech actually changes

Three steps collapse into one. In a classic stack a transcriber turns audio into text, a language model reasons over that text, and a TTS engine (ElevenLabs, Cartesia, Inworld's Realtime TTS-2) speaks the reply. Every stage emits an artifact you can log, diff and assert on.

Speech-to-speech removes the middle. The Vapi docs on the OpenAI Realtime API put it plainly: the Realtime API natively processes audio in and audio out, in contrast to configurations that orchestrate a transcriber, a model and a voice. Any transcript you look at after the fact is reconstructed, often by a different model, after the agent has already decided what to say.

Non-verbal information becomes first-class. OpenAI's own writeup of gpt-realtime highlights that the model can capture non-verbal cues like laughs, switch languages mid-sentence, and shift between "snappy and professional" and "kind and empathetic" on instruction. A transcript smooths all of that out. "I'm so sorry about that" written down looks identical whether the agent delivered it warmly or sprinted through it like a legal disclosure.

Turn-taking is now a modeling decision, not a wrapper. Azure's Realtime API guide describes a semantic VAD mode where the server estimates the probability that the caller is done speaking based on the words themselves, which makes the agent less likely to interrupt. Different models tune that threshold differently. Swap one in and your interruption profile moves, silently, in production.

Pipeline agents give you three artifacts to grade. Speech-to-speech gives you one, and it is the audio.
Pipeline agents give you three artifacts to grade. Speech-to-speech gives you one, and it is the audio.

Why transcript-only tests miss the regressions you care about

The gap between "the transcript reads fine" and "the call sounded fine" is wider than most teams realize. Here is what text-based evaluation routinely misses on a speech-to-speech agent:

  • Tone mismatches. A refund denial delivered in the same cheerful cadence as an upsell. The words are correct. The experience is not.
  • Pacing drift. A model that used to pause for a beat after a long number now barrels through, and nobody notices until a customer calls to complain their account number was read "too fast to catch."
  • Interruption behavior. Semantic VAD tuning changes between model versions, so the agent either cuts callers off or waits an awkward extra half-second. Braintrust's audio-eval cookbook is upfront that evaluating these outputs in practice is still an unsolved problem for most teams.
  • Non-verbal responses. A caller laughs; the agent either laughs back, acknowledges it, or ignores it. On the transcript, all three paths look like one blank turn.
  • Code-switching. The caller flips from English to Spanish mid-sentence. gpt-realtime is designed to handle this cleanly; many transcription-based evals will mis-segment or mis-transcribe the switch and score the wrong failure mode.
  • Pronunciation of names, SKUs and medical terms. The transcript shows the right spelling because the ASR you used to generate it knows the word. The caller heard something else.

The classic post-mortem on a text-only suite is also uncomfortable. You read a passing transcript, listen to the recording, and realize the agent sounded nothing like what the words suggested. By then it has already shipped.

What to actually measure on a speech-to-speech agent

OpenAI's own voice agents evaluation guide tells teams to measure task completion, audible response latency, interruptions, and unwanted silence as independent dimensions, and to keep caller, model, tools and transport consistent across runs. That is the right framing. Translated into a PM-level test checklist for a speech-to-speech agent, you want scores for:

  1. Task outcome. Did the agent do the thing? Did the spoken confirmation match the actual tool call and the actual back-end state?
  2. Pronunciation accuracy. On the specific names, numbers, drug names, SKUs and addresses your callers use. Not generic WER.
  3. Emotional appropriateness. Did the tone match the situation (refund, bereavement, urgent outage) as a human would judge it?
  4. Pace and pauses. Dead air and rushed delivery, measured in seconds, not inferred from the transcript.
  5. Interruption dynamics. How often does the agent talk over the caller, and how long does it wait after a true end of speech?
  6. Vocal stress and prosody. Does the voice carry the right emphasis on the right words, especially when confirming money, dates and identifiers?
  7. Language and accent handling. For agents with mid-call language switching or accented callers, scored per language, not averaged into one number.

Those are measurable, but only against audio. Running them on a transcript gives you a number that does not survive first contact with a real caller.

An illustrative pre-launch suite for a speech-to-speech agent, scored on audio-native metrics.
An illustrative pre-launch suite for a speech-to-speech agent, scored on audio-native metrics.

A pre-launch test suite for a speech-to-speech agent

The suite you ship with should be made of scenarios that a transcript would not catch. Here is a working starting set, assuming an inbound support agent on OpenAI's Realtime API or similar:

ScenarioWhat to assertWhy it matters
Caller interrupts the agent mid-sentenceAgent yields within ~400ms; resumes with the right new turnSemantic VAD behavior shifts with model swaps
Caller switches English to Spanish on the second turnAgent follows the switch; pronunciation of names stays correctgpt-realtime claims mid-sentence switch support; verify on your prompts
Caller laughs after a joke or a friendly lineAgent acknowledges rather than ignoring the beatNon-verbal cues are invisible on text
Caller reads a 16-digit account number quicklyAgent captures it correctly and reads back at a slower, intelligible pacePace is a vocal choice the model makes
Caller is upset about a billing errorAgent tone registers as empathetic, not transactionalTone mismatches are the most common silent regression
Background noise: cafe, car, babyTask completion stays stable; interruption rate does not spikeAudio conditions change VAD behavior
Caller asks for a humanAgent escalates cleanly without extra sales turnsHandoff path tends to drift as prompts are tuned

Every row should be a persona plus a scenario plus a scoring rubric that includes at least one audio-native metric. "The agent sounded appropriately empathetic" is a valid pre-launch gate when it is scored consistently against the same audio by the same model across runs.

Make it a regression test, not a demo

The hardest part of this is not building the first run. It is keeping it alive. Speech-to-speech agents will absorb at least as many model swaps as pipeline agents did, probably more, because each major Realtime release moves the audio behavior. GPT-Realtime-2 is already reported to score 15.2% higher on Big Bench Audio than its predecessor, and anything that moves intelligence that much also moves the voice.

Three habits make the suite survive:

  • Capture real production failures as reusable scenarios. The call where a customer got frustrated by the agent cutting them off is also the best possible regression test for interruption behavior. Replay it on every candidate model.
  • Run the suite on every change. Prompt tweaks, model version bumps, tool additions, voice swaps. Each of these can shift audio behavior independently.
  • Gate deploys automatically. The team that treats the voice suite as a nightly report will drift. The team that blocks production on it will not.

ElevenLabs' agent business is now handling more than 15 million conversations a week, roughly triple what it was in February. Decagon reports that about 90% of listeners could not tell Chord from a human. The ceiling on voice-agent quality is rising fast; the floor underneath you is moving at the same time. Teams that do not test audio directly will keep shipping regressions they only notice on review calls, days later.

Decagon's own number on how convincing a current speech-to-speech agent sounds. The thing most teams are grading on a transcript.
Decagon's own number on how convincing a current speech-to-speech agent sounds. The thing most teams are grading on a transcript.

How Roark handles this

Roark was built for the audio case. Its simulation testing dials your voice agent over real phone calls, PSTN and WebRTC, with personas that control voice, accent, speech pace, emotional register and background-noise environment. It supports 45 languages and accents, which is where the mid-call language-switch scenarios live.

Scoring is audio-native: pronunciation, emotion, vocal stress, pace and pauses, and interruptions, plus 64+ built-in metrics and unlimited custom ones. Those are the signals that disappear from a transcript the moment you move to a speech-to-speech model. Every production call is scored against the same suite, and whatever breaks is filed as an issue automatically.

For regression work, Roark replays captured production calls against your updated agent logic, so a failure you saw once becomes a test you cannot ship over. One-click integrations exist for Vapi, Retell, LiveKit, Pipecat, Bland and ElevenLabs, and simulations can be triggered over HTTP to gate CI. The point is not to add another dashboard. It is to make sure the thing your callers actually hear is the thing you tested.

Speech-to-speech is the direction, and the ecosystem has made that clear in the last two weeks. The teams that treat the audio as the output, not the transcript, are the ones who will ship these agents without flinching.

James Zammit

Written by

James Zammit · Co-founder & CEO @ Roark

Building Roark — the quality platform that simulates, monitors, and auto-improves voice and chat agents.

Bring a recording.
We’ll score it live.

See your own agent measured on the audio it actually produced, in the demo, in real time. Stop guessing whether your voice AI works.

Or start free with $50 in credit · read the docs · support@roark.ai