On October 1, Inworld acquired Ultravox to accelerate its speech-to-speech work. The same week, Decagon shipped Voice 3 built around Chord, its own in-house speech model. OpenAI's gpt-realtime line keeps folding reasoning into the audio model rather than bolting it on. Microsoft, meanwhile, has introduced MAI-Transcribe-2-Streaming aimed at ultra-realistic voice agents. The direction is unambiguous: the STT then LLM then TTS pipeline is being replaced, call by call, with single models that take audio in and emit audio out.
That shift breaks most voice-AI test suites. If your pre-launch QA depends on reading transcripts and asserting against text, you are grading an artifact the production model does not actually produce. The sound of the call is the output now, not a byproduct of it. If you cannot score sound, you cannot sign off on the agent.
What speech-to-speech actually changes
Three steps collapse into one. In a classic stack a transcriber turns audio into text, a language model reasons over that text, and a TTS engine (ElevenLabs, Cartesia, Inworld's Realtime TTS-2) speaks the reply. Every stage emits an artifact you can log, diff and assert on.
Speech-to-speech removes the middle. The Vapi docs on the OpenAI Realtime API put it plainly: the Realtime API natively processes audio in and audio out, in contrast to configurations that orchestrate a transcriber, a model and a voice. Any transcript you look at after the fact is reconstructed, often by a different model, after the agent has already decided what to say.
Non-verbal information becomes first-class. OpenAI's own writeup of gpt-realtime highlights that the model can capture non-verbal cues like laughs, switch languages mid-sentence, and shift between "snappy and professional" and "kind and empathetic" on instruction. A transcript smooths all of that out. "I'm so sorry about that" written down looks identical whether the agent delivered it warmly or sprinted through it like a legal disclosure.
Turn-taking is now a modeling decision, not a wrapper. Azure's Realtime API guide describes a semantic VAD mode where the server estimates the probability that the caller is done speaking based on the words themselves, which makes the agent less likely to interrupt. Different models tune that threshold differently. Swap one in and your interruption profile moves, silently, in production.

Why transcript-only tests miss the regressions you care about
The gap between "the transcript reads fine" and "the call sounded fine" is wider than most teams realize. Here is what text-based evaluation routinely misses on a speech-to-speech agent:
- Tone mismatches. A refund denial delivered in the same cheerful cadence as an upsell. The words are correct. The experience is not.
- Pacing drift. A model that used to pause for a beat after a long number now barrels through, and nobody notices until a customer calls to complain their account number was read "too fast to catch."
- Interruption behavior. Semantic VAD tuning changes between model versions, so the agent either cuts callers off or waits an awkward extra half-second. Braintrust's audio-eval cookbook is upfront that evaluating these outputs in practice is still an unsolved problem for most teams.
- Non-verbal responses. A caller laughs; the agent either laughs back, acknowledges it, or ignores it. On the transcript, all three paths look like one blank turn.
- Code-switching. The caller flips from English to Spanish mid-sentence. gpt-realtime is designed to handle this cleanly; many transcription-based evals will mis-segment or mis-transcribe the switch and score the wrong failure mode.
- Pronunciation of names, SKUs and medical terms. The transcript shows the right spelling because the ASR you used to generate it knows the word. The caller heard something else.
The classic post-mortem on a text-only suite is also uncomfortable. You read a passing transcript, listen to the recording, and realize the agent sounded nothing like what the words suggested. By then it has already shipped.
What to actually measure on a speech-to-speech agent
OpenAI's own voice agents evaluation guide tells teams to measure task completion, audible response latency, interruptions, and unwanted silence as independent dimensions, and to keep caller, model, tools and transport consistent across runs. That is the right framing. Translated into a PM-level test checklist for a speech-to-speech agent, you want scores for:
- Task outcome. Did the agent do the thing? Did the spoken confirmation match the actual tool call and the actual back-end state?
- Pronunciation accuracy. On the specific names, numbers, drug names, SKUs and addresses your callers use. Not generic WER.
- Emotional appropriateness. Did the tone match the situation (refund, bereavement, urgent outage) as a human would judge it?
- Pace and pauses. Dead air and rushed delivery, measured in seconds, not inferred from the transcript.
- Interruption dynamics. How often does the agent talk over the caller, and how long does it wait after a true end of speech?
- Vocal stress and prosody. Does the voice carry the right emphasis on the right words, especially when confirming money, dates and identifiers?
- Language and accent handling. For agents with mid-call language switching or accented callers, scored per language, not averaged into one number.
Those are measurable, but only against audio. Running them on a transcript gives you a number that does not survive first contact with a real caller.

A pre-launch test suite for a speech-to-speech agent
The suite you ship with should be made of scenarios that a transcript would not catch. Here is a working starting set, assuming an inbound support agent on OpenAI's Realtime API or similar:
| Scenario | What to assert | Why it matters |
|---|---|---|
| Caller interrupts the agent mid-sentence | Agent yields within ~400ms; resumes with the right new turn | Semantic VAD behavior shifts with model swaps |
| Caller switches English to Spanish on the second turn | Agent follows the switch; pronunciation of names stays correct | gpt-realtime claims mid-sentence switch support; verify on your prompts |
| Caller laughs after a joke or a friendly line | Agent acknowledges rather than ignoring the beat | Non-verbal cues are invisible on text |
| Caller reads a 16-digit account number quickly | Agent captures it correctly and reads back at a slower, intelligible pace | Pace is a vocal choice the model makes |
| Caller is upset about a billing error | Agent tone registers as empathetic, not transactional | Tone mismatches are the most common silent regression |
| Background noise: cafe, car, baby | Task completion stays stable; interruption rate does not spike | Audio conditions change VAD behavior |
| Caller asks for a human | Agent escalates cleanly without extra sales turns | Handoff path tends to drift as prompts are tuned |
Every row should be a persona plus a scenario plus a scoring rubric that includes at least one audio-native metric. "The agent sounded appropriately empathetic" is a valid pre-launch gate when it is scored consistently against the same audio by the same model across runs.
Make it a regression test, not a demo
The hardest part of this is not building the first run. It is keeping it alive. Speech-to-speech agents will absorb at least as many model swaps as pipeline agents did, probably more, because each major Realtime release moves the audio behavior. GPT-Realtime-2 is already reported to score 15.2% higher on Big Bench Audio than its predecessor, and anything that moves intelligence that much also moves the voice.
Three habits make the suite survive:
- Capture real production failures as reusable scenarios. The call where a customer got frustrated by the agent cutting them off is also the best possible regression test for interruption behavior. Replay it on every candidate model.
- Run the suite on every change. Prompt tweaks, model version bumps, tool additions, voice swaps. Each of these can shift audio behavior independently.
- Gate deploys automatically. The team that treats the voice suite as a nightly report will drift. The team that blocks production on it will not.
ElevenLabs' agent business is now handling more than 15 million conversations a week, roughly triple what it was in February. Decagon reports that about 90% of listeners could not tell Chord from a human. The ceiling on voice-agent quality is rising fast; the floor underneath you is moving at the same time. Teams that do not test audio directly will keep shipping regressions they only notice on review calls, days later.

How Roark handles this
Roark was built for the audio case. Its simulation testing dials your voice agent over real phone calls, PSTN and WebRTC, with personas that control voice, accent, speech pace, emotional register and background-noise environment. It supports 45 languages and accents, which is where the mid-call language-switch scenarios live.
Scoring is audio-native: pronunciation, emotion, vocal stress, pace and pauses, and interruptions, plus 64+ built-in metrics and unlimited custom ones. Those are the signals that disappear from a transcript the moment you move to a speech-to-speech model. Every production call is scored against the same suite, and whatever breaks is filed as an issue automatically.
For regression work, Roark replays captured production calls against your updated agent logic, so a failure you saw once becomes a test you cannot ship over. One-click integrations exist for Vapi, Retell, LiveKit, Pipecat, Bland and ElevenLabs, and simulations can be triggered over HTTP to gate CI. The point is not to add another dashboard. It is to make sure the thing your callers actually hear is the thing you tested.
Speech-to-speech is the direction, and the ecosystem has made that clear in the last two weeks. The teams that treat the audio as the output, not the transcript, are the ones who will ship these agents without flinching.

