All field notes

Voice AI Testing

·

Testing multilingual voice agents before launch

A PM playbook for testing multilingual voice agents: per-language SLAs, code-switching, native-speaker scenarios, and audio-native scoring over real telephony.

James Zammit

James Zammit

Co-founder & CEO @ Roark

9 min read
Testing multilingual voice agents before launch

Microsoft shipped its first real-time speech-to-text model, MAI-Transcribe-2-Streaming, two weeks ago. ElevenLabs' Scribe v2 Realtime now publishes sub-150ms latency across 90+ languages with 93.5% accuracy on the FLEURS benchmark. Japanese retailer Yamada Denki put a multilingual GPT-Realtime agent on its online store and guided around 30,000 shoppers through appliance purchases in two weeks, with 92% positive post-call survey responses.

The ambition is global. The testing is almost always English-first, and the gap is where multilingual voice agents fail silently. This post is a product-manager playbook for the second question: before you flip Spanish, Hindi, or Mandarin on in production, what has to be true, and how do you prove it without shipping on hope?

The vendor sheet is not your dial plan

Every major speech provider now ships a multilingual real-time model, and every provider measures itself on the benchmark that flatters it most. ElevenLabs reports 93.5% accuracy on the FLEURS benchmark across 30 European and Asian languages for Scribe v2 Realtime. AssemblyAI advertises flagship accuracy with native code-switching on Universal-3.5 Pro Streaming across 18 languages. Microsoft's MAI-Transcribe-2-Streaming is positioned at the premium real-time end of the market.

The problem is that none of those numbers describe the audio your agent will actually hear. Bland's own writeup is blunt about this: WER scores from vendor demos are measured on clean, read-speech audio, and live telephony with real accents and background noise can push those numbers 2 to 4 times worse than what the benchmark slide shows. A model that is "93.5% accurate on FLEURS" can be far worse on a Mumbai call center line at 7pm on a Tuesday, and your product dashboard will never show you the gap unless you built it to.

What vendor benchmarks measure vs what your production calls contain
What vendor benchmarks measure vs what your production calls contain

The PM lesson is not that vendors are lying. It is that benchmark results are not acceptance criteria. Acceptance criteria are per-language, per-accent, per-channel measurements of the exact behavior you plan to ship, taken over the exact telephony path your callers will use.

Code-switching is the default, not the edge case

The second trap is treating bilingual callers as an edge case. They are the base case in most non-US markets, and the architecture most teams reach for actively breaks on them.

The familiar approach is to set a language parameter before the call routes to the STT engine and assume callers cooperate. What actually happens, as Bland and others document, is that bilingual speakers switch languages mid-utterance. When the STT engine is locked to a declared language, mid-sentence switches produce corrupted transcripts, the LLM downstream receives garbage input and generates a mismatched response, no error surfaces in your quality dashboard, and the call simply fails silently.

Murf's engineering writeup puts the timing constraint plainly: language detection has to work faster than a sentence completes, and if your system identifies "this caller is speaking French" only at the end of a turn, it is too late to route the ASR correctly for what was said mid-sentence. This is why providers like AssemblyAI now advertise native mid-sentence code-switching with no primary language to declare on their Voice Agent API. The capability exists. Teams still do not test for it.

Recent research is catching up to this. The new CS3-Bench benchmark for Mandarin-English speech-to-speech models exists specifically because older evaluations "often involve only simplistic translations of individual words, failing to capture the natural and contextually grounded patterns of bilingual usage." If the research community is still building benchmarks for this, your internal QA almost certainly has not solved it either.

The four silent failure modes

Before writing a test plan, name the failure modes. Every multilingual regression tends to trace back to one of four patterns, and each one has a different fix.

  1. Accent drift inside a supported language. The ASR claims Spanish support but was trained predominantly on Castilian audio. Mexican or Rioplatense callers get mis-transcribed. The agent asks them to repeat.
  2. Code-switching corruption. A single sentence contains two languages. The transcript loses the switched span. The LLM responds to what it thinks it heard.
  3. TTS pronunciation errors. Numbers, dates, currencies, and proper nouns pronounce correctly in one locale and wrong in another. ElevenLabs' own community docs flag that the apply_text_normalization parameter is missing from some integration layers, causing incorrect pronunciation of numbers, dates, and currencies.
  4. Cultural register mismatch. Japanese and Korean honorifics, Spanish tú/usted, French tu/vous, Hindi aap/tum: the agent uses the wrong one and the conversation gets colder without any single turn being "wrong."

A monitoring dashboard that only shows transcript sentiment or intent-match rate catches none of these cleanly. Accent drift shows up as a lower confidence score and a higher repeat rate. Code-switching corruption shows up as a confused response. Pronunciation errors show up only if a human listens. Register mismatch shows up as a drop in CSAT weeks later.

A PM playbook for multilingual QA

Here is the structure that works. It assumes you are shipping a voice agent in more than one language, or in one language across multiple accent regions, and that you want to stop multilingual bugs from being customer-facing.

1. Set per-language acceptance criteria, not a global SLA

Pick thresholds for each language separately. ASR accuracy varies enough between languages that a single global target hides most of the real quality problems. Rasa's writeup notes that a Portuguese speaker in Brazil sounds different from one in Portugal, and both differ from a speaker in Mozambique, and that each variation requires either dedicated training data or models flexible enough to generalize. Set language-specific thresholds for ASR accuracy, latency (TTS can lag more in some languages than others), task completion, pronunciation accuracy on named-entity words, and recovery rate after a misunderstanding. Publish them per language and track them per language.

2. Use native speakers, not translations

Translated test scripts miss the way real callers actually talk. A script written in English and run through a translator produces sentences no bilingual customer would ever say out loud. For regional formality, local vocabulary ("carro" vs "coche" in Spanish), and realistic disfluencies, you need either native-speaker testers or synthetic callers whose voices and phrasing were trained on real native speech, not translated corpora.

3. Test code-switching deliberately, in both directions

Build a scenario bank where callers switch languages mid-utterance in both directions, for every language pair you serve. English into Spanish and Spanish into English are not the same test. Hinglish, Spanglish, and Franglais each break ASR differently depending on which language the switch originates in. Cover at least a dozen patterns per pair, and include both lexical switches (one word) and clausal switches (a full phrase).

4. Test on the channel you ship on

A model that scores well on your laptop microphone will not score the same way through a telephony codec. Run your test calls through PSTN or WebRTC, not through a text loopback, so you measure the audio your production agent will actually ingest. If your "multilingual QA" is a batch of transcripts fed through an eval harness, you are grading the LLM, not the voice agent.

5. Score the audio, not just the text

Transcript-only evals miss pronunciation, pace, pauses, and emotional register, which is where multilingual quality lives. Rasa's writeup makes the point: the emotional register of voice matters, and a customer calling a support line in frustration sounds very different from someone calmly asking a routine billing question, regardless of language. An audio-native scoring layer that listens for pronunciation errors, stress, and awkward pauses catches regressions that text-only evaluation cannot.

6. Replay real production failures as regressions

The first time a Punjabi-accented caller trips up your agent in production, that call becomes the most valuable test asset you have. Capture it, replay it against new agent versions, and let it fail the build if the fix regresses. This is how multilingual coverage grows past the languages you launched with, instead of starting from zero every quarter.

Illustrative pre-launch simulation covering accent and code-switching scenarios
Illustrative pre-launch simulation covering accent and code-switching scenarios

Where Roark fits

Everything above is doable without Roark, but the plumbing is painful. Roark is a simulation-testing, observability, and reporting platform for production voice agents, and multilingual QA is where its design pays off most directly.

Three capabilities matter for this use case. First, Roark simulates and scores in 45 languages and accents, with personas that define the caller's voice, language, accent, speech pace, emotional register, and background noise environment. That lets you build the native-speaker, per-dialect scenario bank the playbook calls for without hiring 20 linguists for every release. Second, simulations run over real phone calls, PSTN and WebRTC, so the audio your agent hears in the test is the same shape of audio it will hear in production. Third, every live call is scored on audio-native metrics (pronunciation, emotion, pace and pauses, interruptions), not just transcript matching, and failing calls can be replayed as regression tests against the next version of your agent. Teams on Vapi, Retell, LiveKit, Pipecat, Bland, or ElevenLabs can wire it in through a one-click integration.

Pick whichever provider's benchmark looks best on paper. Then prove, in your own languages, through your own telephony, before launch, that the number holds.

The one-page checklist

Before you ship multilingual, confirm:

  • Per-language acceptance criteria written down, with thresholds you can measure in production.
  • Native-speaker scenario bank per language, including regional dialects and formality levels.
  • Code-switching tests in both directions for every language pair you serve.
  • Tests run over PSTN or WebRTC, not text loopback.
  • Audio-native scoring on pronunciation, pace, and emotion, not just transcript match.
  • A production-replay loop that turns the first real failure in each language into a permanent regression test.
  • A dashboard that slices every metric by language and accent, not just aggregate.

If any of those is missing, the next multilingual rollout is a guess. Vendors ship new languages every quarter. Your QA has to keep up, in your own voice.

James Zammit

Written by

James Zammit · Co-founder & CEO @ Roark

Building Roark — the quality platform that simulates, monitors, and auto-improves voice and chat agents.

Bring a recording.
We’ll score it live.

See your own agent measured on the audio it actually produced, in the demo, in real time. Stop guessing whether your voice AI works.

Or start free with $50 in credit · read the docs · support@roark.ai