Outbound voice AI campaigns are having a moment. In the last two weeks, Aircall shipped Outbound Campaigns as a first-class feature that lets you upload a list, pick an agent, and let it work through the contacts end to end. Bland, Retell, Vapi, and Genesys have all leaned harder on outbound in the same window. The pitch is straightforward: your list stops rotting, your SDRs stop dialing dead numbers, your renewals actually get called.
The problem is that "launch a campaign" is a very different failure surface than "launch an agent". An inbound agent handles one call at a time and, when it botches something, it botches something for one caller. A campaign botches the same thing 50,000 times in a row, at $500 to $1,500 of TCPA exposure per unlucky call, before anyone at your company notices. Pre-launch testing has to cover the whole pipeline, not just the conversation.
This post is a checklist for what PM and ops leads should confirm works before the first dial goes out.
Outbound testing is not inbound testing with a different button
Inbound testing usually focuses on a single conversational surface: a caller hits your number, the agent picks up, the agent handles the intent. You can build a small suite of scenarios and grind them.
Outbound moves the failure surface up a layer. The agent still has a conversation, but half a dozen things around the conversation can fail on their own:
- Who is the campaign allowed to call, and does the list respect that?
- What does the agent do when a voicemail picks up? A human? An IVR? Another AI?
- What are the retry rules, and do they hold when the campaign is halfway through a state that just tightened its calling window?
- What happens when the human on the other end says "take me off your list"?
- When the agent decides to transfer to a live rep, does the rep actually have context?

Each of those is testable. Most of them are not tested until the first legal notice or the first CSAT drop.
The six checks you actually have to run before launch
1. The opener, the disclosure, and the recorded consent
The FCC's February 2024 declaratory ruling pulled AI-generated voices into the TCPA's "artificial or prerecorded voice" framework, so outbound AI voice calls to US numbers follow the same consent standards as traditional robocalls. Statutory damages remain $500 per call, up to $1,500 for willful violations, and class certification is available across the whole campaign. The math is unforgiving.

Two things need to be true on every dial:
- The opener names your company, states the purpose, and (as a matter of operational hygiene, even where not yet finally required) discloses that the caller is talking to an AI.
- Consent is loaded into the campaign runner from your source of truth, and any record without documented consent is suppressed at dial time.
Test the opener by running personas that interrupt in the first two seconds, mumble a hello, or launch straight into "who is this?". If your agent skips the disclosure when the human beats it to the punch, you have a compliance bug wearing a UX skin.
2. Voicemail detection and message-leaving behavior
Outbound campaigns rack up voicemails fast. Most platforms let you configure a voicemail-detection strategy. Vapi, for instance, exposes an audio-based or transcript-based detection mode with tradeoffs on latency and accuracy. Whichever you choose, test it against real-world voicemail greetings, including:
- Standard carrier greetings with the beep at variable positions.
- Human "hello... hello?" openers that sound like a live pickup.
- Custom greetings ("hi you've reached...") that end in "leave a message" only after seven seconds.
- International formats.
Then confirm your agent's message is compliant (identify caller, business, callback number, and honor state-specific rules on prerecorded voicemails) and that a voicemail attempt counts against your retry cap correctly.
3. Retry cadence, calling windows, and DNC
Reaching a prospect takes multiple attempts. Bland cites research suggesting an eight-attempt average to actually reach someone on outbound. That's a lot of surface area to get wrong.
Before launch, confirm:
- Calling windows enforce the recipient's local time (8am to 9pm at minimum; several states are stricter).
- Attempts per record per day and per week are hard-capped.
- DNC lists (federal, state, and your internal "do not contact" set) are scrubbed at dial time, not at list-load time, so a mid-campaign opt-out actually stops the next dial.
- A caller who says "stop calling me" during a live call is written back to the DNC set before the next attempt is queued.
The last one is where pipelines usually leak. Verbal opt-outs get logged to the transcript, nobody parses the transcript into a DNC event, and the same person gets the same call three days later.

4. IVR, gatekeepers, and "who is this?"
Outbound agents don't only reach humans and voicemail. They reach IVRs. They reach spouses. They reach assistants asking screening questions. Test whether your agent:
- Recognizes an IVR and either navigates it, waits for a menu option, or hangs up cleanly instead of monologuing at a bot.
- Knows how to answer "who is this?" and "how did you get this number?" without hallucinating a relationship.
- Handles "this is not [name], they're at work" without leaving details it shouldn't.
These are the moments where an outbound agent looks either professional or badly built, and they never show up in a happy-path demo.
5. Escalation and warm transfer
When the outbound agent qualifies a lead or catches a support case that needs a human, the handoff has to hold. That means the trigger fires when it should (not only when the caller yells "human"), the caller isn't dropped into dead air while the transfer resolves, and the receiving rep has enough context that the caller doesn't repeat themselves. Vapi's warm transfer supports summary or full-transcript context via SIP. Telnyx supports conferenced warm transfers where the AI stays on the line as a third participant. Whichever you use, test the failure modes: transfer target doesn't answer, transfer target picks up voicemail, caller drops during the whisper, caller starts talking during the whisper.
6. Post-call write-backs and disposition truth
For every call, the campaign runner needs to know: what happened, what the disposition was, whether consent was reaffirmed or withdrawn, whether the callback was booked, and whether a human takeover is queued. If write-backs are wrong, the campaign will re-dial contacts it shouldn't and skip contacts it should. Test the write-back for every disposition, including the messy ones: partial answer, hangup mid-disclosure, wrong-number correction, transferred-then-dropped.
Build the harness once, then re-run it forever
The checklist above is not something you run once and forget. Every model swap, every prompt change, every telephony-provider update can regress any of these behaviors silently. Aircall is now shipping monthly changelogs on its AI Voice Agents, and the rest of the ecosystem is on similar cadences. If your only pre-launch test is a founder calling their own number twice, you will find out about the regressions from your customers.
The pattern that scales:
- Encode each of the six checks as a set of simulated scenarios with real personas: different accents, background noise, emotional register, and speech pace.
- Run the whole suite on a schedule (nightly, or gating every deploy) against real telephony, not a text loopback. IVRs, voicemail systems, and human-turn dynamics only reveal themselves over actual audio.
- Score every simulated call and every live call against the same metrics so regressions don't hide in the transition from staging to production.
- When a live call fails a metric, capture the recording and turn it into a replayable regression test.
Where Roark fits
Roark is built for exactly this workflow. Simulations dial your voice agent over real PSTN and WebRTC calls with personas that carry distinct voices, accents, emotional register, and background noise, so the outbound checks above get exercised the way real callees exercise them. Both inbound and outbound are supported: Roark can dial your agent's endpoint, or provision a number your agent dials, triggered over HTTP so runs can gate CI.
Every simulated and live call gets scored on Roark's audio-native models, which pick up pronunciation, emotion, vocal stress, pace and pauses, and interruptions. Sixty-four built-in metrics come with the platform, and you can add custom ones for the specific behaviors an outbound campaign has to nail: disclosure delivery timing, voicemail attempt handling, DNC honor after a verbal opt-out. When a check fails, the call is filed as an issue automatically. Real production calls can be captured and replayed against updated agent logic, so the first campaign that surfaces a bad IVR pattern becomes a regression test your next release has to pass.
One-click integrations exist for Vapi, Retell, LiveKit, Pipecat, Bland, and ElevenLabs, so wiring an existing outbound stack in doesn't require rebuilding it. SOC 2 Type II certification and a HIPAA BAA cover the regulated-industry side, which matters for outbound programs in healthcare, insurance, and financial services.
The rule
Never launch an outbound voice AI campaign whose failure modes you have not rehearsed. The economics work when the list is worked cleanly and the compliance chassis holds. They stop working the first time a plaintiff's firm gets a copy of a transcript in which your agent dialed a DNC number, misidentified itself, and left a voicemail without a callback number. Test the campaign, not just the agent.

