Last week ElevenLabs disclosed that its ElevenAgents platform now handles more than 15 million conversations a week, three times the volume it reported in February, with agents booking appointments, renewing insurance policies, and processing refunds for enterprise customers. That number, more than any model benchmark, is why containment rate has quietly become the KPI every voice AI program is judged on. If your agent handles the call end to end without a human, you booked the saving. If it transfers, you did not.
The problem is that containment, measured the way most dashboards measure it, is the single easiest voice AI metric to inflate. A caller who gives up and hangs up looks identical to a caller who got a perfect answer. A caller who accepts a confident wrong answer looks even better. If you are a PM reporting containment to the board, there is a reasonable chance the trend line going up is partly your agent getting worse at listening.
The containment identity most dashboards hide
The United States Postal Service runs one of the largest IVR deployments in the country. In 2021 its Office of Inspector General audited the 1-800-ASK-USPS line and found that callers who hang up before being transferred to an agent, even if their issue is not resolved, are counted as contained calls, so the containment rate is overstated. The headline number had climbed to 78 percent, approaching the agency's 80 percent target, while roughly a quarter of post-call survey respondents reported being very dissatisfied with the experience.
The OIG's fix was an accounting change, not a tech change. Report complete contained calls separately from incomplete abandoned ones. The identity underneath is worth memorizing, because it does not change when the IVR becomes a voice AI:
answered calls = contained + transferred
contained = resolved + abandoned-in-flowMost voice AI platforms report the first line and ship the second line as a single bucket. That is how a rising containment rate can mean your agent got worse: callers stopped waiting for transfer and started hanging up, and every hang-up landed in the contained numerator.

Four ways "contained" lies to you
Containment is cheap to game because the default definition rewards silence. There are four distinct failure modes, and in production you usually have all four at once.
Silent abandonment. The caller hangs up in frustration without ever requesting a transfer. The platform sees "no transfer" and counts the call. In the largest public study of AI receptionists so far, disclosing that the caller is speaking to AI was associated with roughly 20 percent fewer hang-ups, and mentioning that the call is recorded was associated with roughly 30 percent fewer hang-ups. Those are not small moves on a core metric.
Confident wrong answers. The caller gets an answer that sounds right, says "okay thanks," and leaves. Containment records a win. They call back the next day, or worse, they do not call back because they trusted the first answer. One consultancy estimates that adjusting for repeat contacts within 24 hours typically reveals a 5 to 15 percent gap between raw containment and true resolution.
Forced resolution. The agent is instructed never to transfer except on explicit request. The caller gives up asking. Containment is 100 percent on that intent. CSAT, if you measured it, would tell you a different story.
Denominator drift. Vendor demos quote containment on curated test calls with ideal audio. One enterprise audit documented a case where a buyer approved a deployment on a quoted 91 percent containment rate and six months later saw the operating dashboard at 58 percent, because the contract never defined what counted as a contained call. Nothing was wrong with the agent. The spec was wrong.

What to measure instead
The version of containment worth reporting is narrower, and it is a composite, not a single rate. Call it contained resolution. Three conditions, all required:
- The call ended inside the AI without a human transfer.
- The caller's stated intent was actually resolved, verified against downstream system state (CRM record updated, booking confirmed, payment captured, policy renewed).
- The same caller did not re-contact on the same intent within a defined window, typically seven days.
Three sub-metrics make this operational.
| Metric | What it catches | Where to source it |
|---|---|---|
| Resolution validity rate | Confident wrong answers | Human audit sample against ground truth, scaled with LLM graders |
| Repeat contact rate on same intent | False closures, incomplete workflows | CRM contact history joined on caller ID, 24h and 7d windows |
| CSAT split by contained vs transferred | Forced resolution, blind transfers | Post-call survey, segmented not aggregated |
Pair those with by-intent containment, not aggregate. Aggregate containment hides the one intent where your agent is quietly collapsing. Published USPS data showed general inquiry at 93 percent, change of address at 56 percent, and specialist queues often under 5 percent. A single org-wide number averages those into noise.
And watch the gap between contained-call CSAT and transferred-call CSAT. If it is wider than a few points, you are either force-resolving or handing off blind. Both are fixable, neither is visible in aggregate.
Why audio-native scoring matters here
The classic text-based way to catch silent abandonment is to look at call duration and the last turn. Short call, agent spoke last, no transfer, no CRM write: probably an abandon. That heuristic works maybe seventy percent of the time. The other thirty percent are the cases that cost you.
The signals that actually separate "resolved and happy" from "gave up and left" are in the audio, not the transcript. A caller whose pitch rises sharply over the last three turns is not a satisfied customer, no matter what the transcript says. A caller who starts speaking before the agent finishes, three times in a row, is signaling that the agent is too slow or too wordy. A caller whose pace drops to half-speed on a readback is confused and papering over it. None of those show up in the words.
This is where production call scoring has to go beyond "did the transcript match the goal." The underlying call recording carries pronunciation, emotion, vocal stress, pace, pauses, and interruptions, and those are the features that separate a happy contained call from a quiet abandonment. Roark's audio-native metrics are built to score every production call on exactly those features, then file an issue when a call looks contained but sounds abandoned. If the trigger fires often, your containment number has been lying to you.
Testing containment before launch, not after
The uncomfortable truth about containment is that you cannot fix it in production dashboards. By the time a call is in your dashboard, the caller has already hung up and formed an opinion of your brand. Fixing it means stressing containment in simulation, before launch, on the scenarios that break it.
A useful pre-launch suite for containment covers five scenario families:
- The confused caller. Vague opening, mid-call pivot, interrupts the agent. Measures whether the agent escalates cleanly or forces resolution.
- The abandoner. Caller trails off, stops responding, gives one-word answers. Measures whether the agent recognizes disengagement or marks the call closed.
- The repeat intent. Caller mentions they called yesterday. Measures whether the agent recognizes repeat contact as a signal that the first call was not actually resolved.
- The non-native speaker. Accented English, moderate background noise, medium pace. Measures whether containment holds up when ASR confidence drops.
- The hostile caller. Rising sentiment, interruptions, demands for a human. Measures whether the agent transfers on request or stonewalls to protect the containment number.
Each scenario needs to be dialed as a real phone call, not run as a text loop, because the behaviors you are testing for are audio behaviors. Barge-in detection, endpoint prediction, interruption handling, and TTS cancellation all behave differently on PSTN than they do in a browser debugger. Roark runs these as personas over real telephony, scores each run on the same metric suite you use in production, and turns a failed scenario into a replayable regression test.

Instrumenting production honestly
Once the agent is live, the containment reporting job splits in two.
The first job is scoring every call, not sampling. A random 2 percent QA sample is enough to catch systemic issues on high-volume intents and nothing else. Low-volume intents, which is where false containment hides, show up zero or one time in that sample. If you ship every call through an audio-native scoring pipeline, the long-tail intents stop hiding, and your repeat-contact and resolution-validity signals get precise instead of directional.
The second job is reconciling the agent's view of "resolved" against the downstream system of record. The agent thinks it booked the appointment. The scheduling system either has that appointment or it does not. The agent thinks it processed the refund. The payments system either has the credit or does not. Containment that is not reconciled against downstream state is containment on the honor system.
A healthier dashboard, built around the same identity the OIG recommended, looks like this:
| Row | Definition | Good target |
|---|---|---|
| Answered calls | Agent picked up | 100% of inbound |
| Transferred | Handed off to human, with reason code | By intent |
| Contained, complete | Reached a terminal state the agent considers resolved | Report separately |
| Contained, abandoned in flow | Caller hung up mid-flow, no terminal state | Report separately |
| Contained resolution | Complete and reconciled downstream and no repeat contact in 7 days | The number worth reporting |
The containment number worth reporting to your board is row five. Everything above it is diagnostic.
What this looks like for a regulated vertical
Regulated verticals like healthcare and insurance make containment honesty harder, not easier, because the cost of false containment is higher. A missed prior-authorization callback becomes a care gap. A policy renewal that the agent "completed" but the policy admin system never received becomes a lapsed policy. For teams working on patient access or healthcare agent flows, the three-condition definition is not optional: it is the only way to know whether your agent is helping or quietly introducing a new failure mode.
The same logic applies to any outbound-first program where the caller never asked for the conversation in the first place. If an outbound renewal agent "contains" a call where the policyholder hung up at hello, the number is not just overstated, it is pointing the wrong direction.
The shortest version
Containment is the right thing to care about. The default way of measuring it is not. The fix is three steps.
- Report containment, abandonment, and transfer separately. Containment plus escalation should only sum to one hundred when abandonment is reported alongside.
- Treat containment as provisional until repeat contact windows and downstream state agree.
- Score every production call on the audio, not just the transcript, so that silent abandonment and forced resolution surface as issues instead of hiding inside a healthy-looking dashboard.
If voice AI really is scaling to tens of millions of conversations a week, the industry's shared metric needs to stop rewarding the thing that is cheapest to game. The number worth defending is not whether the call avoided a human. It is whether the caller got what they called for.

