Product · Human review

Ground truth for every metric.

Your team reviews real calls, sets the correct answer, and Roark measures how closely each metric agrees. Then it tunes the metric to match your judgment. Evals you can defend, because the people who know the work calibrated them.

Start free with $50 in credit, no card needed. Works with Vapi, Retell, LiveKit, Pipecat + your stack.

Scoring production voice AI for teams at

Google
AT&T
BCG
Spectrum
Aircall
Podium
radiantgraph
Google
AT&T
BCG
Spectrum
Aircall
Podium
radiantgraph

01 · Label, align, tune

Metrics your experts can vouch for.

An LLM judge is a guess until a human checks it. Roark turns your team’s review into ground truth, scores every metric against it, and tunes until the model and your experts agree.

01 · Label

Your team sets the ground truth

Reviewers open a call with its transcript and audio and mark the correct value for every metric. For segment and turn metrics they can anchor the label to the exact moment. Assign several reviewers and Roark rolls their answers into one settled truth, sending real disagreements to adjudication.

Review · refund_call_04283 of 3 labeled
MetricModelYou
Identity verifiedPassFail
Empathy3 / 54 / 5
ResolutionResolvedResolved

transcript + audio in view · two reviewers · disputes settled inline

02 · Align

See how much each metric agrees

For every metric, Roark compares the model score to your team’s ground truth: an agreement rate, Cohen’s kappa, and the exact calls where they diverge. You learn which evals to trust and which need work, in numbers you can put in front of a customer.

Alignment · support_v2128 reviewed
Resolution94%
Empathy88%
Identity verified72%
Overall agreementCohen’s κ 0.81

14 disagreements queued · Identity verified is the one to fix

03 · Tune

Tune the metric to match you

Confirmed labels become examples the metric learns from, the disagreements first. Re-score the reviewed set and watch alignment climb. Nothing changes in production until you publish the version that agrees with your team.

How Roark closes the loop
Tune · Identity verifieddraft v2
Agreement with your team72%91%

6 corrective examples added from your labels. Re-scored on the reviewed set. Publish when it agrees.

…and every metric gets more trustworthy with each review.

02 · Built for review teams

Everything a labeling workflow needs.

Review is a team sport. Roark handles the mechanics so your experts spend their time judging calls, not wrangling a spreadsheet.

Multiple reviewers

Assign a call to several reviewers. The first answer settles it; a second, differing answer opens a dispute.

Inline adjudication

Disagreements surface as Disputed and are resolved in place, so one answer is always the authoritative truth.

Per-metric rubric

Give reviewers the exact rubric for each metric, so labels stay consistent across your whole team.

Moment-level labels

Anchor a label to a quote in the transcript for segment and turn metrics, not just the call as a whole.

Inter-annotator agreement

Measure how much your reviewers agree with each other, and catch a rubric that needs sharpening.

Any metric, any type

Works on everything Roark scores: pass/fail, numeric scores, categories, audio-native and custom alike.

03 · Get started

First call scored in under a minute.

One click on any platform below and production calls stream in on their own, or send any recording with a few lines of code.

Read the quickstart
evaluate.ts
import Roark from '@roarkanalytics/sdk'
const client = new Roark({ bearerToken })
await client.call.create({
recordingUrl, startedAt,
interfaceType: 'PHONE',
callDirection: 'INBOUND',
agent: { customId: 'support_v2' },
}) // scored in seconds
Node · Python, plus a REST API for CI/CD and webhooks the instant a call is scored

Works with

Vapi
Bland
Retell
LiveKit
Pipecat
ElevenLabs
Kore.ai
Google
SOC 2Type IIHIPAABAA available

Enterprise-grade from day one: annual pen tests, SSO/SAML, role-based access, configurable retention.

Security details

Bring a recording.
We’ll score it live.

See your own agent measured on the audio it actually produced, in the demo, in real time. Stop guessing whether your voice AI works.

Or start free with $50 in credit · read the docs · support@roark.ai