Product · Human review
Ground truth for every metric.
Your team reviews real calls, sets the correct answer, and Roark measures how closely each metric agrees. Then it tunes the metric to match your judgment. Evals you can defend, because the people who know the work calibrated them.
Start free with $50 in credit, no card needed. Works with Vapi, Retell, LiveKit, Pipecat + your stack.
Scoring production voice AI for teams at


01 · Label, align, tune
Metrics your experts
can vouch for.
An LLM judge is a guess until a human checks it. Roark turns your team’s review into ground truth, scores every metric against it, and tunes until the model and your experts agree.
Your team sets the ground truth
Reviewers open a call with its transcript and audio and mark the correct value for every metric. For segment and turn metrics they can anchor the label to the exact moment. Assign several reviewers and Roark rolls their answers into one settled truth, sending real disagreements to adjudication.
transcript + audio in view · two reviewers · disputes settled inline
See how much each metric agrees
For every metric, Roark compares the model score to your team’s ground truth: an agreement rate, Cohen’s kappa, and the exact calls where they diverge. You learn which evals to trust and which need work, in numbers you can put in front of a customer.
14 disagreements queued · Identity verified is the one to fix
Tune the metric to match you
Confirmed labels become examples the metric learns from, the disagreements first. Re-score the reviewed set and watch alignment climb. Nothing changes in production until you publish the version that agrees with your team.
How Roark closes the loop6 corrective examples added from your labels. Re-scored on the reviewed set. Publish when it agrees.
…and every metric gets more trustworthy with each review.
02 · Built for review teams
Everything a labeling
workflow needs.
Review is a team sport. Roark handles the mechanics so your experts spend their time judging calls, not wrangling a spreadsheet.
Multiple reviewers
Assign a call to several reviewers. The first answer settles it; a second, differing answer opens a dispute.
Inline adjudication
Disagreements surface as Disputed and are resolved in place, so one answer is always the authoritative truth.
Per-metric rubric
Give reviewers the exact rubric for each metric, so labels stay consistent across your whole team.
Moment-level labels
Anchor a label to a quote in the transcript for segment and turn metrics, not just the call as a whole.
Inter-annotator agreement
Measure how much your reviewers agree with each other, and catch a rubric that needs sharpening.
Any metric, any type
Works on everything Roark scores: pass/fail, numeric scores, categories, audio-native and custom alike.
03 · Get started
First call scored in under a minute.
One click on any platform below and production calls stream in on their own, or send any recording with a few lines of code.
Read the quickstartimport Roark from '@roarkanalytics/sdk'const client = new Roark({ bearerToken })await client.call.create({recordingUrl, startedAt,interfaceType: 'PHONE',callDirection: 'INBOUND',agent: { customId: 'support_v2' },}) // scored in seconds
Works with
Enterprise-grade from day one: annual pen tests, SSO/SAML, role-based access, configurable retention.
Bring a recording.
We’ll score it live.
See your own agent measured on the audio it actually produced, in the demo, in real time. Stop guessing whether your voice AI works.
Or start free with $50 in credit · read the docs · support@roark.ai