A part of the voice stack gets cheap once you can measure its quality. Transcription had word error rate, so its price dropped first. Evaluation has no shared number yet. So it stays expensive.
Text evaluation scores a final answer. A call is not a final answer. It is a live conversation where timing, audio, and tool actions can each go wrong on their own. None of that shows up in the words. So scoring the transcript passes agents that felt broken to the person on the call.
Every serious tool ended up with the same three parts. So the three parts no longer set anyone apart. They are the baseline. The real difference is in the one step the loop does not cover yet.
Test callers push the agent through accents, noise, interruptions, phone menus, and tool steps before a real customer hears it.
Real calls scored on the same things: latency, task done, policy followed. Traces and audio are kept as proof.
People review the riskiest calls. Their labels tell you whether the automatic scores can be trusted.
These companies did not start from the same place, and where they started tells you their edge and their blind spot. Group them by where they came from, not by feature, since the features are converging. You get four types.
Put the four types against the capabilities that matter and the gaps are clear. The top rows are full. The bottom two are mostly half-circles and blanks. That is not a coincidence. Those two are the hard ones.
| Capability | Voice-native QA loop |
Enterprise CX assurance |
General eval / observability |
Voice-agent platforms |
|---|---|---|---|---|
| Simulated calls synthetic callers pre-launch | ● | ● | ◒ | ◒ |
| Audio-native metrics prosody, latency, barge-in | ● | ◒ | ○ | ◒ |
| Production-call replay re-run real traffic | ● | ◒ | ○ | ◒ |
| Trace & tool-call inspection why it failed, span by span | ● | ○ | ● | ◒ |
| Human-calibrated judgment are the auto-scores trustworthy? | ◒ | ◒ | ◒ | ○ |
| Failure → test → release lineage one durable, auditable object | ◒ | ○ | ○ | ○ |
Most of these tools score calls with a model. But a score is also a model output, and it can be wrong. The scoring model can miss things the same way the agent it grades misses things. Adding more metrics does not fix this. It just gives you more numbers to trust without checking.
Grading a call already costs more than running it, so the money in voice is moving toward evaluation. But not evenly. Metric counts and call volumes are easy to copy and are headed for a price war. The value that lasts is the value you earn with each customer over time.
A fair question: maybe agents are fine and this is fear being sold. The benchmarks say otherwise. An agent that does a task well in text loses a lot of that ability once the task runs over audio, and more when the audio is noisy.
Everyone can measure how often the model agrees with a human. Do buyers actually block a launch on it, or is it just a chart no one acts on? This decides whether calibration is a real moat.
“Replay production calls” assumes you are allowed to. Consent, privacy, and data-residency rules may limit how much of that is usable in regulated industries.
The call, the ticket, the test, the metric, or the release? Whichever one becomes the record of truth shapes the market. It is still unsettled.
If the built-in evals in a platform are good enough for most buyers, the voice-native tools get pushed into regulated, high-stakes work only. That is the whole bull and bear case.
Prior piece in this series · The Voice AI Stack. Source of the 4.8× grade-vs-run figure.
Voice-native QA loop · Coval · Hamming · Roark
Adjacent shapes · Cyara & CCaaS QA suites · Braintrust, LangSmith, Arize Phoenix · Retell, Vapi, ElevenLabs
Research anchors · τ-Voice · EVA-Bench · Testing the Testers
Category map from public product pages, docs, pricing, and preprint benchmarks. Company capabilities are self-reported and not verified in a controlled bakeoff; the four-shape taxonomy and all charts are the author's. No product is ranked or recommended. August 2026.