← Back to omkarray.com
Category Analysis · Voice AI Evaluation · August 2026

Was this call
good?

It is the only question in the voice stack without a number. Transcription has word error rate. Latency has milliseconds. The minute has a price. “Good” has none. So a new set of tools is being built to score it. This page maps them.
01 Why evaluation stays expensive Countable first, priced last

A part of the voice stack gets cheap once you can measure its quality. Transcription had word error rate, so its price dropped first. Evaluation has no shared number yet. So it stays expensive.

Exhibit 1. Layers get cheap in the order they became measurable
Each layer got a number. Once it had one, vendors could undercut each other on it. Evaluation still has no number. That is why it still sells on value, not cost.
The point: you cannot start a price war over something nobody can measure. Give “good” a trusted number and evaluation gets cheap too. The whole category is a race to build that number first.
02 What a transcript can't see Text scores the words; voice fails elsewhere

Text evaluation scores a final answer. A call is not a final answer. It is a live conversation where timing, audio, and tool actions can each go wrong on their own. None of that shows up in the words. So scoring the transcript passes agents that felt broken to the person on the call.

Exhibit 2. Eight ways a call fails, and how much of each a transcript catches
Ranked by how much of each failure a text-only score can see. The transcript covers the top of the list and almost none of the bottom. Everything in red is a failure the words do not show. Hover any bar.
In short: a voice evaluation tool is a promise to measure the hard parts: latency, interruptions, silence, and tool actions. The transcript is the easy ten percent.
03 The three parts every tool has Simulate · observe · review

Every serious tool ended up with the same three parts. So the three parts no longer set anyone apart. They are the baseline. The real difference is in the one step the loop does not cover yet.

Exhibit 3. The three parts, and the gap between them
The loop every tool now runs. The three boxes are settled. The dashed step, turning a real failure into a lasting test, is still done by hand in tickets and spreadsheets.
Before launch

Simulate

Test callers push the agent through accents, noise, interruptions, phone menus, and tool steps before a real customer hears it.

In production

Observe

Real calls scored on the same things: latency, task done, policy followed. Traces and audio are kept as proof.

With humans

Review

People review the riskiest calls. Their labels tell you whether the automatic scores can be trusted.

↵ the gap:  reviewed failure → saved test → proof the fix holds. This is the step that pays off over time, and the step no tool has made simple.
04 Four types of companies Grouped by where they started

These companies did not start from the same place, and where they started tells you their edge and their blind spot. Group them by where they came from, not by feature, since the features are converging. You get four types.

Exhibit 4. The four types
The description under each name is the company's own. The risk line is mine. Names are here to place the market, not to rank it.
Shape 1

Voice-native QA loop

Purpose-built for voice: realistic call simulation, audio-native metrics, production monitoring, regression tests.
CovalSimulate / observe / review across voice + chat; metric versioning + human-review calibration; pricing published.
HammingOne-click production-call replay; high-concurrency test calls; 50+ metrics (self-reported).
RoarkProduction issue → drafted fix → simulation → verified deploy; audio-native metrics (self-reported).
Risk: they overlap with each other. Replay, audio metrics, and regression are now claimed by all three.
Shape 2

Enterprise CX assurance

Incumbents in contact-center and IVR testing, extending existing suites toward AI agents.
CyaraEstablished CX / telephony assurance; deep IVR and load testing; enterprise procurement in place.
CCaaS QA suitesIncumbent quality-management vendors adding generative-agent coverage to installed footprints.
Risk to pure-plays: can bundle agent testing into a contract the buyer already signed. Distribution, not product.
Shape 3

General eval & observability

General LLM evaluation and tracing tools with developer adoption, adding voice as an input.
BraintrustDatasets, scorers and experiment tracking for LLM apps; strong developer workflow.
LangSmithTracing + eval tied to a widely-used orchestration ecosystem.
Arize PhoenixOpen-source-rooted LLM observability and evaluation.
Risk to pure-plays: they own the wider AI stack and can add voice. The buyer may not want a separate tool.
Shape 4

Voice-agent platforms

Runtimes that already hold the call data and config, adding “good-enough” evals inside the platform.
Retell / VapiVoice-agent build-and-run platforms with native access to call logs and configuration.
ElevenLabs & runtimesModel / platform players expanding upward into testing and monitoring.
Risk to pure-plays: own the runtime. Adequate bundled evals may remove the reason to buy a standalone tool.
Grouped by origin, not quality. The pure-plays are betting that voice is hard enough that a general tool or a built-in feature will not be trusted for a go-live decision. Whether that holds is the main question in this market.
05 Where they overlap The capability heatmap

Put the four types against the capabilities that matter and the gaps are clear. The top rows are full. The bottom two are mostly half-circles and blanks. That is not a coincidence. Those two are the hard ones.

Exhibit 5. Which type has which capability
● core to the type   ◒ partial or emerging   ○ mostly absent. Based on public positioning, self-reported. The highlighted rows are where the market is still open.
Capability Voice-native
QA loop
Enterprise
CX assurance
General eval
/ observability
Voice-agent
platforms
Simulated calls synthetic callers pre-launch
Audio-native metrics prosody, latency, barge-in
Production-call replay re-run real traffic
Trace & tool-call inspection why it failed, span by span
Human-calibrated judgment are the auto-scores trustworthy?
Failure → test → release lineage one durable, auditable object
A capability everyone calls “core” no longer sets anyone apart. Look at the empty cells instead. Calibration and lineage are where the market is still open.
06 The scores come from a model too And a model can be wrong

Most of these tools score calls with a model. But a score is also a model output, and it can be wrong. The scoring model can miss things the same way the agent it grades misses things. Adding more metrics does not fix this. It just gives you more numbers to trust without checking.

Exhibit 6. The model score vs. the human score
Illustrative. Each dot is a call. Left to right is the model score. Bottom to top is the human score. On the diagonal the model agrees with the human. The shaded corner is where the model says pass and the human says fail. That is the worst case: a false pass. Hover a dot.
What this means: a metric that blocks a release needs a known agreement rate with humans, and it should say “not sure” when it is not sure. Not a bigger dashboard.
More metrics is a marketing claim. Calibrated metrics is a trust claim. Only the second one holds up in a compliance review.
07 Where the durable value is Easy to copy vs. hard to copy

Grading a call already costs more than running it, so the money in voice is moving toward evaluation. But not evenly. Metric counts and call volumes are easy to copy and are headed for a price war. The value that lasts is the value you earn with each customer over time.

Exhibit 7. Grading a call costs more than running it
Per minute: what it costs to run the agent versus what the market charges to evaluate it. From The Voice AI Stack. Testing is now the expensive part.
Exhibit 8. What is easy to copy, and what is not
The parts of a voice evaluation product, ranked by how hard each is to copy. Below the line: easy to copy, headed for a feature race. Above the line: earned per customer, hard to copy.
↑ Above the line: value builds up here
5Raw metric count · “50+ scores”anyone can add one
5Simulated-call volume · “50K concurrent”capacity, not insight
5Protocol / channel coverage · SIP, WebRTC, DTMFstandard within a year
4Trace & audio inspectionconverging fast
3Calibrated, human-checked scoring← the hard part
the line: easy below, earned above
2Evidence lineage · failure → test → release, one objectmostly unbuilt
1Quality policy as code · the go/no-go gate, versionednobody owns it cleanly
↓ Below the line: easy to copy, headed for a price war
The number is how hard it is to copy, not how good it is. 5 means anyone can copy it, so it gets cheap. 1 means it builds up per customer and resists copying. The bottom two rungs, linking a real failure to the test that proves it will not happen again, are the obvious next thing to build. They are also the emptiest cells on the map above.
The features that demo best, like metric counts and call volumes, are the ones headed for a price war. The plain middle, calibrated scoring and evidence lineage, is what builds up per customer, and almost no one has built it.
08 The problem is real What the benchmarks say

A fair question: maybe agents are fine and this is fear being sold. The benchmarks say otherwise. An agent that does a task well in text loses a lot of that ability once the task runs over audio, and more when the audio is noisy.

Exhibit 9. Ability drops from text to audio
Direction, not exact values. The same task, moved from text to clean audio to noisy audio, loses ability at each step. Hover a bar.
Directional read across recent voice benchmarks: τ-Voice, EVA-Bench, and Testing the Testers (which tests the evaluators). Exact numbers vary by setup. The steady finding is that the text-to-voice gap is large and comes mostly from how the agent behaves over audio.
The problem is real. It lives in the audio and timing that transcripts ignore, and the scoring models need checking too. So the demand is clear. The open question is which type of company, and which part of the stack, ends up owning it.
09 Four open questions Not answerable from public info today
Open question

Does calibration actually gate releases?

Everyone can measure how often the model agrees with a human. Do buyers actually block a launch on it, or is it just a chart no one acts on? This decides whether calibration is a real moat.

Open question

Can real audio be legally replayed?

“Replay production calls” assumes you are allowed to. Consent, privacy, and data-residency rules may limit how much of that is usable in regulated industries.

Open question

Which record becomes the source of truth?

The call, the ticket, the test, the metric, or the release? Whichever one becomes the record of truth shapes the market. It is still unsettled.

Open question

Do bundled platform evals win by default?

If the built-in evals in a platform are good enough for most buyers, the voice-native tools get pushed into regulated, high-stakes work only. That is the whole bull and bear case.


Sources & method Public pages & research

Prior piece in this series · The Voice AI Stack. Source of the 4.8× grade-vs-run figure.

Voice-native QA loop · Coval · Hamming · Roark
Adjacent shapes · Cyara & CCaaS QA suites · Braintrust, LangSmith, Arize Phoenix · Retell, Vapi, ElevenLabs
Research anchors · τ-Voice · EVA-Bench · Testing the Testers

Category map from public product pages, docs, pricing, and preprint benchmarks. Company capabilities are self-reported and not verified in a controlled bakeoff; the four-shape taxonomy and all charts are the author's. No product is ranked or recommended. August 2026.