← Back to omkarray.com
PM Analysis · Voice AI Infrastructure · July 2026

The Voice AI Stack:
what's commoditized,
and what isn't

Everything that turns audio into text into audio now costs about two cents a minute. The market sells that minute for ten. This is a layer-by-layer teardown of where the other eight cents goes — built from vendor list prices fetched on a single day, so the comparison is honest.
15 layers scored $0.021 inference floor Prices fetched 26 Jul 2026
How to read the evidence tags. P published vendor price, fetched 26 July 2026 — every one is linked in the sources at the bottom. M modelled by me from assumptions that are stated in the open, so you can attack them. O observed hands-on while building and testing on the platform itself. INR converted at ₹87/USD. Where a modelled number disagrees with a vendor's own per-minute price, I show both and say which one I trust.
$0.021 Commodity cost of
one voice minute
$0.10 Market clearing
price per minute
80–90% Of a retail minute
that isn't inference
1–3% Share of that minute
spent on the LLM
4.8× Cost to evaluate a call
vs. cost to run it
Commodity cost = cheapest credible vendor at each layer, self-hosted orchestration. Market clearing price = the midpoint of five platforms' all-inclusive connected-minute rates. Both derived in §04.
01 —— The waterline Where commoditization has reached

Every layer of the voice stack is being commoditized from the bottom up. The question worth asking is not whether a layer will fall, but how far the water has already risen — because that line tells you exactly where it is still rational to spend engineering effort, and where you are paying rent on something that has already become a dropdown.

As of July 2026 the waterline sits precisely on turn-taking. Below it — carriage, transport, speech recognition, the good-enough language model, orchestration, baseline speech synthesis — you find three or more credible vendors priced within a whisker of each other, an open-source floor underneath them, and switching costs measured in dropdown clicks. Above it — expressive voice, evaluation harnesses, production ops, regulated telephony, domain depth — you find price set by value rather than cost, and switching measured in quarters.

Exhibit A — The commoditization waterline, July 2026
Fifteen layers ordered by how commoditized they are, not by where they sit in the call path. C = commoditization score (5 = many interchangeable vendors at near-marginal cost, switching is a config change; 1 = few substitutes, price set by value, switching means re-engineering or is legally gated).
↑ Above the line — value still accrues here
5PSTN carriage — US & global English · resold minutesbuy on price
5Media transport · WebRTC / SIPopen source is the default
5Streaming speech recognition — Englishplatforms now auto-select for you
5The "good-enough" LLM brain~1–3% of the bill
5Orchestration engine · cascaded pipelinegenuinely $0 to self-host
4Baseline TTS · clear and pleasantone flat price for five vendors
3Turn-taking, endpointing, barge-in← the frontier
waterline · July 2026
3Native speech-to-speech · realtime audio2–3 players, priced competitively
2Premium TTS · expressive, cloned, Indic-naturalbuyers still pay 11×
2Eval & simulation harnesscommoditizing next
2Production ops & the iterate loopunderserved by everyone
2Indic STT/TTS at 8 kHz · with code-switchingprice is commodity, quality is not
1Compliance & trust · SOC 2, HIPAA, PCIsold as flat rent
1PSTN carriage — India & regulated marketsmonths + a legal entity
1Domain SOPs & system-of-record depthleast commoditized layer in the stack
↓ Below the line — a config change, not a strategy
Note the one inversion. Regulated telephony sits at the physical bottom of the call path — the same layer that scores a 5 in the United States — yet scores a 1 in India. Same technology, opposite market. That is the tell that its scarcity is legal, not technical, and therefore that no model release will dissolve it.

Two things follow immediately. First, the layers below the line are where a buyer should be ruthless about price and a builder should refuse to spend a single sprint. Second, and less obvious: the line is moving upward at a knowable rate, which means a moat you are building today at layer n has a shelf life you can estimate. §10 puts dates on it.

02 —— Why layers fall in this order A rule that predicts the sequence

The ordering above is not arbitrary and it is not simply "cheap things first." Look at what the fallen layers have in common and a single property does most of the explanatory work:

The measurement rule. A layer commoditizes in the order its quality becomes measurable by a number that buyers agree on. Once quality is a scalar, vendors compete on it, converge on it, and are left competing on price. Until quality is a scalar, taste survives — and taste carries margin.

Run the stack against it. Speech recognition has word error rate: a single number, publicly leaderboarded, reproducible on shared audio. It commoditized hardest and fastest. Text-to-speech has mean opinion score, which is a survey, not a measurement — so TTS split cleanly in two: a commodity floor where "clear and intelligible" is effectively binary, and a durable premium tier for expressiveness, where the metric is still a human shrug. Turn-taking only recently acquired numbers — false-barge-in rate, endpoint latency, semantic end-of-turn accuracy — and it is commoditizing right now, in public, at exactly the pace the rule predicts. Evaluation has no agreed scalar at all: nobody can hand you a number for "was this call good" that a second vendor would reproduce. It is the most expensive add-on on any price list. And compliance is not measured, it is attested — audits and contracts, not benchmarks — which is why it never rides a cost curve at all.

This is my framework rather than a vendor's, so treat it as a hypothesis with a prediction attached. The prediction is uncomfortably specific and appears in §10: the eval layer's price collapses within months of the first credible public benchmark for conversational quality, not gradually before it. If you want to know when your eval moat expires, watch for the leaderboard, not the competitors.

Exhibit B — The full gradient, with evidence
Sorted most → least commoditized. "Vendors inside 2× on price" is the operational test: if three or more credible suppliers sit within a factor of two of each other for the same job, the layer has cleared.
C Layer Vendors inside 2× on price Switching cost Verdict
5PSTN carriage
US & global English
Twilio $0.0085/min inbound · LiveKit $0.01 · resold at $0.015configBuy on price. There is no strategy here.
5Media transport
WebRTC / SIP
LiveKit OSS ($0) · Daily · Agora · LiveKit Cloud $0.01/mindaysSolved. WebSocket-only is now a defect, not a design choice.
5Streaming STT
English
$0.0025 · $0.003 · $0.0048 · $0.0065 — five vendors in a 2.6× banddropdownOver. At least one platform doesn't even let you choose — it auto-selects on latency. O
5"Good-enough" LLM
the brain
Gemini 2.5 Flash-Lite $0.10/1M · GPT-5.4-nano $0.20/1M · Sarvam-30B ₹2.5/1MdropdownFree in practice. Compete on tool-calling reliability and p95, never on price.
5Orchestration
cascaded engine
LiveKit Agents · Pipecat · Bolna's MIT-licensed engine — all open sourceweeksGenuinely $0. Anyone charging $0.05/min here is exposed.
4Baseline TTS
clear & pleasant
Aura-1 $15/1M chars · Bulbul v2 ~$17 · Aura-2 $30 — plus five vendors sold at one identical platform pricedropdownCommodity at "good enough". Premium is a different product, not a better one.
3Turn-taking
endpointing, barge-in
Krisp · Deepgram Flux · AssemblyAI native · LiveKit — offered as a menu inside the major platforms Odropdown to pick,
weeks to tune
The frontier. Vendor choice commoditized inside a year; correct tuning is still craft.
3Native speech-to-speech
realtime audio
~$0.018/min vs ~$0.128/min — a 7× gap between two suppliersre-architecturePrice competition has arrived, but two or three players is a market, not a commodity.
2Premium TTS
expressive / cloned
ElevenLabs at $165/1M chars holds a 5.5–11× premium over the commodity tiervoice = brandBuyers still pay 11×. Real differentiation survives here.
2Eval & simulation
the test harness
Best-in-class ships AI-persona multi-turn simulation and batch regression; a direct competitor ships manual review only Orebuild suitesPriced as the most expensive add-on on the sheet. Commoditizing next.
2Production ops
the iterate loop
Nobody has closed the loop well — the gap is industry-wide OhighUnderserved across every platform tested.
2Indic STT/TTS
8 kHz + code-switch
Sarvam · in-house Indian stacks · AI4Bharat lineagequality-gated,
not price-gated
Price is already commodity at ₹30/hr. Quality at 8 kHz is not.
1Compliance & trust
SOC 2 · HIPAA · PCI
Sold as flat rent: $1,000–2,000/mo add-ons, or a $30,000/yr floorre-audit,
re-contract
Not a technology. Does not commoditize on a model curve.
1PSTN carriage
India & regulated markets
DLT registration, 140/160-series numbering, KYC documents Omonths +
legal entity
Hard moat. Where Indian players structurally beat US platforms.
1Domain SOPs
system-of-record depth
Native helpdesk/CRM integration is simply missing at three otherwise-strong platforms Orebuild logicThe least commoditized layer in the entire stack.
Layers are chosen to be mutually exclusive: a vendor's price is counted against exactly one of them, so no cost is double-attributed when the layers are summed into a minute in §04.
03 —— Dispersion is the measurement One unit, one axis, six layers

"Commoditized" is usually asserted. It can be measured. If a layer has truly cleared, the credible suppliers will have converged: the ratio between the cheapest and dearest vendor doing the same job collapses toward one. If a layer hasn't cleared, that ratio stays wide, because buyers are still paying for something price alone doesn't capture.

The obstacle is that these layers are quoted in incompatible units — per minute, per hour, per million characters, per million tokens. So everything below is normalized to one dollar figure per minute of live conversation M, using a single consistent model of what a minute contains: four assistant turns, roughly 2,000 input tokens and 60 output tokens per turn, and about 265 characters of synthesized speech (an agent that holds roughly a third of a 150-words-per-minute conversation). Those assumptions are stated so you can break them; §06 shows where they bend.

Exhibit C — Price dispersion by layer, all normalized to $/minute
Each dot is one shipping vendor model at its published rate. Log scale — equal distances are equal ratios, which is the point. The shaded band on each row marks the quality-equivalent tier where vendors are genuinely substitutable; the ratio at right is the spread inside that band. Hover any dot for the vendor and price.
inside the substitutable tier priced above it — a different product quality-equivalent band
Speech-synthesis and language-model rows are converted from per-character and per-token list prices M; telephony, transport and speech-recognition rows are published per-minute or per-hour rates P. The open-source orchestration floor is a true $0 and cannot be drawn on a log axis — it is marked to the left of the scale break. Full price tables, unconverted, are in §05.

The chart makes the argument better than the prose can. The three tightest bands — telephony at 1.8×, baseline speech synthesis at 2.3×, English speech recognition at 3.0× — are exactly the layers below the waterline. The three widest — orchestration at 5.5× above a free floor, speech-to-speech at 7.1×, and the full range of speech synthesis at 11× — are exactly the layers where the argument about value is still live.

The clearest single signal on the page

One price for five suppliers

A major platform prices five different speech-synthesis vendors at one identical $0.015/min, and a sixth — the premium one — at $0.040. When a reseller stops distinguishing between suppliers in its own price list, it is telling you those suppliers are fungible. That is commoditization stated by someone with every incentive to deny it.

The inversion nobody expects

Platforms take their voice margin on the cheap tier

Work the same numbers backwards M. That flat $0.015/min sits about 3.7× above the underlying commodity rate of ~$0.004/min. The $0.040/min premium line sits at roughly 0.9× — essentially at cost. The expensive voice is the honest one; the margin is hidden in the cheap tier, where you assume there is none.

That inversion is worth sitting with, because it generalizes. Platform margin does not live where prices look high. It lives where a buyer has stopped comparing — and buyers stop comparing precisely once a layer feels commoditized. The moment you accept that a layer is a commodity is the moment you stop auditing its price.

04 —— The anatomy of a voice minute Where ten cents actually goes

One minute of inbound support, built three ways, every figure drawn from list prices fetched on the same day. The first column is what the components cost. The second and third are what you can buy the same minute for.

Exhibit D — One inbound minute, three ways
Stacked to the same linear scale. The vertical rule marks the inference floor — the total cost of every component needed to actually conduct the conversation. Hover any segment for its rate.
telephony speech recognition language model speech synthesis orchestration & platform fee
Commodity build: cheapest credible vendor per layer, open-source orchestration self-hosted, with a nominal allowance for the compute it runs on. Retail stacks are one platform's own published component rates, where speech recognition is folded into the engine fee and therefore not separately visible P.

Cross-check the middle bar against what the market actually charges for a whole bundled minute and it holds: an Indian platform sells an all-inclusive connected minute at ₹6 (about $0.10), another lands between $0.07 and $0.15, a third publishes a headline band of $0.07–0.31, and an enterprise contract bottoms out around $0.07 at volume O. The market clears at roughly ten cents against a two-cent component floor.

Any business plan whose savings thesis is "we'll switch to a cheaper model" is optimizing the smallest line on the bill. In a commodity-built minute the language model is one to three percent of the cost. You could make the brain free and the minute would barely move.

There is a second reading of the same chart that matters more for anyone selling minutes. Notice that a whole bundled retail minute at $0.10 costs less than two and a half times the premium voice line alone in the third bar. Platform pricing is a packaging decision far more than it is a cost floor — which is exactly why it is vulnerable.

05 —— The verified price sheet Fetched 26 July 2026

Everything above rests on these numbers, so here they are unconverted, in the units the vendors publish them in. Every row is a list price from the vendor's own pricing page on a single day.

Exhibit E.1 — Streaming speech recognition
Five credible vendors inside a 2.6× band. This layer is finished.
Vendor / modelPublishedPer minuteNote
AssemblyAI Universal-Streaming$0.15/hr$0.0025Cheapest credible streaming ASR
OpenAI gpt-4o-mini-transcribe$0.0030
Deepgram Nova-3 mono$0.0048$0.0042 on the Growth tier
Deepgram Nova-3 multilingual$0.0058
Sarvam STT (Indic)₹30/hr~$0.0060At parity with English ASR
Deepgram Flux — turn-aware$0.0065+35% over Nova-3 — the premium is for endpointing, not transcription
AssemblyAI Universal-3.5 Pro Realtime$0.45/hr$0.0075
The 35% Flux premium is the most informative number in this table. The market has unbundled "hearing the words" from "knowing when the human stopped talking," and only the second one still carries margin. Transcription is done as a business.
Watch-out: at least one vendor bills WebSocket session duration rather than audio duration — idle connections bill. For bursty inbound traffic the real cost runs above the headline.
Exhibit E.2 — Speech synthesis
Not one market. A commodity floor and a premium tier that has never had to discount.
Vendor / modelPer 1M charsvs cheapest≈ per minute M
Deepgram Aura-1$151.0×$0.004
Sarvam Bulbul v2₹15/10k ≈ $171.1×$0.005
Deepgram Aura-2$302.0×$0.008
Sarvam Bulbul v3 (beta)₹30/10k ≈ $342.3×$0.009
ElevenLabs — Creator$916.1×$0.024
ElevenLabs — Pro / Scale / Business$16511.0×$0.044
ElevenLabs — Starter$20013.3×$0.053
The pricing stops improving with volume. Above the entry Creator tier, Pro, Scale and Business all land at roughly the same $165/1M. A vendor that does not need to discount at scale is a vendor that does not feel commodity pressure — and that, not the headline rate, is the evidence that this tier hasn't fallen.
Exhibit E.3 — The language model
Priced per million tokens; converted at 8,000 input and 240 output tokens per minute of conversation M.
Model$/1M in$/1M out≈ per minute M
Sarvam-30B₹2.5 ≈ $0.029₹10 ≈ $0.115~$0.0003
Gemini 2.5 Flash-Lite$0.10$0.40~$0.0009
GPT-5.4-nano$0.20$1.25~$0.002
Gemini 3.5 Flash-Lite$0.30$2.50~$0.003
GPT-5.4-mini$0.75$4.50~$0.007
Gemini 3.6 Flash$1.50$7.50~$0.014
GPT-5.6-terra / GPT-5.4$2.50$15.00~$0.024
GPT-5.5$5.00$30.00~$0.047
Shaded rows are the "good-enough" band — models that reliably hold a support conversation and call tools. Inside that band the spread is about 3.3×, and the whole band costs less than a third of a cent a minute.
Exhibit E.4 — Telephony, transport and platform fees
ItemPriceWhat it tells you
Twilio — US inbound local$0.0085/minThe true carriage floor
Twilio — US outbound local$0.0140/min
Twilio — US toll-free inbound$0.0220/min
Twilio — US local number$1.15/mo
LiveKit Cloud — telephony inbound$0.01/minResold carriage, near cost
LiveKit Cloud — agent session$0.01/minManaged orchestration, priced at a fifth of the platforms
LiveKit media server + Agents framework$0 — open sourceThe floor under the entire orchestration layer
Retell — voice engine$0.055/minThree independent vendors converging on ~$0.05 for a layer with a $0 floor is not a cost floor. It is a price umbrella.
Vapi — platform fee$0.05/min
Bolna — platform fee~$0.05–0.06/min
06 —— Two structural taxes Why the cheap paths stay expensive

Two forces keep the commodity layers from collapsing all the way to zero. Both are structural rather than commercial, which means no amount of vendor competition removes them.

The context tax — and an honest disagreement with my own model

My model puts a good-enough brain at roughly $0.0009 per minute. One platform's own price list puts the same class of model at $0.081 per minute — about six times higher P. I am not going to paper over that gap, because both readings are actionable and I cannot yet distinguish them from public data.

Either platforms mark the commodity brain up several-fold, or a real production voice context is three to six times heavier than a naive estimate — long system prompts, retrieved knowledge-base chunks, and a full transcript re-sent on every single turn. My instinct is that the truth is mostly the second: a voice agent re-uploads its entire conversation with each reply, so token consumption grows quadratically with call length while the conversation only grows linearly. Either way the lesson is identical and it is the single most useful operational takeaway in this piece:

Context size, not model choice, is your language-model cost lever. Trimming a system prompt or capping transcript history moves the bill by multiples. Switching from one good-enough model to another moves it by fractions of a cent.

The audio tax — visible in the vendors' own price lists

Speech-to-speech models promise to delete the pipeline: no separate recognition, no separate synthesis, no seam. The reason they cannot simply undercut a cascaded stack is printed in the price list of the company that makes both.

Exhibit F — The same model, charged differently for text and audio
Input token price, one vendor, two modalities, same model family. Audio costs more per token by construction — an audio second carries far more tokens than a second of reading.
Prices per million input tokens P. The multiple, not the absolute level, is the durable fact: it survives every generation of price cut because it reflects how much information a second of audio actually is.

Stack that tax on top of the fact that a speech-to-speech model re-sends accumulating audio history and the economics resolve cleanly. The cheapest realtime audio model lands at about $0.018/min — genuine parity with a cascaded pipeline. The most expensive lands at about $0.128/min, roughly ten times a cascaded build and a seven-fold gap between two suppliers of the same capability M. One vendor has collapsed speech-to-speech pricing sevenfold below the other; that is price competition arriving, not a commodity forming.

So the prediction here is narrow and, I think, safe: speech-to-speech wins on naturalness well before it wins on cost, and cascaded pipelines stay the production default for anything touching money or compliance — because a cascaded stack hands you an inspectable transcript at every hop, and a native audio model does not. Expect blended architectures, not a switchover.

07 —— The margin confession Priced by the people who know

Here is the most useful thing on any voice platform's pricing page: the add-ons. Each one is a line item the vendor believes it can charge for separately, which makes the price list a direct readout of what the market still treats as scarce. Measured against the two-cent commodity minute, the readout is stark.

Exhibit G — Per-minute add-ons, priced as multiples of a whole commodity minute
The reference line is 1.0× — the entire cost of conducting the call. Anything to the right of it costs more than the conversation itself.
All add-on rates are published per-minute prices from one platform's pricing page P, divided by the $0.021 commodity minute derived in §04.

Read the top bar again. Evaluating a call is priced at nearly five times the cost of running it, and roughly twice the platform's own voice-engine fee. In almost every other part of software, testing is the cheap part — you run the test suite a thousand times because each run costs nothing. In voice AI the relationship is inverted, and that inversion is the single clearest statement anyone in this industry has made about where the scarcity is.

The rest of the margin is sold as rent rather than usage, which tells you it is access being priced, not compute:

Exhibit H — Margin sold as flat rent
Add-onPriceWhat the pricing model reveals
HIPAA mode$2,000/moCompliance is sold as access, not usage — the marginal cost of a compliant minute is the same as a non-compliant one
Zero Data Retention$1,000/moMutually exclusive with HIPAA mode on the same account O — a packaging decision, not a technical constraint
HIPAA (different vendor, very different scale)$1,000/moThe same pattern reappears at a company an order of magnitude smaller
Concurrency above 20 channels$8/mo per slotCapacity guarantees still command rent
"Stable server cluster"+$0.02/minReliability itself is a paid upgrade — priced at roughly one entire commodity minute
Enterprise floor at one vendor$30,000/yrCompliance, no-code tooling and owned telephony bundled as a single access fee O
Read as a portfolio: inference sold at cost, margin taken on evaluation, compliance, reliability and capacity. That is a precise map of the un-commoditized surface of this industry — and the vendors drew it themselves, in public, on their own pricing pages.
08 —— The four things that aren't commoditized And why each one holds

Strip out everything the price sheet has settled, and four things are left standing. None of them is a model. Two of them are not even technology.

Moat 01 · geographic

Regulated telephony and local-language quality

India requires DLT registration, 140/160-series numbering and KYC documentation before you can originate a call at all. Onboarding runs into it immediately: one platform demanded Truecaller verification, another a legal business name and registration certificate before it would provision anything O. Meanwhile one major US platform provisions numbers in the US and Canada only, and another's free numbers are US-only O. No model release changes this. It is the structural reason Indian platforms win Indian deployments that US platforms cannot serve at any price.

The nuance the price sheet forces: Indic is no longer expensive, merely uneven. Indic speech recognition at ₹30/hr is at parity with English, and Indic synthesis undercuts the mid-tier English vendors. The moat migrated from price to paperwork — and to quality at 8 kHz with Hinglish code-switching. Anyone still pricing Indic as a premium is charging for a moat that already eroded.

Moat 02 · legal

Compliance and trust

SOC 2, HIPAA, PCI DSS, GDPR and ISO 27001 are audits and contracts, not capabilities. They do not ride a model cost curve because there is no model involved — which is precisely why they are the most durable line on any price list.

The split this produces is sharp and visible. One platform holds the full certification set and charges a $30,000/yr floor for access to it. A technically comparable competitor asserts no certifications it publicly holds, answering the question with "contact support" O. Near-identical technical stacks, opposite market access. Every regulated buyer is filtered by a document, not a benchmark.

Moat 03 · measurement

The eval harness and the iterate loop

This is the widest capability spread anywhere in the space. At the top: AI-persona multi-turn simulation, batch regression through an API, A/B traffic splitting, and production quality assurance that reports resolution rate and flags hallucinations. At the bottom, among otherwise credible competitors: manual transcript review, no batch simulation, no A/B, no regression suite at all O.

When one dimension ranges from best-in-class to nonexistent across platforms that are otherwise near-identical, that dimension is where the real product work is. The $0.10/min quality-assurance line confirms it commercially. By the measurement rule in §02, this moat holds only until someone publishes a benchmark everybody accepts.

Moat 04 · integration

Domain SOPs and system-of-record depth

The most telling gap in the entire landscape: native helpdesk and CRM integration is missing at three otherwise strong platforms O — all of them pushing customers toward webhooks and Zapier for the one thing an inbound support desk cannot function without.

Meanwhile the platforms that do ship it ship it deeply: help-centre knowledge-base import at one, two-way Salesforce and HubSpot sync at another. Nobody's model quality decides these deals. Their integration surface does. This is the least commoditized layer in the stack and the one where a specialist can still beat a better-funded generalist outright.

Notice what these four have in common: not one of them gets cheaper when inference gets cheaper. Three of them get more valuable, because a world of cheap capable agents is a world with more calls to evaluate, more regulations to satisfy, and more systems of record to write into.

09 —— The in-house builder's trap A pattern, not a company

There is a failure mode I keep running into when large product organizations build voice agents internally, and it is the exact inverse of what you would expect. In-house teams are usually accused of reinventing infrastructure. What they actually do is subtler and more expensive: they build the un-commoditized layers rather well, and then dramatically under-buy the commodity ones.

The archetype looks like this. The team has invested real effort in a scenario harness — synthetic test audio, an adversarial test model, generated conversation scenarios — and in an operations cockpit that beats commercial platforms outright on the parts that matter after launch: product insights, repeat-caller detection, per-line-of-business summarization. Those are precisely the layers above the waterline. That is money spent correctly.

And then the commodity layers sit frozen at whatever was wired in during the first sprint:

Exhibit I — The under-buying pattern
Each row is a commodity layer left at its first-sprint default, against what the market now offers for it.
Typical in-house stateWhat the commodity market offers todayThe move
One inference provider, two modelsEight or more providers; the good-enough brain at ~$0.001/minAdd a cheap-fast tier and a fallback. Pure latency and cost win, no lock-in incurred.
One fixed premium voiceCommodity synthesis at $15–17/1M chars against $165/1M for the premium tier — roughly 10×Keep the premium voice where brand actually matters; move high-volume lines of business to the commodity tier.
No language selection — often an empty dropdownIndic speech recognition at ₹30/hr, at price parity with EnglishPopulate it. This is a configuration gap, not a cost problem — and in India the caller base is multilingual by default.
No turn-taking or barge-in controlsNow an off-the-shelf dropdown across every major orchestration frameworkBuy it; do not engineer it. This layer is commoditizing underneath you as you read this.
One fixed telephony provider, outbound onlyRegulated-India telephony is a moat you already sit insideThe one place a fixed choice is defensible — but shipping outbound-only forfeits the moat's actual value, which is inbound.
Composite drawn from hands-on evaluation of in-house and commercial platforms O. The pattern is consistent enough to be worth stating as a rule.
The strategic picture this produces is unusually clean, and it is good news: the weak dimensions are all cheap commodity fixes, and the strong ones are already the right places to be spending. An in-house team in this position should aggressively buy down its commodity layers and keep investing in scenarios and operations.

The trap, stated plainly: a fixed default at a commodity layer feels like a decision that was made. It usually isn't — it is a decision that was never revisited after the layer beneath it collapsed in price. The audit is worth running annually, and it takes an afternoon.

A commoditization map is only worth writing if it makes claims that can later be checked and found wrong. So here is the waterline as a moving object, with dates attached and — more importantly — with the specific observation that would prove each prediction false.

Exhibit J — The waterline over time
The height of the line is the highest layer that has cleared. Solid to the left of today, dashed and hollow to the right — those are forecasts, not observations.
Dates for cleared layers are when three or more credible vendors first sat inside a 2× price band with an open-source floor beneath them. Forecast dates are mine M.
Exhibit K — The prediction ledger
Written down so they can be checked, which is the only kind of prediction worth making.
HorizonPredictionWhy I believe itWhat would falsify it
≤ 12 mo
by mid-2027
Semantic endpointing is bundled free into every speech-recognition vendor and stops being a differentiator. It is already a vendor dropdown inside the major platforms O, and one vendor is currently selling it as a 35%-premium SKU. Premium SKUs for a feature four competitors also ship do not survive a year. The endpointing premium still exists at mid-2027, or turn-taking quality visibly diverges between vendors on a shared benchmark.
12–24 mo
by mid-2028
Simulation and batch regression become table stakes, and the $0.10/min quality-assurance price collapses toward $0.01–0.02. Three platforms shipped multi-turn simulation within about a year of each other, and one already meters evaluation at ₹2/min O. Convergent shipping at that pace is what a layer looks like just before it clears. AI quality assurance still prices above $0.05/min in mid-2028 — which would mean conversational quality resisted measurement longer than the measurement rule predicts.
12–24 mo Platform orchestration fees compress from ~$0.05/min toward $0.02–0.03. A managed vendor already clears the same function at $0.01/min and open source clears it at $0. Three independent vendors landing on exactly $0.05 is a price umbrella, and umbrellas break. Fees hold at $0.05+ through 2028, which would mean the bundled reliability and support around orchestration is worth more than I think.
ongoing Speech-to-speech wins naturalness before it wins cost; cascaded stays the default for regulated and money-touching flows. The cheapest realtime audio model already reaches cascaded parity at ~$0.018/min, but the 2–3× audio-token tax is structural and transcript inspectability is weak. A native audio model ships per-turn inspectable transcripts and undercuts a cascaded stack on price. Both, not either.
The meta-prediction, from §02: the eval layer's price will not decay smoothly — it will fall off a cliff within a couple of quarters of the first public benchmark for conversational quality that buyers accept. If you are betting on an evaluation moat, the leaderboard is the thing to watch, not your competitors.
11 —— So what Three readers, three conclusions
If you're buying a platform

Stop diligencing model menus

Every serious platform now offers the same recognition, synthesis and language models, and the caller cannot tell which one you picked. Treat a vendor differentiating on "we support 20+ models" as differentiating on a commodity.

Spend the diligence where the spread is real: the eval harness, the production ops loop, compliance posture, and telephony for your specific geography. Those four are where otherwise-similar platforms range from best-in-class to nonexistent.

If you're pricing a platform

Your minute is your exposure

The ~$0.05/min orchestration fee is the most exposed revenue line in the industry — a five-times markup on a layer with a genuine $0 floor and two credible substitutes. The add-ons are the opposite: defensible, and priced by scarcity you did not manufacture.

Margin is migrating from the minute to the outcome — to evaluation, compliance, reliability guarantees and vertical integration. Price accordingly, before the umbrella breaks and you have to.

If you're building in-house

Buy the bottom of the stack with both hands

Everything below the waterline should be a purchase decision revisited annually, not an architecture decision made once. Everything above it is where your team's time converts into something a vendor cannot sell you: your evaluation suite, your operations loop, your domain SOPs, and your integration into the system of record that actually resolves the customer's problem. The failure mode is not building too much — it is building the right things while quietly overpaying for the wrong ones.

12 —— What I couldn't verify Read this before quoting anything above

A price sheet dated to a single day is only useful if it is equally explicit about what isn't on it. The following was not verified in this pass and is deliberately absent rather than guessed at. Several of these are load-bearing.

GapWhy it matters
Indic quality benchmarksNo word-error-rate or opinion-score evidence for Hinglish code-switching at 8 kHz. The claim "Indic price is commodity, Indic quality is not" rests on pricing plus hands-on impressions, not on a measured comparison. This is the highest-value gap on the list — it is the one that decides real deployment choices in India.
Open-weight speech modelsOpen recognition and synthesis models, and their standings on public leaderboards, were not checked. The open-source floor beneath the recognition and baseline-synthesis layers is therefore asserted from price evidence only, not from measured quality.
Contact-centre incumbentsThe established CCaaS vendors bundling voice agents into existing seats — and outcome-based pricing models — are unassessed. That is commoditization arriving from above, and this analysis only looks upward from the infrastructure.
Funding, M&A and consolidationNo coverage. Any claim here about who outlasts whom would be unsupported.
Two vendors' per-minute ratesOne synthesis vendor publishes only credit bundles, no per-character rate and no latency claim — it appears in this analysis solely as filtered through a reseller's price. Another publishes synthesis credits but no conversational per-minute rate.
India regulatory change in 2026Current rules on automated and AI-originated calls, and caller-name presentation rollout, were not re-checked. The telephony moat claim rests on hands-on onboarding friction O — good evidence the moat exists, weaker evidence of its present legal detail.
Every M figure is mine, not a vendor'sThe token and character assumptions are stated in §03 precisely so they can be attacked. The six-fold disagreement documented in §06 is the weakest link in the arithmetic, and I have flagged it rather than smoothed it.
Prices move. Everything here is a snapshot of 26 July 2026 and should be re-fetched before any of it is used in a decision. The structure of the argument — dispersion as measurement, the measurement rule, margin migrating to the outcome — should outlast the specific numbers.

Read next — the rest of the Voice AI series
This piece is the market-level view. The three below are the platform-level teardowns it generalizes from, and the thesis is the argument for the layer this analysis says is scarcest.
Bolna AI →200K calls/day, sub-600ms pipeline, and an MIT-licensed orchestration engine — the open-source floor under layer five, shipped by a company that sells the layer above it.
Smallest.ai →The full-stack counter-argument: own the models and the commodity curve works for you. 20× cost reduction to $0.01/min.
Ringg.ai →Distribution-first, ~85% assembled from vendor APIs — the purest test of whether a GTM moat outlasts the commodity curve.
Mixpanel for Voice AI →The thesis this analysis keeps arriving at from the pricing side: the observability and evaluation layer is the scarce one, and nobody has built it properly.

Sources — all list prices fetched 26 July 2026: deepgram.com/pricing · assemblyai.com/pricing · elevenlabs.io/pricing · cartesia.ai/pricing · twilio.com/voice/pricing/us · ai.google.dev/gemini-api/docs/pricing · developers.openai.com/api/docs/pricing · livekit.io/pricing · docs.sarvam.ai — pricing · retellai.com/pricing · vapi.ai/pricing.
Open-source orchestration floor: livekit/agents · pipecat-ai/pipecat · bolna-ai/bolna.

Hands-on observations O come from building and testing on these platforms directly, July 2026. Modelled figures M use the token and character assumptions stated in §03. Colour scales on every chart were validated for colour-vision deficiency separation before publication. Bengaluru, July 2026.