Every layer of the voice stack is being commoditized from the bottom up. The question worth asking is not whether a layer will fall, but how far the water has already risen — because that line tells you exactly where it is still rational to spend engineering effort, and where you are paying rent on something that has already become a dropdown.
As of July 2026 the waterline sits precisely on turn-taking. Below it — carriage, transport, speech recognition, the good-enough language model, orchestration, baseline speech synthesis — you find three or more credible vendors priced within a whisker of each other, an open-source floor underneath them, and switching costs measured in dropdown clicks. Above it — expressive voice, evaluation harnesses, production ops, regulated telephony, domain depth — you find price set by value rather than cost, and switching measured in quarters.
Two things follow immediately. First, the layers below the line are where a buyer should be ruthless about price and a builder should refuse to spend a single sprint. Second, and less obvious: the line is moving upward at a knowable rate, which means a moat you are building today at layer n has a shelf life you can estimate. §10 puts dates on it.
The ordering above is not arbitrary and it is not simply "cheap things first." Look at what the fallen layers have in common and a single property does most of the explanatory work:
Run the stack against it. Speech recognition has word error rate: a single number, publicly leaderboarded, reproducible on shared audio. It commoditized hardest and fastest. Text-to-speech has mean opinion score, which is a survey, not a measurement — so TTS split cleanly in two: a commodity floor where "clear and intelligible" is effectively binary, and a durable premium tier for expressiveness, where the metric is still a human shrug. Turn-taking only recently acquired numbers — false-barge-in rate, endpoint latency, semantic end-of-turn accuracy — and it is commoditizing right now, in public, at exactly the pace the rule predicts. Evaluation has no agreed scalar at all: nobody can hand you a number for "was this call good" that a second vendor would reproduce. It is the most expensive add-on on any price list. And compliance is not measured, it is attested — audits and contracts, not benchmarks — which is why it never rides a cost curve at all.
This is my framework rather than a vendor's, so treat it as a hypothesis with a prediction attached. The prediction is uncomfortably specific and appears in §10: the eval layer's price collapses within months of the first credible public benchmark for conversational quality, not gradually before it. If you want to know when your eval moat expires, watch for the leaderboard, not the competitors.
| C | Layer | Vendors inside 2× on price | Switching cost | Verdict |
|---|---|---|---|---|
| 5 | PSTN carriage US & global English | Twilio $0.0085/min inbound · LiveKit $0.01 · resold at $0.015 | config | Buy on price. There is no strategy here. |
| 5 | Media transport WebRTC / SIP | LiveKit OSS ($0) · Daily · Agora · LiveKit Cloud $0.01/min | days | Solved. WebSocket-only is now a defect, not a design choice. |
| 5 | Streaming STT English | $0.0025 · $0.003 · $0.0048 · $0.0065 — five vendors in a 2.6× band | dropdown | Over. At least one platform doesn't even let you choose — it auto-selects on latency. O |
| 5 | "Good-enough" LLM the brain | Gemini 2.5 Flash-Lite $0.10/1M · GPT-5.4-nano $0.20/1M · Sarvam-30B ₹2.5/1M | dropdown | Free in practice. Compete on tool-calling reliability and p95, never on price. |
| 5 | Orchestration cascaded engine | LiveKit Agents · Pipecat · Bolna's MIT-licensed engine — all open source | weeks | Genuinely $0. Anyone charging $0.05/min here is exposed. |
| 4 | Baseline TTS clear & pleasant | Aura-1 $15/1M chars · Bulbul v2 ~$17 · Aura-2 $30 — plus five vendors sold at one identical platform price | dropdown | Commodity at "good enough". Premium is a different product, not a better one. |
| 3 | Turn-taking endpointing, barge-in | Krisp · Deepgram Flux · AssemblyAI native · LiveKit — offered as a menu inside the major platforms O | dropdown to pick, weeks to tune | The frontier. Vendor choice commoditized inside a year; correct tuning is still craft. |
| 3 | Native speech-to-speech realtime audio | ~$0.018/min vs ~$0.128/min — a 7× gap between two suppliers | re-architecture | Price competition has arrived, but two or three players is a market, not a commodity. |
| 2 | Premium TTS expressive / cloned | ElevenLabs at $165/1M chars holds a 5.5–11× premium over the commodity tier | voice = brand | Buyers still pay 11×. Real differentiation survives here. |
| 2 | Eval & simulation the test harness | Best-in-class ships AI-persona multi-turn simulation and batch regression; a direct competitor ships manual review only O | rebuild suites | Priced as the most expensive add-on on the sheet. Commoditizing next. |
| 2 | Production ops the iterate loop | Nobody has closed the loop well — the gap is industry-wide O | high | Underserved across every platform tested. |
| 2 | Indic STT/TTS 8 kHz + code-switch | Sarvam · in-house Indian stacks · AI4Bharat lineage | quality-gated, not price-gated | Price is already commodity at ₹30/hr. Quality at 8 kHz is not. |
| 1 | Compliance & trust SOC 2 · HIPAA · PCI | Sold as flat rent: $1,000–2,000/mo add-ons, or a $30,000/yr floor | re-audit, re-contract | Not a technology. Does not commoditize on a model curve. |
| 1 | PSTN carriage India & regulated markets | DLT registration, 140/160-series numbering, KYC documents O | months + legal entity | Hard moat. Where Indian players structurally beat US platforms. |
| 1 | Domain SOPs system-of-record depth | Native helpdesk/CRM integration is simply missing at three otherwise-strong platforms O | rebuild logic | The least commoditized layer in the entire stack. |
"Commoditized" is usually asserted. It can be measured. If a layer has truly cleared, the credible suppliers will have converged: the ratio between the cheapest and dearest vendor doing the same job collapses toward one. If a layer hasn't cleared, that ratio stays wide, because buyers are still paying for something price alone doesn't capture.
The obstacle is that these layers are quoted in incompatible units — per minute, per hour, per million characters, per million tokens. So everything below is normalized to one dollar figure per minute of live conversation M, using a single consistent model of what a minute contains: four assistant turns, roughly 2,000 input tokens and 60 output tokens per turn, and about 265 characters of synthesized speech (an agent that holds roughly a third of a 150-words-per-minute conversation). Those assumptions are stated so you can break them; §06 shows where they bend.
The chart makes the argument better than the prose can. The three tightest bands — telephony at 1.8×, baseline speech synthesis at 2.3×, English speech recognition at 3.0× — are exactly the layers below the waterline. The three widest — orchestration at 5.5× above a free floor, speech-to-speech at 7.1×, and the full range of speech synthesis at 11× — are exactly the layers where the argument about value is still live.
A major platform prices five different speech-synthesis vendors at one identical $0.015/min, and a sixth — the premium one — at $0.040. When a reseller stops distinguishing between suppliers in its own price list, it is telling you those suppliers are fungible. That is commoditization stated by someone with every incentive to deny it.
Work the same numbers backwards M. That flat $0.015/min sits about 3.7× above the underlying commodity rate of ~$0.004/min. The $0.040/min premium line sits at roughly 0.9× — essentially at cost. The expensive voice is the honest one; the margin is hidden in the cheap tier, where you assume there is none.
That inversion is worth sitting with, because it generalizes. Platform margin does not live where prices look high. It lives where a buyer has stopped comparing — and buyers stop comparing precisely once a layer feels commoditized. The moment you accept that a layer is a commodity is the moment you stop auditing its price.
One minute of inbound support, built three ways, every figure drawn from list prices fetched on the same day. The first column is what the components cost. The second and third are what you can buy the same minute for.
Cross-check the middle bar against what the market actually charges for a whole bundled minute and it holds: an Indian platform sells an all-inclusive connected minute at ₹6 (about $0.10), another lands between $0.07 and $0.15, a third publishes a headline band of $0.07–0.31, and an enterprise contract bottoms out around $0.07 at volume O. The market clears at roughly ten cents against a two-cent component floor.
There is a second reading of the same chart that matters more for anyone selling minutes. Notice that a whole bundled retail minute at $0.10 costs less than two and a half times the premium voice line alone in the third bar. Platform pricing is a packaging decision far more than it is a cost floor — which is exactly why it is vulnerable.
Everything above rests on these numbers, so here they are unconverted, in the units the vendors publish them in. Every row is a list price from the vendor's own pricing page on a single day.
| Vendor / model | Published | Per minute | Note |
|---|---|---|---|
| AssemblyAI Universal-Streaming | $0.15/hr | $0.0025 | Cheapest credible streaming ASR |
| OpenAI gpt-4o-mini-transcribe | — | $0.0030 | |
| Deepgram Nova-3 mono | — | $0.0048 | $0.0042 on the Growth tier |
| Deepgram Nova-3 multilingual | — | $0.0058 | |
| Sarvam STT (Indic) | ₹30/hr | ~$0.0060 | At parity with English ASR |
| Deepgram Flux — turn-aware | — | $0.0065 | +35% over Nova-3 — the premium is for endpointing, not transcription |
| AssemblyAI Universal-3.5 Pro Realtime | $0.45/hr | $0.0075 |
| Vendor / model | Per 1M chars | vs cheapest | ≈ per minute M |
|---|---|---|---|
| Deepgram Aura-1 | $15 | 1.0× | $0.004 |
| Sarvam Bulbul v2 | ₹15/10k ≈ $17 | 1.1× | $0.005 |
| Deepgram Aura-2 | $30 | 2.0× | $0.008 |
| Sarvam Bulbul v3 (beta) | ₹30/10k ≈ $34 | 2.3× | $0.009 |
| ElevenLabs — Creator | $91 | 6.1× | $0.024 |
| ElevenLabs — Pro / Scale / Business | $165 | 11.0× | $0.044 |
| ElevenLabs — Starter | $200 | 13.3× | $0.053 |
| Model | $/1M in | $/1M out | ≈ per minute M |
|---|---|---|---|
| Sarvam-30B | ₹2.5 ≈ $0.029 | ₹10 ≈ $0.115 | ~$0.0003 |
| Gemini 2.5 Flash-Lite | $0.10 | $0.40 | ~$0.0009 |
| GPT-5.4-nano | $0.20 | $1.25 | ~$0.002 |
| Gemini 3.5 Flash-Lite | $0.30 | $2.50 | ~$0.003 |
| GPT-5.4-mini | $0.75 | $4.50 | ~$0.007 |
| Gemini 3.6 Flash | $1.50 | $7.50 | ~$0.014 |
| GPT-5.6-terra / GPT-5.4 | $2.50 | $15.00 | ~$0.024 |
| GPT-5.5 | $5.00 | $30.00 | ~$0.047 |
| Item | Price | What it tells you |
|---|---|---|
| Twilio — US inbound local | $0.0085/min | The true carriage floor |
| Twilio — US outbound local | $0.0140/min | |
| Twilio — US toll-free inbound | $0.0220/min | |
| Twilio — US local number | $1.15/mo | |
| LiveKit Cloud — telephony inbound | $0.01/min | Resold carriage, near cost |
| LiveKit Cloud — agent session | $0.01/min | Managed orchestration, priced at a fifth of the platforms |
| LiveKit media server + Agents framework | $0 — open source | The floor under the entire orchestration layer |
| Retell — voice engine | $0.055/min | Three independent vendors converging on ~$0.05 for a layer with a $0 floor is not a cost floor. It is a price umbrella. |
| Vapi — platform fee | $0.05/min | |
| Bolna — platform fee | ~$0.05–0.06/min |
Two forces keep the commodity layers from collapsing all the way to zero. Both are structural rather than commercial, which means no amount of vendor competition removes them.
My model puts a good-enough brain at roughly $0.0009 per minute. One platform's own price list puts the same class of model at $0.081 per minute — about six times higher P. I am not going to paper over that gap, because both readings are actionable and I cannot yet distinguish them from public data.
Either platforms mark the commodity brain up several-fold, or a real production voice context is three to six times heavier than a naive estimate — long system prompts, retrieved knowledge-base chunks, and a full transcript re-sent on every single turn. My instinct is that the truth is mostly the second: a voice agent re-uploads its entire conversation with each reply, so token consumption grows quadratically with call length while the conversation only grows linearly. Either way the lesson is identical and it is the single most useful operational takeaway in this piece:
Speech-to-speech models promise to delete the pipeline: no separate recognition, no separate synthesis, no seam. The reason they cannot simply undercut a cascaded stack is printed in the price list of the company that makes both.
Stack that tax on top of the fact that a speech-to-speech model re-sends accumulating audio history and the economics resolve cleanly. The cheapest realtime audio model lands at about $0.018/min — genuine parity with a cascaded pipeline. The most expensive lands at about $0.128/min, roughly ten times a cascaded build and a seven-fold gap between two suppliers of the same capability M. One vendor has collapsed speech-to-speech pricing sevenfold below the other; that is price competition arriving, not a commodity forming.
So the prediction here is narrow and, I think, safe: speech-to-speech wins on naturalness well before it wins on cost, and cascaded pipelines stay the production default for anything touching money or compliance — because a cascaded stack hands you an inspectable transcript at every hop, and a native audio model does not. Expect blended architectures, not a switchover.
Here is the most useful thing on any voice platform's pricing page: the add-ons. Each one is a line item the vendor believes it can charge for separately, which makes the price list a direct readout of what the market still treats as scarce. Measured against the two-cent commodity minute, the readout is stark.
Read the top bar again. Evaluating a call is priced at nearly five times the cost of running it, and roughly twice the platform's own voice-engine fee. In almost every other part of software, testing is the cheap part — you run the test suite a thousand times because each run costs nothing. In voice AI the relationship is inverted, and that inversion is the single clearest statement anyone in this industry has made about where the scarcity is.
The rest of the margin is sold as rent rather than usage, which tells you it is access being priced, not compute:
| Add-on | Price | What the pricing model reveals |
|---|---|---|
| HIPAA mode | $2,000/mo | Compliance is sold as access, not usage — the marginal cost of a compliant minute is the same as a non-compliant one |
| Zero Data Retention | $1,000/mo | Mutually exclusive with HIPAA mode on the same account O — a packaging decision, not a technical constraint |
| HIPAA (different vendor, very different scale) | $1,000/mo | The same pattern reappears at a company an order of magnitude smaller |
| Concurrency above 20 channels | $8/mo per slot | Capacity guarantees still command rent |
| "Stable server cluster" | +$0.02/min | Reliability itself is a paid upgrade — priced at roughly one entire commodity minute |
| Enterprise floor at one vendor | $30,000/yr | Compliance, no-code tooling and owned telephony bundled as a single access fee O |
Strip out everything the price sheet has settled, and four things are left standing. None of them is a model. Two of them are not even technology.
India requires DLT registration, 140/160-series numbering and KYC documentation before you can originate a call at all. Onboarding runs into it immediately: one platform demanded Truecaller verification, another a legal business name and registration certificate before it would provision anything O. Meanwhile one major US platform provisions numbers in the US and Canada only, and another's free numbers are US-only O. No model release changes this. It is the structural reason Indian platforms win Indian deployments that US platforms cannot serve at any price.
The nuance the price sheet forces: Indic is no longer expensive, merely uneven. Indic speech recognition at ₹30/hr is at parity with English, and Indic synthesis undercuts the mid-tier English vendors. The moat migrated from price to paperwork — and to quality at 8 kHz with Hinglish code-switching. Anyone still pricing Indic as a premium is charging for a moat that already eroded.
SOC 2, HIPAA, PCI DSS, GDPR and ISO 27001 are audits and contracts, not capabilities. They do not ride a model cost curve because there is no model involved — which is precisely why they are the most durable line on any price list.
The split this produces is sharp and visible. One platform holds the full certification set and charges a $30,000/yr floor for access to it. A technically comparable competitor asserts no certifications it publicly holds, answering the question with "contact support" O. Near-identical technical stacks, opposite market access. Every regulated buyer is filtered by a document, not a benchmark.
This is the widest capability spread anywhere in the space. At the top: AI-persona multi-turn simulation, batch regression through an API, A/B traffic splitting, and production quality assurance that reports resolution rate and flags hallucinations. At the bottom, among otherwise credible competitors: manual transcript review, no batch simulation, no A/B, no regression suite at all O.
When one dimension ranges from best-in-class to nonexistent across platforms that are otherwise near-identical, that dimension is where the real product work is. The $0.10/min quality-assurance line confirms it commercially. By the measurement rule in §02, this moat holds only until someone publishes a benchmark everybody accepts.
The most telling gap in the entire landscape: native helpdesk and CRM integration is missing at three otherwise strong platforms O — all of them pushing customers toward webhooks and Zapier for the one thing an inbound support desk cannot function without.
Meanwhile the platforms that do ship it ship it deeply: help-centre knowledge-base import at one, two-way Salesforce and HubSpot sync at another. Nobody's model quality decides these deals. Their integration surface does. This is the least commoditized layer in the stack and the one where a specialist can still beat a better-funded generalist outright.
Notice what these four have in common: not one of them gets cheaper when inference gets cheaper. Three of them get more valuable, because a world of cheap capable agents is a world with more calls to evaluate, more regulations to satisfy, and more systems of record to write into.
There is a failure mode I keep running into when large product organizations build voice agents internally, and it is the exact inverse of what you would expect. In-house teams are usually accused of reinventing infrastructure. What they actually do is subtler and more expensive: they build the un-commoditized layers rather well, and then dramatically under-buy the commodity ones.
The archetype looks like this. The team has invested real effort in a scenario harness — synthetic test audio, an adversarial test model, generated conversation scenarios — and in an operations cockpit that beats commercial platforms outright on the parts that matter after launch: product insights, repeat-caller detection, per-line-of-business summarization. Those are precisely the layers above the waterline. That is money spent correctly.
And then the commodity layers sit frozen at whatever was wired in during the first sprint:
| Typical in-house state | What the commodity market offers today | The move |
|---|---|---|
| One inference provider, two models | Eight or more providers; the good-enough brain at ~$0.001/min | Add a cheap-fast tier and a fallback. Pure latency and cost win, no lock-in incurred. |
| One fixed premium voice | Commodity synthesis at $15–17/1M chars against $165/1M for the premium tier — roughly 10× | Keep the premium voice where brand actually matters; move high-volume lines of business to the commodity tier. |
| No language selection — often an empty dropdown | Indic speech recognition at ₹30/hr, at price parity with English | Populate it. This is a configuration gap, not a cost problem — and in India the caller base is multilingual by default. |
| No turn-taking or barge-in controls | Now an off-the-shelf dropdown across every major orchestration framework | Buy it; do not engineer it. This layer is commoditizing underneath you as you read this. |
| One fixed telephony provider, outbound only | Regulated-India telephony is a moat you already sit inside | The one place a fixed choice is defensible — but shipping outbound-only forfeits the moat's actual value, which is inbound. |
The trap, stated plainly: a fixed default at a commodity layer feels like a decision that was made. It usually isn't — it is a decision that was never revisited after the layer beneath it collapsed in price. The audit is worth running annually, and it takes an afternoon.
A commoditization map is only worth writing if it makes claims that can later be checked and found wrong. So here is the waterline as a moving object, with dates attached and — more importantly — with the specific observation that would prove each prediction false.
| Horizon | Prediction | Why I believe it | What would falsify it |
|---|---|---|---|
| ≤ 12 mo by mid-2027 |
Semantic endpointing is bundled free into every speech-recognition vendor and stops being a differentiator. | It is already a vendor dropdown inside the major platforms O, and one vendor is currently selling it as a 35%-premium SKU. Premium SKUs for a feature four competitors also ship do not survive a year. | The endpointing premium still exists at mid-2027, or turn-taking quality visibly diverges between vendors on a shared benchmark. |
| 12–24 mo by mid-2028 |
Simulation and batch regression become table stakes, and the $0.10/min quality-assurance price collapses toward $0.01–0.02. | Three platforms shipped multi-turn simulation within about a year of each other, and one already meters evaluation at ₹2/min O. Convergent shipping at that pace is what a layer looks like just before it clears. | AI quality assurance still prices above $0.05/min in mid-2028 — which would mean conversational quality resisted measurement longer than the measurement rule predicts. |
| 12–24 mo | Platform orchestration fees compress from ~$0.05/min toward $0.02–0.03. | A managed vendor already clears the same function at $0.01/min and open source clears it at $0. Three independent vendors landing on exactly $0.05 is a price umbrella, and umbrellas break. | Fees hold at $0.05+ through 2028, which would mean the bundled reliability and support around orchestration is worth more than I think. |
| ongoing | Speech-to-speech wins naturalness before it wins cost; cascaded stays the default for regulated and money-touching flows. | The cheapest realtime audio model already reaches cascaded parity at ~$0.018/min, but the 2–3× audio-token tax is structural and transcript inspectability is weak. | A native audio model ships per-turn inspectable transcripts and undercuts a cascaded stack on price. Both, not either. |
Every serious platform now offers the same recognition, synthesis and language models, and the caller cannot tell which one you picked. Treat a vendor differentiating on "we support 20+ models" as differentiating on a commodity.
Spend the diligence where the spread is real: the eval harness, the production ops loop, compliance posture, and telephony for your specific geography. Those four are where otherwise-similar platforms range from best-in-class to nonexistent.
The ~$0.05/min orchestration fee is the most exposed revenue line in the industry — a five-times markup on a layer with a genuine $0 floor and two credible substitutes. The add-ons are the opposite: defensible, and priced by scarcity you did not manufacture.
Margin is migrating from the minute to the outcome — to evaluation, compliance, reliability guarantees and vertical integration. Price accordingly, before the umbrella breaks and you have to.
Everything below the waterline should be a purchase decision revisited annually, not an architecture decision made once. Everything above it is where your team's time converts into something a vendor cannot sell you: your evaluation suite, your operations loop, your domain SOPs, and your integration into the system of record that actually resolves the customer's problem. The failure mode is not building too much — it is building the right things while quietly overpaying for the wrong ones.
A price sheet dated to a single day is only useful if it is equally explicit about what isn't on it. The following was not verified in this pass and is deliberately absent rather than guessed at. Several of these are load-bearing.
| Gap | Why it matters |
|---|---|
| Indic quality benchmarks | No word-error-rate or opinion-score evidence for Hinglish code-switching at 8 kHz. The claim "Indic price is commodity, Indic quality is not" rests on pricing plus hands-on impressions, not on a measured comparison. This is the highest-value gap on the list — it is the one that decides real deployment choices in India. |
| Open-weight speech models | Open recognition and synthesis models, and their standings on public leaderboards, were not checked. The open-source floor beneath the recognition and baseline-synthesis layers is therefore asserted from price evidence only, not from measured quality. |
| Contact-centre incumbents | The established CCaaS vendors bundling voice agents into existing seats — and outcome-based pricing models — are unassessed. That is commoditization arriving from above, and this analysis only looks upward from the infrastructure. |
| Funding, M&A and consolidation | No coverage. Any claim here about who outlasts whom would be unsupported. |
| Two vendors' per-minute rates | One synthesis vendor publishes only credit bundles, no per-character rate and no latency claim — it appears in this analysis solely as filtered through a reseller's price. Another publishes synthesis credits but no conversational per-minute rate. |
| India regulatory change in 2026 | Current rules on automated and AI-originated calls, and caller-name presentation rollout, were not re-checked. The telephony moat claim rests on hands-on onboarding friction O — good evidence the moat exists, weaker evidence of its present legal detail. |
| Every M figure is mine, not a vendor's | The token and character assumptions are stated in §03 precisely so they can be attacked. The six-fold disagreement documented in §06 is the weakest link in the arithmetic, and I have flagged it rather than smoothed it. |
| Bolna AI → | 200K calls/day, sub-600ms pipeline, and an MIT-licensed orchestration engine — the open-source floor under layer five, shipped by a company that sells the layer above it. |
| Smallest.ai → | The full-stack counter-argument: own the models and the commodity curve works for you. 20× cost reduction to $0.01/min. |
| Ringg.ai → | Distribution-first, ~85% assembled from vendor APIs — the purest test of whether a GTM moat outlasts the commodity curve. |
| Mixpanel for Voice AI → | The thesis this analysis keeps arriving at from the pricing side: the observability and evaluation layer is the scarce one, and nobody has built it properly. |