Strategy

The unit economics of an enterprise voice AI programme

Voice AI unit economics measure the contribution margin on every resolved contact, not just cost per call. Dilr Voice is built so enterprises can model the fully loaded cost to serve, from inference and telephony to human fallback, and see how that margin moves as call volume and resolution rate rise across the programme.

Most voice AI business cases die on the same number: cost per call. A vendor quotes a few pence a minute, someone multiplies it by call volume, and a savings figure lands on a slide. Then the programme goes live, a third of contacts still reach a human, quality assurance eats a headcount, and the finance team asks why the promised margin never showed up in the accounts.

The problem is not the technology. It is the unit. Cost per call tells you what one interaction costs to run. It does not tell you what one resolved customer problem costs to serve, or how that cost moves as volume grows and the agent gets better at finishing contacts on its own. That second view is the contribution margin of a voice AI programme, and it is the number a voice AI agents deployment actually lives or dies on.

This guide is shipped by the team behind Dilr Voice, enterprise voice AI built for regulated deployments. Or see DATS, our five-stage AI consulting system.

This matters because the macro evidence says most enterprises never get to a defensible margin at all. McKinsey's State of AI (November 2025) found that around 88% of organisations now use AI somewhere, yet only about 33% have it in production and roughly 14% report material EBIT impact. BCG's Widening AI Value Gap (September 2025) put the share of "future-built" companies capturing real value at around 5%. The gap between using AI and profiting from it is, in accounting terms, a margin problem nobody modelled.

What are the unit economics of a voice AI programme?

The unit economics of a voice AI programme are the contribution margin earned on each resolved contact: the price or internal value of handling a customer interaction, minus the fully loaded cost to serve it. For Dilr Voice, that cost stack runs from model inference and telephony through human fallback, quality assurance and downstream rework. Unit economics answer one question cost per call cannot: does each resolved contact make money, and does it make more money at scale?

The distinction is not academic. A programme can post a low cost per minute and still run a negative contribution margin once escalations and rework are loaded in. Conversely, a programme with an unremarkable per-minute cost can throw off strong margins if it resolves most contacts without a human. The static unit cost of a call is one input. The margin is the output, and it is the output that pays for the AI operating model you build around it.

Why is cost per call the wrong number to optimise?

Cost per call optimises the wrong denominator. It rewards driving down the price of every interaction, including the ones the agent fails to finish. If you shave a penny off the per-minute rate but resolution stays flat, you have made unresolved contacts cheaper, not better. The number that compounds is cost per resolved contact, because an unresolved contact is paid for twice: once by the agent, then again by the human who cleans it up.

Where enterprise AI value leaks out
88%Use AI71%Gen-AI wkly33%In prod14%EBIT impact6%AI-mature
Share of enterprises reaching each stage of AI value capture, 2025-2026. Source: McKinsey, The State of AI (Nov 2025)

That funnel is the reason unit economics matter. The drop from 88% using AI to 14% seeing EBIT impact is mostly programmes that optimised activity, not margin. If your reporting stops at cost per call, you are measuring the widest bar and ignoring the narrowest one. To fix the denominator, first fix how you count a resolution, which is why margin work always pairs with honest first-contact resolution measurement.

What goes into the fully loaded cost to serve a resolved contact?

The fully loaded cost has four layers: real-time inference, telephony, human fallback and assurance. Inference is the speech-to-text, text-to-speech and language-model calls that run every second of the call. Telephony is the carrier minutes. Human fallback is the cost of every contact that escalates. Assurance is quality sampling, monitoring, retraining and the rework created when the agent gets something wrong. For Dilr Voice, only the first two are truly per-minute; the rest are per-contact or fixed.

The per-minute layers are concrete and published. Carrier minutes from Twilio run about $0.0085 per minute for an inbound US local call and $0.0140 outbound (2026 pay-as-you-go rates). Streaming speech-to-text from Deepgram Nova-3 is about $0.0077 per minute. A conversational voice layer from ElevenLabs is billed from roughly $0.08 per minute on annual business plans. Add a language model and you land, illustratively, near £0.09 to £0.12 per minute of live conversation once converted and blended.

From cost per minute to margin per resolved contact
01Inference + telephonyPer minute, published vendor rates02Cost per handled callRate x average handle time03Cost per resolved contactLoad human fallback for escalations04Loaded cost to serveAdd QA, monitoring, rework05Contribution marginValue of a resolved contact minus cost
The build-up that turns a per-minute rate into a contribution margin. Each step adds a cost layer that cost-per-call reporting ignores.

The trap is that the two per-minute layers are the smallest part of the stack and the only part vendors put on the price sheet. The layers that decide the margin, human fallback and assurance, are the ones the hidden costs of voice AI TCO work exists to surface. A margin model that stops at inference and telephony is modelling roughly a third of the real cost to serve.

How does voice AI gross margin change as the resolution rate rises?

Resolution rate is the single largest lever on voice AI gross margin, because every unresolved contact carries the full human cost on top of the AI cost. As the agent resolves more contacts without escalation, the expensive human layer is spread across fewer contacts, and the cost to serve falls faster than any per-minute saving could deliver. For Dilr Voice programmes, moving resolution from 50% to 80% typically changes the margin more than any inference price cut on offer.

The table below is an illustrative model, not sourced figures. It assumes an AI attempt costs about £0.60 fully loaded per contact (derived from the published rates above plus assurance), and that an escalated contact reaches a human at about £8, the midpoint of published UK estimates of roughly £7 to £12 per inbound call. Cost per resolved contact is the AI attempt plus the escalation cost carried by the contacts that do not resolve.

Resolution rate (modelled)Escalations per 100Illustrative cost per resolved contactMargin vs an £8 human contact
50%50£4.60£3.40
65%35£3.40£4.60
80%20£2.20£5.80
90%10£1.40£6.60

Read across the rows and the shape is the whole thesis: a thirty-point gain in resolution nearly triples the contribution margin, while the underlying per-minute rate never changed. This is why Andreessen Horowitz observed, in its analysis of AI economics, that "we have seen a surprisingly consistent pattern in the financial data of AI companies, with gross margins often in the 50-60% range". Inference is a genuine cost of goods sold that never trends to zero, so margin has to be earned on the resolution side, not conjured from a cheaper model.

How does call volume change the margin?

Volume changes the margin through the fixed layer, not the variable one. Inference, telephony and human fallback scale roughly linearly with contacts, so they do not improve with size. Platform licensing, integration engineering, monitoring and the team that tunes the agent are largely fixed, so their cost per contact falls as volume rises. The blended human-AI cost curve therefore bends downward with scale, but only until a fresh fixed cost, a new language or a new channel, resets it.

This is where a programme quietly earns or loses its margin. The same diagnostic logic underpins our AI placement diagnostic, a fixed-fee assessment used before any deployment commitment, precisely because the fixed layer has to be sized against realistic volume, not a launch-day pilot. Spread a £30,000 annual engineering and platform cost over 50,000 contacts and it adds £0.60 each; spread it over 500,000 and it adds £0.06. Volume does not cut the variable cost, but it decides whether the fixed cost is a rounding error or a margin killer.

Where is the break-even point against human handling?

Break-even is the volume and resolution combination at which the loaded cost to serve a contact by AI falls below the cost of a human. Using the illustrative figures above, an agent resolving 80% of contacts serves each one for about £2.20 against a human at £8 to £12, so it breaks even quickly on variable cost. Whether it clears its fixed cost too depends on volume. For Dilr Voice, break-even is a curve, not a single number.

Two programmes with identical per-minute rates can sit on opposite sides of break-even because one runs at 60% resolution across 20,000 contacts and the other at 85% across 400,000. The first may never clear its fixed layer; the second clears it in the first quarter. This is why unit economics and the total programme ROI case have to be modelled together: the cost curve sets the floor, and the benefit case sits on top of it. Use the DATS five-stage methodology to size both before committing budget.

How is unit economics different from ROI, TCO and pricing?

Unit economics is the cost-side margin curve; ROI, TCO and pricing are adjacent but distinct views. ROI compares total programme benefit to total cost and is a benefit-attribution exercise. TCO enumerates every cost line, including the hidden ones, without expressing a margin. Pricing is what a vendor charges you, covered in per-minute versus per-resolution pricing. Unit economics is the only one that models margin per resolved contact as volume and resolution move.

Keeping these separate stops a common finance-review failure: presenting a TCO table as if it were a margin, or an ROI headline as if it were a unit cost. Each answers a different board question. If your finance team needs to allocate the cost across departments, that is a chargeback and cost-allocation problem, not a unit-economics one, and mixing them produces a model nobody trusts. The clean sequence is TCO to size costs, unit economics to find the margin curve, then the business case to defend the investment.

Want to see this in production? Try Dilr Voice live, book a scoping call, see our DATS methodology, or read about our approach to placing AI where the margin actually moves.

What is the best way to model voice AI unit economics in 2026?

The best model in 2026 is a resolution-driven contribution-margin model, built on published vendor rates and your own escalation and rework data, refreshed quarterly. It should express cost per resolved contact, not per call, and show the margin curve across a realistic resolution range. Dilr Voice exposes the resolution and escalation data this model needs, but the discipline matters more than the platform: a good model on a rival agent beats a missing one.

That verdict is scoped, and there are cases where a lighter approach wins. For a small, single-use-case deployment on Vapi, Retell AI, Bland AI, Synthflow or PolyAI, a simple cost-per-call spreadsheet may be enough to make the decision, and building a full margin model would be over-engineering. The contribution-margin model earns its keep once volume is high, escalation is material, or a regulator or CFO will scrutinise the numbers. Integrations with Twilio, Salesforce and HubSpot then determine how cleanly you can pull the resolution and rework data the model depends on.

Does voice AI have a per-call cost that SaaS never had?

Yes, and this is the structural reason margin discipline matters more for voice AI than for software. Traditional SaaS is built once and served many times at near-zero marginal cost, which is why it posts 80% gross margins. A voice AI contact runs live inference every second, so there is a real cost of goods sold on every interaction. Dilr Voice treats that inference line as a permanent variable cost to be managed, not a one-off to be amortised away.

What gross margin should a voice AI programme target?

There is no universal target, but a defensible programme should model a clear positive contribution margin per resolved contact and know its break-even volume. Rather than chase a headline percentage, anchor on the resolution rate that moves your margin most and the fixed cost that scale must absorb. Dilr Voice programmes are sized against a modelled curve, so the target is the curve's shape, resolution up and cost per resolved contact down, not a number lifted from a SaaS comparison.

How often should you rebuild the unit-economics model?

Rebuild it quarterly, and immediately after any material change: a new language model, a new channel, a pricing change from a telephony or inference vendor, or a shift in escalation rate. Inference prices move fast, and a model built on last year's rates will misstate the margin. For Dilr Voice deployments run through the AI execution office, the unit-economics model is a living artefact reviewed with the operating numbers, not a one-off slide.

Service
AI Execution Office
Service
AI Operating Model
Product
Dilr Voice
Talk to the operators

Model the margin before you sign.

30-min scoping call · No deck · Confidential. We will build the contribution-margin curve with your real resolution and escalation data, and tell you where break-even actually sits.

Written by the Dilr.ai engineering team, practitioners who ship enterprise AI in production. Follow us on LinkedIn for shipping notes, or subscribe via the RSS feed. More strategy work in the strategy category, or start with our enterprise voice AI guide.

voice AI unit economics gross margin enterprisecontribution margin per resolved callcost to serve voice automationvoice AI unit economics redditbest voice AI cost model 2026blended human AI cost curveDilr Voice economics

Questions this article answers

What are the unit economics of a voice AI programme?

The unit economics of a voice AI programme are the contribution margin earned on each resolved contact: the price or internal value of handling a customer interaction, minus the fully loaded cost to serve it. For Dilr Voice, that cost stack runs from model inference and telephony through human fallback, quality assurance and downstream rework. Unit economics answer one question cost per call cannot: does each resolved contact make money, and does it make more money at scale?

Why is cost per call the wrong number to optimise?

Cost per call optimises the wrong denominator. It rewards driving down the price of every interaction, including the ones the agent fails to finish. If you shave a penny off the per-minute rate but resolution stays flat, you have made unresolved contacts cheaper, not better. The number that compounds is cost per resolved contact, because an unresolved contact is paid for twice: once by the agent, then again by the human who cleans it up.

What goes into the fully loaded cost to serve a resolved contact?

The fully loaded cost has four layers: real-time inference, telephony, human fallback and assurance. Inference is the speech-to-text, text-to-speech and language-model calls that run every second of the call. Telephony is the carrier minutes. Human fallback is the cost of every contact that escalates. Assurance is quality sampling, monitoring, retraining and the rework created when the agent gets something wrong. For Dilr Voice, only the first two are truly per-minute; the rest are per-contact or fixed.

How does voice AI gross margin change as the resolution rate rises?

Resolution rate is the single largest lever on voice AI gross margin, because every unresolved contact carries the full human cost on top of the AI cost. As the agent resolves more contacts without escalation, the expensive human layer is spread across fewer contacts, and the cost to serve falls faster than any per-minute saving could deliver. For Dilr Voice programmes, moving resolution from 50% to 80% typically changes the margin more than any inference price cut on offer.

How does call volume change the margin?

Volume changes the margin through the fixed layer, not the variable one. Inference, telephony and human fallback scale roughly linearly with contacts, so they do not improve with size. Platform licensing, integration engineering, monitoring and the team that tunes the agent are largely fixed, so their cost per contact falls as volume rises. The blended human-AI cost curve therefore bends downward with scale, but only until a fresh fixed cost, a new language or a new channel, resets it.

Where is the break-even point against human handling?

Break-even is the volume and resolution combination at which the loaded cost to serve a contact by AI falls below the cost of a human. Using the illustrative figures above, an agent resolving 80% of contacts serves each one for about £2.20 against a human at £8 to £12, so it breaks even quickly on variable cost. Whether it clears its fixed cost too depends on volume. For Dilr Voice, break-even is a curve, not a single number.

How is unit economics different from ROI, TCO and pricing?

Unit economics is the cost-side margin curve; ROI, TCO and pricing are adjacent but distinct views. ROI compares total programme benefit to total cost and is a benefit-attribution exercise. TCO enumerates every cost line, including the hidden ones, without expressing a margin. Pricing is what a vendor charges you, covered in per-minute versus per-resolution pricing. Unit economics is the only one that models margin per resolved contact as volume and resolution move.

What is the best way to model voice AI unit economics in 2026?

The best model in 2026 is a resolution-driven contribution-margin model, built on published vendor rates and your own escalation and rework data, refreshed quarterly. It should express cost per resolved contact, not per call, and show the margin curve across a realistic resolution range. Dilr Voice exposes the resolution and escalation data this model needs, but the discipline matters more than the platform: a good model on a rival agent beats a missing one.

AI consulting (DATS)

Place AI where the P&L moves

The DATS system runs from a fixed-fee placement diagnostic through to embedded delivery, so AI reaches production instead of staying a pilot.

Related articles

← Previous
Voice AI dialler modes: predictive, progressive or preview?

One email, once a month. No hype. Just what we learned shipping.