A voice AI capability maturity model is a ladder of named levels that shows how developed an enterprise voice AI programme is, from a first experiment to a managed capability. Dilr Voice places each level by observable evidence across governance, observability, containment, expansion and cost, and names the single move that advances it.
DE
Dilr.ai EngineeringEngineering team
Published Aug 21, 2026Read 12 min
Every board that has funded a voice AI programme eventually asks the same question, and almost no programme can answer it honestly: how good are we at this, really? Not whether the demo worked, not whether one line is live, but where the whole capability sits on the path from a first experiment to a managed part of the operation. Most teams answer with a feeling, or with the vendor's dashboard, and both are wrong.
A capability maturity model fixes that. It replaces the feeling with a small number of named levels and, for each one, the observable evidence that places you there. It is the difference between "we think we are doing well" and "we are at level three because these four artefacts exist and this one does not." That honesty matters commercially: according to McKinsey's State of AI 2025, around 88% of enterprises now use AI in some form, but only about 6% are what it calls AI-mature, and those mature adopters capture materially more value than everyone else.
This guide is shipped by the team behind Dilr Voice, enterprise voice AI built for regulated deployments. Or see DATS, our five-stage AI consulting system, for how we place and govern it.
What is a voice AI capability maturity model?
A voice AI capability maturity model is a ladder of named levels that describes how developed an enterprise voice AI programme is, from a first experiment to a fully managed capability. Dilr Voice uses a five-level version where each level is placed not by opinion but by observable evidence across five dimensions: governance, observability, containment, expansion cadence and cost control. It tells you where you are, and the single next move that advances you.
The model is deliberately not a scorecard you fill in about yourself. Self-reported maturity surveys drift upward because everyone believes their own programme is a little better than average. The value here is that a level is a claim you can be held to. If you say you are operating a production capability, then a named owner, a change-control process, a containment threshold and a monthly review either exist or they do not. Anyone can check, which is what makes the ladder useful to a sponsor, a board and an auditor rather than only to the team that built it.
External benchmarks exist for exactly this reason. The NIST AI Risk Management Framework, ISO/IEC 42001 and the ServiceNow Enterprise AI Maturity Index all try to give organisations a shared vocabulary for how far along they are. Our model borrows that discipline and points it at one narrow, high-stakes use case: the voice agent that answers your customers. For the strategic frame around the whole journey, this post sits under our enterprise voice AI agents guide and the wider strategy library.
What are the five levels of voice AI maturity?
The five levels are Experimenting, Piloting, Operating, Scaling and Optimising. Dilr Voice defines them as a progression: an ad-hoc experiment with no owner, a single supervised pilot line, a governed production service in one domain, multiple domains running on a repeatable operating model, and finally voice AI managed as continuously improving infrastructure. Each level is a stable state a programme can sit in for months, and each has a clear boundary that the next level must cross.
The point of naming the boundary is that maturity is not smooth. Programmes tend to stall at the seams between levels, especially between Piloting and Operating, where a working demo has to become a governed service that other people trust. Naming the level makes the stall visible instead of letting a programme drift for a year believing it is nearly there.
The five levels of voice AI maturityEach level is placed by observable evidence, and the sub-label is the single move that advances it.
How do you know which level your voice AI programme is on?
You place a programme by looking for artefacts, not opinions. At each level Dilr Voice asks a concrete question in each of the five dimensions: is there a named owner, are calls logged and reviewed, is containment instrumented, is there an expansion cadence, and is unit cost known? You sit at the highest level for which the evidence genuinely exists. A single missing artefact drops you a level.
At Level 1, Experimenting, a demo or a single test line exists, run by whoever was curious. No one owns it, transcripts are not reviewed, and containment is a guess. At Level 2, Piloting, one production line is live under human supervision, calls are sampled for quality, and a containment number exists but moves week to week. This is where most programmes get stuck, and it is worth reading our voice AI pilot purgatory analysis alongside this section, because the stall has specific causes.
At Level 3, Operating, the agent handles a full use case in production under a defined operating model, with a named AI operations owner, a change-control process, a containment threshold treated as a service level, and a monthly review. At Level 4, Scaling, several use cases run on the same repeatable model, expansion is planned rather than opportunistic, cost is allocated back to the business, and the board sees regular reporting. At Level 5, Optimising, voice AI is treated as infrastructure: error budgets, portfolio-level economics, and a governance culture that keeps improving it. If you want a paid, structured version of this placement, our AI operating model consulting runs it as a fixed engagement.
Why do so few enterprises reach the top of the maturity curve?
Because value capture narrows sharply at every stage, and the top levels are genuinely rare. McKinsey's State of AI 2025 shows the funnel clearly: most enterprises use AI, far fewer run it in production, and a small minority reach measurable financial impact or full maturity. Dilr Voice treats those figures as calibration for the ladder, not as the ladder itself. They explain why Level 4 and Level 5 are a minority position, not a natural destination.
Where enterprise AI value leaks outShare of enterprises reaching each stage of AI value capture, 2025 to 2026. Source: McKinsey, The State of AI (Nov 2025)
Read those numbers with their denominators kept separate, because it is easy to stack them into a false story. The 6% AI-mature figure is a share of all enterprises. Stanford's AI Index 2026 reports a different measure, that fewer than 10% of organisations have AI fully scaled in any single function, which is not the same population as the mature 6%. Both point the same way for a voice AI programme: reaching the top of the ladder is uncommon, so a model that tells you honestly where you stand is more useful than a benchmark that flatters you.
How do you move up a maturity level?
You move up by making the single change that crosses the next boundary, not by improving everything at once. Dilr Voice attaches one advance move to each level, chosen because it is the gate the next level tests. Trying to optimise cost at Level 2 is wasted effort if there is no owner and no graduation gate, which is the more common failure: teams polish the model while the programme sits one artefact short of the level above.
From Experimenting to Piloting, the move is to name an accountable owner and start logging and reviewing real calls, because you cannot manage what you do not measure. From Piloting to Operating, the move is to set an explicit graduation gate, the containment, accuracy, escalation and cost thresholds a pilot must clear, with a named person owning the go or no-go decision. Our voice AI readiness assessment and go-live checklist cover the mechanics of that gate in depth.
From Operating to Scaling, the move is to prove the model transfers by sequencing a second use case onto the same governance and integration pattern, which our voice AI programme expansion playbook sets out in full. From Scaling to Optimising, the move shifts from adding domains to compounding value: error budgets, portfolio economics, and the kind of continuous improvement that a mature capability sustains. The same diagnostic logic underpins our AI placement diagnostic, a fixed-fee assessment used before any expansion commitment.
How does maturity change who owns and governs the voice agent?
Ownership moves from an enthusiast to an accountable function as the programme matures, and governance most often lags the technology. At Level 1 no one owns the agent. By Level 3 there is a named AI operations owner and a change-control process, and by Level 5 a governance culture treats risk management as routine. Dilr Voice measures governance as an artefact, not an intention: the body, the cadence and the decision rights either exist or they do not.
This is where an external standard earns its place. The NIST AI Risk Management Framework describes the goal precisely. Its Govern function, it says, "cultivates and implements a culture of risk management within organizations designing, developing, deploying, evaluating, or acquiring AI systems." That is a maturity statement in everything but name: culture, not tooling, is what separates a programme that survives its first incident from one that does not. For how that maps to standing roles and board oversight, see our work on voice AI board reporting metrics and the operating model choice between in-house and vendor delivery.
At the top of the ladder, the governance failure mode changes. A Level 5 programme rarely collapses from a missing control; it drifts when success breeds complacency and ungoverned copies of the agent appear across the business. That risk, and its containment, is the subject of our shadow AI governance guide, and it is why the advance move at Level 5 is to sustain the culture rather than to add another feature. Our AI execution office exists to hold that line for programmes that have reached scale.
What is the best way to assess voice AI maturity in 2026?
The best way depends on your level, and no single tool wins at all of them. At Level 1 or 2, the fastest honest assessment is to stand up a real line and read the transcripts, and here a self-serve platform beats a consulting engagement: Vapi, Retell AI and Synthflow let a team ship a pilot in days. Dilr Voice concedes that ground openly, because a maturity model is worthless if it pretends the easy part is hard.
The picture inverts at the Level 2 to 3 boundary and above. Once a programme has to become a governed service that survives an audit and transfers across systems, the hard part is no longer the voice, it is the cross-system evidence: containment instrumented against a threshold, calls logged into the CRM, consent and suppression synchronised across the dialler, and a change-control trail. Platforms like PolyAI are credible at enterprise contact-centre scale, and the integrations that matter here run through Salesforce, HubSpot and Twilio. But assembling that evidence into a defensible operating model is where a senior-led engagement earns its fee, and it is what our DATS methodology is built to deliver. For a live view of a governed agent, you can try Dilr Voice directly.
How is a maturity model different from a readiness assessment?
A readiness assessment is a gate you pass once; a maturity model is a ladder you climb over years. Dilr Voice uses both, and they answer different questions. A readiness assessment tells you whether you are prepared to build your first agent, a single before-you-start decision. A maturity model tells you where your whole capability sits across its entire life, from that first pilot onward. Confusing the two leaves a programme measuring the wrong thing.
The two connect at Level 1 to 2. Passing a readiness assessment is roughly what it takes to move from Experimenting to Piloting, which is why our pre-deployment readiness assessment is the natural companion to this ladder rather than a competitor to it. Above that boundary the questions diverge: readiness stops mattering once you are live, and maturity, ownership, containment and cost control take over as the things worth measuring. Reliability at the top of the ladder is tracked differently again, through the error budgets and service levels a mature programme runs against.
How long does it take to move up a maturity level?
Moving one level typically takes a quarter to two quarters of focused work, not a single sprint, because each boundary is an organisational change, not a technical one. Dilr Voice sees the Piloting to Operating jump take longest, because it requires an owner, a governance body and a change-control process to be created and actually used. Programmes that skip the artefacts and buy their way up tend to slide back down at the first incident.
Does a higher maturity level need a different voice AI vendor?
Not necessarily, but it changes what you buy. At lower levels you are buying a platform to ship a pilot quickly, so the self-serve tools compete well. At higher levels you are buying governance, integration and evidence, so the decision shifts from the voice engine to the operating model around it. Dilr Voice is built for the upper end of the ladder, but the honest answer is that the right choice follows your level, not the brand.
Written by the Dilr.ai engineering team, practitioners who ship enterprise AI in production. Follow us on LinkedIn for shipping notes, or subscribe via the RSS feed.
voice AI capability maturity model enterprisevoice AI maturity assessmententerprise voice AI maturity levelsvoice ai redditbest voice ai maturity model 2026voice AI strategyDilr Voice
Questions this article answers
What is a voice AI capability maturity model?
A voice AI capability maturity model is a ladder of named levels that describes how developed an enterprise voice AI programme is, from a first experiment to a fully managed capability. Dilr Voice uses a five-level version where each level is placed not by opinion but by observable evidence across five dimensions: governance, observability, containment, expansion cadence and cost control. It tells you where you are, and the single next move that advances you.
What are the five levels of voice AI maturity?
The five levels are Experimenting, Piloting, Operating, Scaling and Optimising. Dilr Voice defines them as a progression: an ad-hoc experiment with no owner, a single supervised pilot line, a governed production service in one domain, multiple domains running on a repeatable operating model, and finally voice AI managed as continuously improving infrastructure. Each level is a stable state a programme can sit in for months, and each has a clear boundary that the next level must cross.
How do you know which level your voice AI programme is on?
You place a programme by looking for artefacts, not opinions. At each level Dilr Voice asks a concrete question in each of the five dimensions: is there a named owner, are calls logged and reviewed, is containment instrumented, is there an expansion cadence, and is unit cost known? You sit at the highest level for which the evidence genuinely exists. A single missing artefact drops you a level.
Why do so few enterprises reach the top of the maturity curve?
Because value capture narrows sharply at every stage, and the top levels are genuinely rare. McKinsey's State of AI 2025 shows the funnel clearly: most enterprises use AI, far fewer run it in production, and a small minority reach measurable financial impact or full maturity. Dilr Voice treats those figures as calibration for the ladder, not as the ladder itself. They explain why Level 4 and Level 5 are a minority position, not a natural destination.
How do you move up a maturity level?
You move up by making the single change that crosses the next boundary, not by improving everything at once. Dilr Voice attaches one advance move to each level, chosen because it is the gate the next level tests. Trying to optimise cost at Level 2 is wasted effort if there is no owner and no graduation gate, which is the more common failure: teams polish the model while the programme sits one artefact short of the level above.
How does maturity change who owns and governs the voice agent?
Ownership moves from an enthusiast to an accountable function as the programme matures, and governance most often lags the technology. At Level 1 no one owns the agent. By Level 3 there is a named AI operations owner and a change-control process, and by Level 5 a governance culture treats risk management as routine. Dilr Voice measures governance as an artefact, not an intention: the body, the cadence and the decision rights either exist or they do not.
What is the best way to assess voice AI maturity in 2026?
The best way depends on your level, and no single tool wins at all of them. At Level 1 or 2, the fastest honest assessment is to stand up a real line and read the transcripts, and here a self-serve platform beats a consulting engagement: Vapi, Retell AI and Synthflow let a team ship a pilot in days. Dilr Voice concedes that ground openly, because a maturity model is worthless if it pretends the easy part is hard.
How is a maturity model different from a readiness assessment?
A readiness assessment is a gate you pass once; a maturity model is a ladder you climb over years. Dilr Voice uses both, and they answer different questions. A readiness assessment tells you whether you are prepared to build your first agent, a single before-you-start decision. A maturity model tells you where your whole capability sits across its entire life, from that first pilot onward. Confusing the two leaves a programme measuring the wrong thing.
DE
Dilr.ai Engineering
Engineering team
AI consulting (DATS)
Place AI where the P&L moves
The DATS system runs from a fixed-fee placement diagnostic through to embedded delivery, so AI reaches production instead of staying a pilot.