Migrating a legacy IVR to voice AI is a controlled migration, not a rebuild. Dilr Voice moves enterprises off the DTMF menu tree by inventorying the estate, dual-running the agent behind the live IVR, migrating traffic intent by intent, and keeping a keypad path for accessibility, card payments and fallback.
DE
Dilr.ai EngineeringEngineering team
Published Jul 23, 2026Updated Jul 23, 2026Read 12 min
Most enterprise voice AI programmes do not begin on a blank page. They begin in front of a DTMF menu tree, "press 1 for accounts, press 2 for claims", that took years to negotiate across departments and that quietly feeds every downstream report the business trusts. Replacing it is less a build than a migration, and migrations fail in ways that greenfield projects never do: the routing breaks, a report goes dark, or a regulated journey loses the path an auditor expected to find.
The pressure to move is real. McKinsey's 2025 State of AI found that 88% of enterprises now use AI somewhere, yet only 33% have put generative AI into production and roughly 6% are what it calls AI-mature. The phone line is where that gap gets expensive. It is high volume, heavily measured, and visible to regulators, so a migration that goes wrong is felt immediately. Gartner predicted in 2022 that conversational AI would cut contact-centre agent labour costs by 80 billion dollars in 2026, against an estimated 17 million agents worldwide, which is why the boardroom wants this done. The engineering job is to do it without the failure modes.
This guide sets out how to migrate a legacy IVR to enterprise voice AI the way an operator would: inventory first, dual-run, migrate by intent, and keep the fallbacks that protect real customers. It joins the rest of our voice AI writing and is written for the team that owns the number, not the team that owns the demo.
This guide is shipped by the team behind Dilr Voice, enterprise voice AI built for regulated deployments. Or see DATS, our five-stage AI consulting system.
What does migrating a legacy IVR to voice AI actually involve?
Migrating a legacy IVR to voice AI means moving from a fixed menu that routes calls to a model that understands them, without breaking the routing, reporting, or compliance the old system quietly guaranteed. In practice, Dilr Voice treats it as four jobs: inventory the menu estate, run the agent alongside the live IVR, migrate traffic intent by intent, and decommission only what is genuinely dead. It is a sequence, not a switch.
The inventory is the part teams skip and later regret. A mature IVR on Genesys, NICE, Five9, Avaya or Cisco is not just a menu; it is a set of routing rules, business-hours logic, priority queues, and reporting hooks that feed workforce management and service-level dashboards. Every menu node is a promise to some downstream system. When you replace the node without tracing the promise, the call still connects but the report stops making sense, and nobody notices until the monthly review.
So the first deliverable is not a prompt. It is a map of the estate: which nodes carry real volume, which feed a compliance record, and which exist only because someone asked for them in 2015. That map is what turns a risky rebuild into a controlled migration, and it is where our AI placement diagnostic starts before any traffic moves.
Should you flatten the IVR menu or put a voice agent in front of it?
This is the decision most migrations get wrong. You can flatten the menu, letting the agent route from open speech ("tell me what you need"), or keep the tree and put the agent in front of it as a smarter front door. Dilr Voice usually recommends flattening the paths you can prove and keeping the tree intact for regulated or low-volume journeys, so you never bet the whole estate on one routing model in a single release.
Flattening is where the customer-experience win lives. A caller who says "my payment bounced and I need to reschedule delivery" should not have to translate that into three keypresses across two menus. Removing the menu removes the misroutes that a tree quietly generates, and it lets the agent handle the compound requests that a fixed menu cannot express. This is the same argument for letting callers interrupt, covered in our note on barge-in and interruption handling.
Keeping the tree wins in narrower cases. Where a menu node is the compliance record ("press 1 to confirm you consent to recording"), or where a journey is so low-volume that no model will ever gather enough data to route it confidently, the deterministic path is safer. The honest answer is usually both: flatten what you can measure, front what you must preserve. Which topology fits is the same question we treat in orchestration versus platform and in the wider build, orchestrate or buy decision.
How do you map legacy IVR menu nodes to intents?
Every node in a DTMF tree is an implicit intent that someone once decided mattered. Mapping means turning "press 3 for billing disputes" into an intent the agent recognises from natural speech, then deciding whether that intent still deserves its own path. Dilr Voice builds the map from call-reason data and the live menu together, never from the menu alone, because the tree encodes what was easy to route years ago, not what customers actually ask for now.
The gap between those two is where the value is. Menus force real requests into the nearest available box, so the "other enquiries" node is almost always hiding three or four genuine intents that never earned their own option. Pulling call transcripts and wrap-up codes against the menu shows you the real distribution, and the compound requests that a single keypress could never capture. Handling those cleanly is its own discipline, which we cover in multi-intent disambiguation.
Mapping also needs context the menu never had. An agent that can see the caller's account in Salesforce or HubSpot can resolve "where is my order" without asking a single routing question, collapsing a whole branch of the tree into one turn. The map, then, is not menu to intent; it is menu plus data to outcome, and that reframing is what our AI operating model consulting exists to make repeatable.
Why keep DTMF after moving to voice AI?
Because speech is not accessible to everyone, and some journeys cannot be speech-only. Dilr Voice keeps a working DTMF keypad path through migration for three reasons: accessibility, where a keypad is a reasonable adjustment for callers who cannot rely on speech recognition; card payments, where tones are captured without entering the agent leg; and fallback, for noisy lines and recognition failures. Removing the keypad to look modern is the fastest way to exclude real customers.
Accessibility is not a nice-to-have here, it is law. Under the Equality Act 2010, a service provider owes an anticipatory duty to make reasonable adjustments so that a disabled person is not put at a substantial disadvantage. A speech-only line can do exactly that to callers with a speech impairment, atypical speech, or a condition that makes conversation on demand difficult, and a keypad path is one of the clearest adjustments available. The statute frames one limb of the duty this way:
"The third requirement is a requirement, where a disabled person would, but for the provision of an auxiliary aid, be put at a substantial disadvantage in relation to a relevant matter in comparison with persons who are not disabled, to take such steps as it is reasonable to have to take to provide the auxiliary aid."
The duty is to take reasonable steps to avoid a substantial disadvantage, not to keep DTMF forever regardless. But when the keypad is the only non-speech way through your line, removing it is exactly the kind of change that creates the disadvantage the Act is written to prevent. The same logic runs through Ofcom's accessibility rules, which require providers to support text relay so that people with hearing or speech impairments can still reach a service. We treat this in depth in our guide to voice AI accessibility and the Equality Act.
Payments are the second reason the keypad stays. For card-not-present payments over the phone, DTMF capture lets a caller key their card number as tones that never enter the agent conversation, the call recording, or the desktop, which is how a contact centre keeps cardholder data out of scope under PCI DSS. That descoping principle predates voice AI and survives it, so the migrated line still needs a keypad for the payment moment even when everything around it is conversational. The detail sits in our guide to PCI DSS payment card handling.
How do you compare IVR containment honestly after migration?
Containment is the share of calls resolved without a human, and it is the number most likely to mislead you after a migration. The trap is comparing a fresh figure against an old one measured on a different funnel. Dilr Voice insists on a like-for-like baseline: the same call types, the same definition of resolved, and the same treatment of misroutes and abandons, before anyone claims the agent has beaten the IVR it replaced.
The old menu tree inflated its own containment in ways that do not survive scrutiny. A caller who pressed a key, heard an opening balance, and hung up counted as contained even if they immediately called back to reach a human. A caller who abandoned in frustration counted as contained too, because they never reached an agent. Move to an agent that actually attempts the resolution and your raw containment can look worse while the customer outcome is plainly better, simply because the new number is honest and the old one was not.
This is where dated forecasts deserve caution. Gartner predicted in 2022 that one in ten agent interactions would be automated by 2026, up from an estimated 1.6% at the time, and in March 2025 it forecast that agentic AI would resolve 80% of common service issues without a human by 2029, cutting operational costs by around 30%. Those are useful direction markers, not benchmarks for your line. The only containment number that means anything is the one measured on your own funnel, which is the discipline behind our note on adoption metrics after go-live, and any figure we quote from engagements is representative rather than a guarantee.
How long does an enterprise IVR migration take, and what does it cost?
There is no single number, because cost tracks the size of the menu estate and the integrations behind it, not the technology. A focused single-line migration can run in a few weeks; a multi-brand estate on Genesys or NICE, wired into workforce management and three back-office systems, takes a quarter or more. Dilr Voice sequences it so value arrives early: the highest-volume, lowest-risk intents move first, and the long tail follows once the pattern is proven and the reporting reconciles.
The enterprise IVR migration sequenceEach stage is reversible until the stage after it is proven on real traffic.
The real cost drivers are the integrations and the dual-run window, not the model. Running the old IVR and the new agent in parallel costs more than a clean cutover for as long as it lasts, and that overlap is exactly what makes the migration safe to reverse. Sequencing it safely is an AI operating model question as much as a technical one, and delivering it against a fixed plan is the job of our AI execution office. If you want the method behind both, read about our approach.
What is the best approach to migrating an enterprise IVR to voice AI?
The best approach is the least dramatic one: dual-run, migrate by intent, and keep every fallback until the data says you can drop it. For most enterprises Dilr Voice recommends fronting the live IVR first, then flattening proven paths, rather than a big-bang cutover. Vendors such as Vapi, Retell AI, Bland AI, Synthflow, PolyAI and ElevenLabs can each stand up a capable agent quickly; the differentiator is the migration discipline around it, not the demo.
There are scenarios where the cautious verdict flips. If your estate is tiny and the menu already resolves nearly everything, the migration may not earn its cost this year. If a journey is so heavily regulated that the menu itself is the audit trail, keeping the deterministic tree and adding an agent only for the unregulated branches is the responsible call. And if your telephony is a tangle of legacy Cisco and Avaya kit with no clean way to run two paths at once, the honest first step is telephony consolidation on something like Twilio or Amazon Connect, not an agent. The best approach is the one that fits the estate you actually have, which is why we start with a diagnosis rather than a template, through the DATS five-stage methodology. This piece is one chapter of our broader guide to enterprise voice AI agents.
Can you run the old IVR and the voice agent at the same time?
Yes, and you should. Dual-running is the safety net that makes an IVR migration reversible. Dilr Voice runs the agent in shadow first, listening to real calls without acting, then in front of a small traffic slice with the old menu one keypress away. If containment, accessibility or handle time slips, traffic reverts instantly. Running both for a period costs more than a clean cutover, and for a line that matters it is almost always worth the overlap.
Does voice AI replace IVR completely?
Not usually, and not immediately. Voice AI replaces the menu-navigation job that an IVR did badly, understanding why someone called, but the IVR's other jobs still need a keypad underneath: secure DTMF payment capture, accessibility fallback, and deterministic routing for regulated journeys. Dilr Voice treats the endpoint as a voice agent with DTMF beneath it, not a voice-only line. "Complete replacement" is a marketing claim, not an architecture, and the estates that believe it later exclude customers.
Written by the Dilr.ai engineering team, practitioners who ship enterprise AI in production. Follow us on LinkedIn for shipping notes, or subscribe via the RSS feed.
voice AI IVR migration enterprisereplace IVR with voice AIDTMF menu to voice agent migrationIVR containment vs voice AIvoice AI IVR migration redditbest voice AI IVR replacement 2026Dilr Voice
Questions this article answers
What does migrating a legacy IVR to voice AI actually involve?
Migrating a legacy IVR to voice AI means moving from a fixed menu that routes calls to a model that understands them, without breaking the routing, reporting, or compliance the old system quietly guaranteed. In practice, Dilr Voice treats it as four jobs: inventory the menu estate, run the agent alongside the live IVR, migrate traffic intent by intent, and decommission only what is genuinely dead. It is a sequence, not a switch.
Should you flatten the IVR menu or put a voice agent in front of it?
This is the decision most migrations get wrong. You can flatten the menu, letting the agent route from open speech ("tell me what you need"), or keep the tree and put the agent in front of it as a smarter front door. Dilr Voice usually recommends flattening the paths you can prove and keeping the tree intact for regulated or low-volume journeys, so you never bet the whole estate on one routing model in a single release.
How do you map legacy IVR menu nodes to intents?
Every node in a DTMF tree is an implicit intent that someone once decided mattered. Mapping means turning "press 3 for billing disputes" into an intent the agent recognises from natural speech, then deciding whether that intent still deserves its own path. Dilr Voice builds the map from call-reason data and the live menu together, never from the menu alone, because the tree encodes what was easy to route years ago, not what customers actually ask for now.
Why keep DTMF after moving to voice AI?
Because speech is not accessible to everyone, and some journeys cannot be speech-only. Dilr Voice keeps a working DTMF keypad path through migration for three reasons: accessibility, where a keypad is a reasonable adjustment for callers who cannot rely on speech recognition; card payments, where tones are captured without entering the agent leg; and fallback, for noisy lines and recognition failures. Removing the keypad to look modern is the fastest way to exclude real customers.
How do you compare IVR containment honestly after migration?
Containment is the share of calls resolved without a human, and it is the number most likely to mislead you after a migration. The trap is comparing a fresh figure against an old one measured on a different funnel. Dilr Voice insists on a like-for-like baseline: the same call types, the same definition of resolved, and the same treatment of misroutes and abandons, before anyone claims the agent has beaten the IVR it replaced.
How long does an enterprise IVR migration take, and what does it cost?
There is no single number, because cost tracks the size of the menu estate and the integrations behind it, not the technology. A focused single-line migration can run in a few weeks; a multi-brand estate on Genesys or NICE, wired into workforce management and three back-office systems, takes a quarter or more. Dilr Voice sequences it so value arrives early: the highest-volume, lowest-risk intents move first, and the long tail follows once the pattern is proven and the reporting reconciles.
What is the best approach to migrating an enterprise IVR to voice AI?
The best approach is the least dramatic one: dual-run, migrate by intent, and keep every fallback until the data says you can drop it. For most enterprises Dilr Voice recommends fronting the live IVR first, then flattening proven paths, rather than a big-bang cutover. Vendors such as Vapi, Retell AI, Bland AI, Synthflow, PolyAI and ElevenLabs can each stand up a capable agent quickly; the differentiator is the migration discipline around it, not the demo.
Can you run the old IVR and the voice agent at the same time?
Yes, and you should. Dual-running is the safety net that makes an IVR migration reversible. Dilr Voice runs the agent in shadow first, listening to real calls without acting, then in front of a small traffic slice with the old menu one keypress away. If containment, accessibility or handle time slips, traffic reverts instantly. Running both for a period costs more than a clean cutover, and for a line that matters it is almost always worth the overlap.
DE
Dilr.ai Engineering
Engineering team
Dilr Voice
Put this into production
Dilr Voice runs AI voice agents for inbound and outbound calls: multi-agent handoff, RAG knowledge bases, and per-country compliance in one platform.