The voice AI programme post-mortem: a lessons-learned guide
In short
A voice AI programme post-mortem is a structured, blameless review that captures what a deployment taught you before the team rotates. Dilr.ai's guide covers when to run it, how to keep it blameless, the estimate-versus-actual questions that matter, and how the lessons feed the next funding cycle rather than dying in a document.
DE
Dilr.ai EngineeringEngineering team
Published Aug 14, 2026Read 11 min
Most enterprise voice AI programmes generate their most valuable asset in the last two weeks and then throw it away. That asset is not the agent, the dashboards or the containment rate. It is the honest account of what the team assumed, what broke, which decisions aged badly, and how far the real numbers drifted from the business case. In the McKinsey State of AI 2025 survey, 88% of enterprises now use AI somewhere, yet only about 6% capture material EBIT impact. The gap between those two figures is mostly organisational learning that never got written down.
The stakes are not academic. Gartner predicts that at least 30% of generative AI projects will be abandoned after proof of concept by the end of 2025, citing poor data quality, weak risk controls, escalating costs and unclear business value. RAND, after interviewing 65 data scientists and engineers with at least five years of experience, found that more than 80% of AI projects fail, twice the rate of non-AI IT projects. As Gartner's Rita Sallam put it, "executives are impatient to see returns on GenAI investments, yet organizations are struggling to prove and realize value." A programme post-mortem is the discipline that turns one team's expensive mistakes into the next team's starting position, and it is the practice most voice AI programmes skip.
This guide is shipped by the team behind Dilr Voice, enterprise voice AI built for regulated deployments. Or see DATS, our five-stage AI consulting system, which builds the review into every engagement.
What is a voice AI programme post-mortem?
A voice AI programme post-mortem is a structured, blameless review of a whole deployment, run at a natural boundary such as the end of a phase or the year-one mark, to capture what the programme actually taught you. Dilr.ai treats it as a governance step, not a retrospective ritual: it records the assumptions that broke, the decisions that aged badly, and the gap between forecast and actual, so the next funding cycle starts from evidence rather than memory.
It is deliberately different from a technical incident review. When a single call flow fails at 2am, you run the blameless post-mortem inside your voice AI incident response runbook and fix the system that let it happen. A programme post-mortem zooms out to the entire initiative: the sourcing decision, the integration roadmap, the change-management plan, the vendor relationship and the economics. Both are blameless and both feed prevention, but one asks "why did this outage happen" and the other asks "what did this programme teach us about how we deploy voice AI at all". Confusing the two is why so many teams run neither well.
Why do so many voice AI programmes lose their hardest lessons?
Programmes lose their lessons because the people who hold them leave the room. By the time a voice AI deployment reaches steady state, the pilot squad has rotated, the launch engineer has moved on, and the sponsor has claimed the win. Nobody owns the memory. The result is a value gap: McKinsey's 2025 data shows most enterprises now use AI, but only a small fraction reach production and fewer still capture EBIT impact, and the difference is rarely the model.
Where enterprise AI value leaks outShare of enterprises reaching each stage of AI value capture: most use AI, few reach production, fewer still see EBIT impact. Source: McKinsey, The State of AI (2025)
The pattern repeats because the incentives point the wrong way. A launch is celebrated; a candid review of what the launch cost and what it missed feels like inviting criticism. So the programme documents its outputs, its go-live date and its headline containment number, and quietly buries the estimate-versus-actual truth. That is how an organisation ends up making the same integration mistake on its second voice agent that it made on its first, and why a disciplined DATS placement review so often finds the same unlearned lesson sitting in two different business units. The shadow AI governance problem, a recurring theme across our strategy guides, has the same root: knowledge that was never captured centrally.
When should you run a voice AI programme post-mortem?
Run a programme post-mortem at four moments: the end of each delivery phase, immediately after a major incident once the fire is out, at the year-one review before the renewal decision, and before any expansion into a second use case. Each trigger asks a different question, so Dilr.ai recommends scheduling the phase and year-one reviews in advance rather than waiting for a crisis to force one. A review you have to argue for is a review that will be skipped.
The year-one trigger is the one most often conflated with something else. A post-mortem is not the same as a voice AI business case refresh: the business case refresh re-forecasts the ROI and decides renewal on the numbers, while the post-mortem explains why the numbers landed where they did and what that means for how you run the next programme. Similarly, the pre-expansion trigger feeds your programme expansion playbook rather than replacing it. The post-mortem supplies the honest inputs; the other documents make the decisions. Run them in that order and each one gets sharper.
How do you run a blameless post-mortem that gets honest input?
You get honest input by making it structurally safe to be honest. A blameless review separates the account of what happened from any judgement of who is at fault, so contributors describe decisions as they looked at the time rather than defending them in hindsight. Google's Site Reliability Engineering team, which popularised the practice for production failures, states the principle directly, and the same rule is what makes a programme review usable rather than performative.
"For a postmortem to be truly blameless, it must focus on identifying the contributing causes of the incident without indicting any individual or team for bad or inappropriate behavior."
Google wrote that about a single outage, and it governs the post-incident review inside your incident runbook. Lift it to the programme level and it means the same thing: a blameless review starts from the assumption that everyone involved had good intentions and acted on the information they had. You cannot fix people, but you can fix the sourcing process, the integration sequence and the change plan that set them up to make a wrong call. The practical mechanics are unglamorous: a written pre-read circulated before the meeting so the loudest voice does not set the narrative, a neutral facilitator who did not own delivery, and a rule that every "decision that aged badly" is recorded against the information available on the day, not the information you have now. This is the same discipline our AI operating model consulting installs as a standing cadence rather than a one-off.
What should a voice AI post-mortem actually capture?
A voice AI post-mortem should capture four things: the assumptions that broke, the decisions that aged badly, the estimate-versus-actual on cost and timeline, and the containment and quality numbers against what you promised the board. Everything else is colour. Dilr.ai structures the review around those four ledgers because they are the ones that change the next decision, and because they are specific enough that a successor team can act on them without having been in the room.
The sequence below is the method we run, and it deliberately ends by handing the output forward rather than filing it.
The programme post-mortem sequenceEach stage produces an input the next funding decision can use; the review is not finished until stage five hands it forward.
The estimate-versus-actual ledger is the one that earns the meeting. Vendors routinely quote seven to twelve months to value; independent reviews put meaningful ROI closer to two to four years, and the delta is exactly the kind of assumption a post-mortem exists to surface. Capture the real integration effort against the plan, whether the Twilio or CRM connection took the fortnight you scoped or the quarter it actually needed, and whether the change-management and workforce redeployment work was resourced or wished for. The point is not to relitigate the sourcing decision between building in-house or buying a managed platform, which your operating model analysis already made; it is to record how that decision actually played out so the next one is better informed.
For the small minority of programmes that do capture and act on their lessons, the payoff shows up in the numbers. Gartner's survey of 822 business leaders found that organisations realising value from AI reported the following average gains.
What value-capturing AI programmes reportAverage gains reported by leaders realising value from AI, from Gartner's survey of 822 business leaders (2024). Source: Gartner, GenAI project abandonment (Jul 2024)
How do the lessons feed the next funding cycle?
Lessons feed the next cycle only when the review has an owner, a named decision it informs, and a deadline that sits before that decision is made. A post-mortem that produces a document but no change to the next business case has failed, however honest. Dilr.ai ties each finding to a specific downstream action: a revised integration estimate for the next use case, a hard prerequisite added to the go-live checklist, or a governance change that stops a repeat.
This is not a novel idea; it is how the public sector is required to spend. HM Treasury's Green Book on appraisal and evaluation in central government makes evaluation a formal stage of the spending cycle precisely so that findings feed back into future decisions rather than dying with the project. Enterprises rarely enforce the same loop for AI, which is why the same voice AI mistakes recur across business units. The fix is procedural: the post-mortem output becomes a required input to the next funding paper, reviewed by whoever holds the AI execution office or the equivalent programme governance seat, and the use-case prioritisation roadmap is re-ranked in light of it. When learning is a gate rather than a nicety, the abandonment rate that Gartner and RAND describe starts to fall.
The same governance discipline is what we build into Dilr Voice deployments, where evaluation and feedback run as a standing cadence rather than a task bolted on at the end.
What is the best format for a voice AI programme post-mortem in 2026?
The best format in 2026 is proportionate: match the depth of the review to the size of the decision it feeds. There is no single template that wins every time, and a team that runs the same three-hour workshop for a minor phase and a year-one renewal will burn goodwill on the first and under-invest on the second. The honest answer is a tiered approach, with a clear concession about where a lighter or heavier tool is right.
For a single production incident, do not run a programme post-mortem at all: the Phase 06 blameless review in your incident runbook is the correct, lighter tool, and reaching for a full programme review there is overkill that delays the fix. For a routine phase boundary, a ninety-minute structured review against the four ledgers is enough. Reserve the full format, written pre-reads, a neutral facilitator, estimate-versus-actual reconstruction and a board-facing findings paper, for the year-one review and pre-expansion gates where real money rides on the decision. On tooling, the format is deliberately vendor-neutral: whether you built on Vapi or Retell AI, bought a managed platform such as PolyAI, or run Dilr Voice, the four questions are identical, because the failures the RAND research documents are organisational, not technical. Where a competitor wins is speed: a small team shipping a low-risk internal agent is right to run a fifteen-minute retro and move on, and pretending otherwise is process for its own sake.
Is a programme post-mortem the same as an incident post-mortem?
No. An incident post-mortem examines a single failure, its detection, containment and cause, and lives inside operational incident response. A programme post-mortem examines the whole initiative, its sourcing, economics, integration and change management, and feeds decisions about how you deploy voice AI next. They share the blameless principle but differ in scope, cadence and audience. Run the incident review within days of an outage, the programme review at phase and year-one boundaries; each answers a question the other cannot.
Who should own the voice AI programme post-mortem?
The programme owner should own it, not the vendor and not the pilot team, because ownership must sit with whoever carries the next funding decision. In practice that is the AI programme lead or the operating-model owner, supported by a neutral facilitator who did not deliver the work. The vendor contributes evidence but cannot mark its own homework. Dilr.ai's view is blunt: if no named person owns the review and its actions, the lessons will not survive the next reorganisation.
Written by the Dilr.ai engineering team, practitioners who ship enterprise AI in production. Follow us on LinkedIn for shipping notes, or subscribe via the RSS feed.
voice AI programme post-mortem enterprisevoice AI lessons learnedvoice AI programme retrospectiveblameless post-mortem AIvoice AI programme review redditbest voice AI programme review 2026Dilr Voice
Questions this article answers
What is a voice AI programme post-mortem?
A voice AI programme post-mortem is a structured, blameless review of a whole deployment, run at a natural boundary such as the end of a phase or the year-one mark, to capture what the programme actually taught you. Dilr.ai treats it as a governance step, not a retrospective ritual: it records the assumptions that broke, the decisions that aged badly, and the gap between forecast and actual, so the next funding cycle starts from evidence rather than memory.
Why do so many voice AI programmes lose their hardest lessons?
Programmes lose their lessons because the people who hold them leave the room. By the time a voice AI deployment reaches steady state, the pilot squad has rotated, the launch engineer has moved on, and the sponsor has claimed the win. Nobody owns the memory. The result is a value gap: McKinsey's 2025 data shows most enterprises now use AI, but only a small fraction reach production and fewer still capture EBIT impact, and the difference is rarely the model.
When should you run a voice AI programme post-mortem?
Run a programme post-mortem at four moments: the end of each delivery phase, immediately after a major incident once the fire is out, at the year-one review before the renewal decision, and before any expansion into a second use case. Each trigger asks a different question, so Dilr.ai recommends scheduling the phase and year-one reviews in advance rather than waiting for a crisis to force one. A review you have to argue for is a review that will be skipped.
How do you run a blameless post-mortem that gets honest input?
You get honest input by making it structurally safe to be honest. A blameless review separates the account of what happened from any judgement of who is at fault, so contributors describe decisions as they looked at the time rather than defending them in hindsight. Google's Site Reliability Engineering team, which popularised the practice for production failures, states the principle directly, and the same rule is what makes a programme review usable rather than performative.
What should a voice AI post-mortem actually capture?
A voice AI post-mortem should capture four things: the assumptions that broke, the decisions that aged badly, the estimate-versus-actual on cost and timeline, and the containment and quality numbers against what you promised the board. Everything else is colour. Dilr.ai structures the review around those four ledgers because they are the ones that change the next decision, and because they are specific enough that a successor team can act on them without having been in the room.
How do the lessons feed the next funding cycle?
Lessons feed the next cycle only when the review has an owner, a named decision it informs, and a deadline that sits before that decision is made. A post-mortem that produces a document but no change to the next business case has failed, however honest. Dilr.ai ties each finding to a specific downstream action: a revised integration estimate for the next use case, a hard prerequisite added to the go-live checklist, or a governance change that stops a repeat.
What is the best format for a voice AI programme post-mortem in 2026?
The best format in 2026 is proportionate: match the depth of the review to the size of the decision it feeds. There is no single template that wins every time, and a team that runs the same three-hour workshop for a minor phase and a year-one renewal will burn goodwill on the first and under-invest on the second. The honest answer is a tiered approach, with a clear concession about where a lighter or heavier tool is right.
Is a programme post-mortem the same as an incident post-mortem?
No. An incident post-mortem examines a single failure, its detection, containment and cause, and lives inside operational incident response. A programme post-mortem examines the whole initiative, its sourcing, economics, integration and change management, and feeds decisions about how you deploy voice AI next. They share the blameless principle but differ in scope, cadence and audience. Run the incident review within days of an outage, the programme review at phase and year-one boundaries; each answers a question the other cannot.
DE
Dilr.ai Engineering
Engineering team
AI consulting (DATS)
Place AI where the P&L moves
The DATS system runs from a fixed-fee placement diagnostic through to embedded delivery, so AI reaches production instead of staying a pilot.