Voice AI Baseline: The Pre-Pilot Measurement Guide
In short
A pre-pilot baseline is a dated record of how your contact centre performs before a voice AI agent goes live. Dilr Voice deployments fix handle time, first-contact resolution, containment, transfer rate, cost per contact and CSAT across a full cycle, so uplift can be proved against a counterfactual rather than claimed after the fact.
DE
Dilr.ai EngineeringEngineering team
Published Sep 4, 2026Read 12 min
Every voice AI business case promises a number: shorter calls, higher first-contact resolution, lower cost to serve. Then the pilot goes live, the agent starts taking calls, and six months later someone in finance asks the only question that matters. Did it actually work, and by how much? For most programmes there is no honest answer, because nobody wrote down what the operation looked like before anything changed. The uplift is real, but it cannot be proved, because the starting point is gone.
This is the pre-pilot baseline problem, and it is the quietest way a good deployment loses its own business case. Across McKinsey's 2025 State of AI research, around 88% of organisations report using AI, but only about 14% say it is moving earnings, and roughly 6% are genuinely mature. A large part of that gap is not model quality. It is measurement. Programmes that never baselined their old process cannot separate the effect of the agent from seasonality, a good quarter, or a change in the mix of calls, so the value stays anecdotal and the next budget round treats it as a cost.
This guide is shipped by the team behind Dilr Voice, enterprise voice AI built for regulated deployments. Or see DATS, our five-stage AI consulting system.
This article is the discipline of baselining a contact centre before a voice AI agent takes a single live call: which metrics to capture, over what window, how to keep the comparison fair, and how the finished baseline becomes the reference the business case and the go-live scorecard both depend on. It is deliberately a method, not a benchmark table. The numbers that matter are yours, measured in your operation, not a sector average that flatters or scares you.
What is a pre-pilot baseline for a voice AI programme?
A pre-pilot baseline is a measured, dated record of how your contact centre performs today, captured before a voice AI agent handles any live traffic. For a Dilr Voice deployment it fixes the current values of handle time, first-contact resolution, containment, transfer rate, cost per contact and customer satisfaction over a defined window, so any change after go-live can be attributed to the agent rather than to seasonality or a shifting call mix.
The word baseline is used loosely across the industry, so it helps to be precise. It is not a target, and it is not a forecast. It is the observed state of the operation you are about to change, recorded while the old process is still running and still reconstructable. Once traffic moves to the agent, the previous distribution of calls, handle times and transfers is gone, and any later attempt to describe the starting point is an estimate. A baseline captured in advance replaces that estimate with a measurement, which is the difference between a business case you can defend and one you have to argue.
Why do most voice AI programmes fail to prove their uplift?
Most programmes fail to prove uplift because they measure the result without fixing the counterfactual, the picture of what the operation would have done without the agent. When containment climbs and calls get shorter, that improvement has to be read against something, and if the something was never recorded the credit is contestable. A sceptical finance reviewer can say the quarter was quieter or the call mix simpler, and without a baseline there is no evidence to answer them.
The discipline is not unique to voice AI. The Magenta Book, HM Treasury's guidance for evaluating government programmes, puts the principle plainly: "The costs and benefits of each policy option must be assessed relative to the counterfactual." A voice AI pilot is an intervention like any other. The counterfactual is the human and IVR operation you baselined before you switched anything on, and it is the only fair thing to measure the pilot against. Skip it, and you are left describing a number with no honest denominator.
The scale of the problem shows up in the macro data. Deployment is easy and near-universal now, but the drop from deployment to proven earnings impact is severe, and a good part of that fall is programmes that cannot evidence what they changed.
Where enterprise AI value leaks outShare of enterprises reaching each stage of AI value capture, 2025 to 2026; the fall from deployment to proven EBIT impact is where unbaselined programmes disappear. Source: McKinsey, The State of AI (Nov 2025)
This is where a disciplined AI placement approach earns its keep. Fixing the baseline before automation is not busywork; it is what turns the post-launch numbers into an argument the board will accept. Our ROI attribution guide covers the credit-allocation half of the same problem, once the baseline exists to attribute against.
Which metrics should you baseline before a voice AI pilot?
Baseline the operational metrics the agent will move and the ones a finance reviewer will ask about, not every number your platform can export. Six carry most of the weight for a Dilr Voice deployment: average handle time, first-contact resolution, containment, transfer rate, cost per contact, and customer satisfaction. Capture each with a written definition, a value, the window it covers and the source it came from, because a metric without its definition is not comparable after go-live.
The six are worth being exact about, since half the disputes later come from a definition drift nobody noticed.
Average handle time. Talk time plus after-call work, per contact. Decide up front whether hold and wrap are included, because the agent will change that split.
First-contact resolution. The share of contacts resolved without a repeat within a set window. Fix the window now, or the metric is meaningless. Our first-contact resolution measurement guide goes deeper on this single metric.
Containment or deflection. The share of contacts handled without a human, in your current IVR or self-service. This is the number the agent most directly attacks, and the containment rate benchmark explains why an aggregate figure hides more than it shows.
Transfer rate. The share of contacts passed to a human, and to which queue. Baseline it per destination, not in total.
Cost per contact. Fully loaded, including the telephony, the platform and the labour the contact consumes. Finance will rebuild this from their own ledger, so agree the method before you publish a number.
Customer satisfaction. Whatever instrument you already run, CSAT or otherwise, kept identical across the baseline and the pilot so the comparison holds.
Resist the urge to baseline a routing distribution here as well. Capturing which calls go where before go-live is real and necessary, but it belongs to the launch and adoption phase; our post-go-live adoption metrics guide owns the routing baseline and the first-fortnight read, and duplicating it now only muddies which artefact is authoritative.
How long should you measure the baseline?
Measure long enough to cover a full business cycle, so the baseline reflects your normal range rather than a fortnight that was calm or chaotic. For most contact centres that means at least one month, ideally a quarter, spanning a peak and a trough. A baseline built from a quiet week overstates the operation and sets a target the pilot cannot fairly meet; one from a spike does the opposite. Capture the shape of demand, not a snapshot.
Two practical constraints pull against a long window. The programme has momentum and nobody wants to wait a quarter to start, and the operation may be changing anyway, which erodes the very stability you are trying to record. The resolution is to start the baseline the moment the pilot is approved, run it in parallel with the build, and freeze it before the first live call. That way the window is as long as the build allows without adding a single day to the timeline. If your demand is strongly seasonal, note the season the baseline covers explicitly, because a January baseline compared against a July pilot is a comparison of two operations, not one.
How do you control the baseline so the comparison is fair?
Control the baseline by recording anything that could be mistaken for the agent's effect, so the comparison is like for like. Four distortions do most of the damage: seasonality, the mix of contact types, agent tenure, and one-off events such as an outage or a campaign. None can be removed, but all can be documented and, where the design allows, held constant. A baseline that names its distortions is harder to argue with than one pretending the period was clean.
The same diagnostic logic underpins our AI execution office, where a standing team runs exactly this kind of measurement discipline as a recurring task rather than a one-off exercise.
The most reliable control is structural. Rather than switching all traffic at once, hold back a comparable slice of contacts on the old process and run the agent alongside it, so the two are measured over the same period under the same conditions. Where a live holdout is impractical, the fallback is a careful before-and-after read against the frozen baseline, with the known distortions written down and adjusted for. Measure per queue rather than in aggregate, because an aggregate uplift can hide a queue that got worse, and a fair baseline is one that can show both. This is the difference between measuring change and inferring it.
How does the baseline feed the business case and the go-live scorecard?
The baseline becomes a single artefact, the baseline pack, that two later documents consume directly. Each metric appears once, with its definition, its measured value, the window it covers, its source and a named owner. The business case reads the pack to state the gap it intends to close in your own numbers, and the go-live scorecard reads the same pack to set the threshold the pilot must beat before more traffic is trusted. One measured source, two decisions.
The pre-pilot baseline sequenceEach step feeds the next; the pack is the single source the business case and the scorecard both read.
The pack pulls from systems you already run: the ACD and telephony reporting in a platform like Twilio, Genesys or Amazon Connect, the IVR logs, the quality scorecards, workforce management for shrinkage, and finance for the cost line. Keeping the baseline distinct from what comes next matters. It is not a benefits register, which tracks realised value quarter by quarter after the fact, and it is not the business case model itself, which turns the gap into pounds and a payback period. The baseline is the measured floor all of those stand on, and wiring it in early is a core part of any credible AI operating model. You can see the wider method on our strategy blog and in the DATS methodology.
What is the best way to run a voice AI baseline in 2026?
The best approach in 2026 is a short, structured baseline run in parallel with the build, owned by a named analyst, measured per queue, and frozen before the first live call. Tooling matters less than discipline. Self-serve platforms such as Vapi, Retell AI, Bland AI and Synthflow let you launch an agent in an afternoon, and for a low-risk internal line nobody will audit, a formal baseline may be overkill. Be honest about which scenario you are in.
For a regulated, board-visible deployment the trade is the other way. When the uplift will be quoted to a finance committee or a regulator, the baseline is the evidence, and a governed platform such as Dilr Voice or a specialist like PolyAI is built to preserve the audit trail that a self-serve tool treats as optional. The honest verdict: match the rigour to the stakes. A throwaway agent needs no baseline; a programme whose value will be challenged needs one that a hostile reviewer cannot dismiss. If you are unsure which you are running, you are probably running the second. The go-live checklist sets out where the frozen baseline sits in the wider launch sequence.
The same measurement discipline runs through our AI operating model consulting, where the baseline is treated as a standing asset rather than a one-off task, and revisited whenever the operation itself changes.
Can you baseline after the pilot has started?
Not properly. A baseline has to be captured while the old process is still running, because once traffic moves to the agent the previous distribution of calls, handle times and transfers is gone. Anything assembled afterwards is a reconstruction, and a reconstructed starting point is exactly the weakness a sceptical reviewer will attack. If a pilot has already begun without a baseline, the honest move is to say so and rely on a structured holdout from here.
Is a voice AI baseline the same as benefits realisation?
No. A baseline is the one-off measurement of the operation before the pilot, taken while the old process still runs. Benefits realisation is the ongoing tracking of value after go-live, quarter by quarter, against that baseline. They meet at the same reference point, but they are different disciplines with different owners. The baseline is captured once and frozen; the benefits register is a living document reviewed for the life of the programme.
What data protection applies to baseline call data?
Baseline data often includes call recordings and customer records, so the same UK GDPR duties apply as to any processing. UK GDPR's data minimisation principle, which the ICO enforces, means a baseline should use aggregated operational metrics wherever possible rather than retaining raw recordings, and any recordings used should sit under your existing lawful basis and retention schedule. Treat the baseline dataset as in scope for normal privacy governance, not a temporary exception because it is only for measurement.
30-min scoping call · No deck · Confidential. We will help you fix a baseline your finance committee cannot argue with, before a single call is automated.
Written by the Dilr.ai engineering team, practitioners who ship enterprise AI in production. Follow us on LinkedIn for shipping notes, or subscribe via the RSS feed.
voice AI pre-pilot baseline measurementcontact centre baseline before AIvoice AI baseline metricsmeasure AI uplift voicevoice ai baseline redditbest voice AI baseline approach 2026Dilr Voice
Questions this article answers
What is a pre-pilot baseline for a voice AI programme?
A pre-pilot baseline is a measured, dated record of how your contact centre performs today, captured before a voice AI agent handles any live traffic. For a Dilr Voice deployment it fixes the current values of handle time, first-contact resolution, containment, transfer rate, cost per contact and customer satisfaction over a defined window, so any change after go-live can be attributed to the agent rather than to seasonality or a shifting call mix.
Why do most voice AI programmes fail to prove their uplift?
Most programmes fail to prove uplift because they measure the result without fixing the counterfactual, the picture of what the operation would have done without the agent. When containment climbs and calls get shorter, that improvement has to be read against something, and if the something was never recorded the credit is contestable. A sceptical finance reviewer can say the quarter was quieter or the call mix simpler, and without a baseline there is no evidence to answer them.
Which metrics should you baseline before a voice AI pilot?
Baseline the operational metrics the agent will move and the ones a finance reviewer will ask about, not every number your platform can export. Six carry most of the weight for a Dilr Voice deployment: average handle time, first-contact resolution, containment, transfer rate, cost per contact, and customer satisfaction. Capture each with a written definition, a value, the window it covers and the source it came from, because a metric without its definition is not comparable after go-live.
How long should you measure the baseline?
Measure long enough to cover a full business cycle, so the baseline reflects your normal range rather than a fortnight that was calm or chaotic. For most contact centres that means at least one month, ideally a quarter, spanning a peak and a trough. A baseline built from a quiet week overstates the operation and sets a target the pilot cannot fairly meet; one from a spike does the opposite. Capture the shape of demand, not a snapshot.
How do you control the baseline so the comparison is fair?
Control the baseline by recording anything that could be mistaken for the agent's effect, so the comparison is like for like. Four distortions do most of the damage: seasonality, the mix of contact types, agent tenure, and one-off events such as an outage or a campaign. None can be removed, but all can be documented and, where the design allows, held constant. A baseline that names its distortions is harder to argue with than one pretending the period was clean.
How does the baseline feed the business case and the go-live scorecard?
The baseline becomes a single artefact, the baseline pack, that two later documents consume directly. Each metric appears once, with its definition, its measured value, the window it covers, its source and a named owner. The business case reads the pack to state the gap it intends to close in your own numbers, and the go-live scorecard reads the same pack to set the threshold the pilot must beat before more traffic is trusted. One measured source, two decisions.
What is the best way to run a voice AI baseline in 2026?
The best approach in 2026 is a short, structured baseline run in parallel with the build, owned by a named analyst, measured per queue, and frozen before the first live call. Tooling matters less than discipline. Self-serve platforms such as Vapi, Retell AI, Bland AI and Synthflow let you launch an agent in an afternoon, and for a low-risk internal line nobody will audit, a formal baseline may be overkill. Be honest about which scenario you are in.
Can you baseline after the pilot has started?
Not properly. A baseline has to be captured while the old process is still running, because once traffic moves to the agent the previous distribution of calls, handle times and transfers is gone. Anything assembled afterwards is a reconstruction, and a reconstructed starting point is exactly the weakness a sceptical reviewer will attack. If a pilot has already begun without a baseline, the honest move is to say so and rely on a structured holdout from here.
DE
Dilr.ai Engineering
Engineering team
AI consulting (DATS)
Place AI where the P&L moves
The DATS system runs from a fixed-fee placement diagnostic through to embedded delivery, so AI reaches production instead of staying a pilot.