Voice AI

Voice AI Caller Identity Verification: Enterprise Guide

Dilr Voice is an enterprise voice AI platform built to verify a caller's identity before any sensitive action. This guide explains non-biometric caller identity verification: knowledge and possession factors, risk-based step-up, safe failure paths, and the audit trail auditors expect, grounded in UK Finance fraud data and the Payment Services Regulations 2017.

DILR.AI ENGINEERING Caller ID&V for voice AI Knowing who is on the line before anything sensitive happens 01 Identify the claim 02 Assess the action risk 03 Step up: knowledge + possession 04 Pass and log, or fail safely £576.4m UK APP fraud losses, 2025 17% of APP cases begin over the phone

Every enterprise voice AI programme eventually hits the same wall: the agent is about to do something that matters, moving money, changing a delivery address, resetting a password, disclosing an account balance, and it has no reliable idea who is actually on the line. A confident voice reading out a name and postcode is not proof of identity. It is the opening move of most social engineering scripts.

The stakes are no longer theoretical. UK Finance reported that authorised push payment (APP) fraud losses rose to £576.4 million in 2025, up 19 per cent on the year, across 248,070 cases. Of those cases, 17 per cent began over the phone, and those telephone-originated scams accounted for 28 per cent of the losses because they tend to be higher value. The phone channel is where the expensive frauds start, which is exactly the channel a voice agent answers.

Caller identity verification is the control that decides whether an autonomous voice agent is a productivity gain or a fraud liability. This guide covers the non-biometric verification workflow: what to check, when to step up, what to do when a caller fails, and the audit trail a regulator will ask for. It deliberately leaves voice biometrics to one side, because most enterprises can and should get the workflow right before they touch a voiceprint.

This guide is shipped by the team behind Dilr Voice, enterprise voice AI built for regulated deployments. Or see DATS, our five-stage AI consulting system for placing AI where the risk and the return actually sit.

What is caller identity verification for a voice AI agent?

Caller identity verification is the process a voice AI agent uses to gain confidence that the person on the line is who they claim to be, before it performs any action that depends on that identity. It combines something the caller knows, something they hold, and the risk of the action requested, into a single yes-or-no gate. Dilr Voice treats it as a per-action decision, not a one-time check at the start of a call.

The distinction that trips most teams up is between recognition and verification. A caller who states an account number has been recognised, the system knows which record they are pointing at, but they have not been verified. Verification is the deliberate act of raising confidence to a level that matches the sensitivity of the requested action. Reading out a balance needs less confidence than moving funds to a new payee. Designing an AI voice agent that treats those two requests identically is how programmes leak both fraud losses and customer trust, and it is a failure mode our DATS five-stage system is built to surface early.

Why does the phone channel need identity verification in 2026?

The phone channel needs verification because it carries the most costly frauds and the weakest default controls. UK Finance recorded £576.4 million in APP fraud in 2025, and telephone-originated cases, though only 17 per cent of the volume, drove 28 per cent of the losses. Unauthorised fraud added a further £703.4 million, and criminals increasingly compromise one-time passcodes through social engineering. A voice agent that authorises actions without verifying is an open door.

The chart below shows why the phone deserves specific attention rather than a copied-across web control. Losses concentrate in the channels where a human, or a convincing agent, can be talked into acting.

Share of UK APP fraud losses by origin channel, 2025
32%Online28%Telecoms24%Unknown8%Other7%Email
Telephone-originated scams are only 17 per cent of APP cases but 28 per cent of losses, a far higher loss-per-case than the online channel where most cases begin. Source: UK Finance, Annual Fraud Report 2026

There is a second reason the timing matters. Since 7 October 2024, the Payment Systems Regulator has required in-scope payment firms to reimburse victims of Faster Payments APP scams, up to a maximum of £85,000 per claim, with the cost split between sending and receiving firms. When verification fails and money leaves, the sending or receiving payment firm now often carries the bill directly. That reimbursement exposure turns caller verification from a compliance nicety into a line item, and it is a strong argument for building the control into your voice AI operating model rather than bolting it on after an incident.

How does a voice AI agent verify a caller without voice biometrics?

A voice agent verifies without biometrics by combining independent factors: something the caller knows, such as a memorable detail or a transaction they can describe, and something they possess, such as a registered device that receives a one-time passcode. It reserves the strongest factors for the highest-risk actions, and it never treats caller line identity (the number shown) as proof, because that number can be spoofed. Dilr Voice orchestrates these factors as a scored decision.

The regulated benchmark here is strong customer authentication. Under the UK Payment Services Regulations 2017, payment firms must authenticate using two or more independent elements drawn from knowledge, possession and inherence. The regulation defines the standard precisely:

"authentication based on the use of two or more elements that are independent, in that the breach of one element does not compromise the reliability of any other element, and designed in such a way as to protect the confidentiality of the authentication data"

Those three categories, knowledge (something the caller knows), possession (something the caller holds), and inherence (something the caller is), are the design vocabulary even for enterprises that fall outside the regulation. A utility, a housing association or an insurer is not a payment service provider, but the same taxonomy tells them how to layer factors. In practice a well-built voice AI agent leans on knowledge plus possession, sends a one-time passcode through an integration such as Twilio, and confirms it in-call, keeping voice biometrics as a separate and consent-heavy path covered in our guide to voice biometric data security. Knowledge-based checks alone, the security questions many contact centres still rely on, are the weakest link: the answers are often discoverable, and treating them as sufficient for a payment is the gap fraudsters walk through.

When should a voice agent step up verification?

A voice agent should step up verification the moment the requested action crosses a risk threshold, not at a fixed point in the call. Reading general account information may need only a light knowledge check. Moving money, changing contact details, adding a payee or disclosing sensitive data should each trigger a stronger, possession-based factor. The principle is proportionality: verification effort rises with the harm a mistaken action would cause. Dilr Voice binds each intent to a required confidence level.

The same diagnostic logic underpins our AI placement diagnostic, a fixed-fee assessment used before any deployment commitment, which maps each customer intent to the verification it warrants. The flow below shows the ladder in practice: the agent identifies the claimed identity, scores the risk of what is being asked, and only then decides whether the light-touch check it already has is enough or whether it must escalate to a stronger factor before proceeding.

The step-up verification ladder
01Identify the claimCaller states who they are02Score the action riskRead-only vs money-movement03Low risk: light knowledge checkProceed and log04High risk: step upAdd a possession factor05Confirm one-time passcodeRegistered device only06Pass: complete and recordFail: safe fallback
Verification effort is bound to the risk of the requested action, not applied as a single gate at the start of the call.

Two design rules keep the ladder honest. First, step-up state should not carry across actions without re-checking: a caller verified to read a balance has not earned the right to move funds. Second, the thresholds belong in configuration that is versioned and approved, not buried in a prompt an operator can edit unseen, which is why we treat them as part of release management for a voice agent. Getting the thresholds right is a judgement call about your own fraud exposure, and it is one our AI operating model consulting works through with risk and compliance teams rather than leaving to engineering alone.

What happens when a caller fails verification?

When a caller fails verification, the agent must fail safely: it stops short of the sensitive action, does not disclose why the check failed, and routes the caller down a pre-agreed path rather than improvising. Depending on risk, that path is a retry with a different factor, a warm transfer to a trained human, or a hard stop with a callback to a number already on file. Dilr Voice never lets a failed check quietly degrade into a completed action.

The information-leak risk is subtle and important. An agent that says "that answer was wrong, try your date of birth instead" is coaching an attacker through the factors one by one. The correct behaviour is a neutral response and a controlled escalation, ideally a warm transfer that carries context so the customer does not repeat everything to a human. Failure paths also need their own audit records: a spike in failed verifications on a particular action is an early fraud signal, and instrumenting it is part of good voice AI observability. The flow below sets out the decision.

Handling a failed verification
01Verification failsNeutral response only02Assess remaining riskWhat was being asked03Low risk: offer one retryDifferent factor04High risk: warm transferContext carried to a human05Record the failed attemptFeed fraud monitoring
A failed check ends in a controlled outcome, never a leaked hint or a silently completed action.

Handled well, the failure path protects the genuine customer too. Real people forget details, ring from a new phone, or fail a check for entirely innocent reasons, and a punitive dead-end drives them to the branch or the complaints queue. A humane fallback, with a clear route to a person, is both a fraud control and a customer-experience decision, and balancing the two is exactly the kind of trade-off our DATS five-stage AI methodology is built to work through.

What do auditors and regulators expect from voice AI verification?

Auditors and regulators expect a documented, consistent and evidenced verification process, not a best-effort one. They want to see which factors were used, why an action required a step-up, and a tamper-evident log of every decision tied to the call. For payment firms the Financial Conduct Authority and the Payment Systems Regulator set the bar; for everyone else, the Information Commissioner's Office expects you to verify identity using the least data necessary. Dilr Voice produces that audit trail by default.

Data minimisation is the constraint teams most often miss. Under UK GDPR Article 5(1)(c), you must collect only the personal data you actually need, so verification should confirm identity without hoarding new sensitive attributes: a passcode confirmation and a scoped knowledge check, not a growing file of memorable data. The same discipline underpins our guide to data minimisation and redaction. On the financial side, the reimbursement regime raises the cost of weak verification directly: the PSR's rules mean in-scope firms reimburse APP scam victims up to £85,000 per claim, and UK Finance reports that banks paid out £354.3 million to APP victims in 2025, equal to 61 per cent of losses. A verification failure that lets a scam through is now, increasingly, a bill the firm pays. If a verified interaction still results in a breach of personal data, the Article 33 notification clock starts regardless, so the audit trail matters twice over.

The same governance logic runs through our AI execution office, which stands up the controls, evidence and reporting a regulated voice deployment needs before it scales.

What is the best way to verify caller identity on voice AI in 2026?

The best way to verify caller identity in 2026 is a layered, risk-based workflow that combines a knowledge factor with a possession factor, steps up by action, fails safely, and logs every decision, rather than any single clever technique. The right platform choice depends on your regulatory exposure and call mix. For most regulated enterprises, a governed platform beats a build-your-own stack, but there are cases where the reverse is true, and an honest answer names them.

Build-your-own toolkits such as Vapi, Retell AI and Bland AI give engineering teams maximum control and are a reasonable choice for a low-volume, low-risk line where no money moves and no sensitive data is disclosed, because the verification burden is genuinely light. Once real value or regulated data is on the line, the calculus shifts: a governed platform like Dilr Voice or PolyAI, wired into your identity and passcode infrastructure through Twilio, and into your systems of record through Salesforce, HubSpot or Stripe, gives you the audit trail, the versioned thresholds and the failure-path controls that auditors expect out of the box. The concession is real, and worth stating plainly: if you have a single simple use case and strong in-house security engineering, building it yourself can be cheaper and perfectly compliant. If you are running many actions across many teams under FCA or ICO scrutiny, buying the governance is almost always the better trade, and our DATS methodology exists to make that build-versus-buy call with evidence.

Want to see this in production? Try Dilr Voice live, stand up an AI execution office, read the enterprise voice AI guide, or browse more voice AI engineering notes.

Is knowledge-based verification still acceptable on its own?

Knowledge-based verification, the classic security questions, is no longer acceptable on its own for anything sensitive. Answers such as a mother's maiden name or a first school are often discoverable or already breached, so a voice agent that treats them as sufficient for a payment is relying on a factor attackers can research. Dilr Voice permits knowledge checks only as one layer, always paired with a possession factor when the action carries real risk.

Does caller ID (CLI) count as identity verification?

No. Caller line identity, the number that appears when a call arrives, does not count as verification because it can be spoofed cheaply, and UK Finance data shows telephone scams tend to be higher in value than online ones. A voice AI agent may use CLI as a weak signal to personalise or to raise suspicion, but it must never treat a matching number as proof of who is calling. Genuine verification always requires an active factor the caller supplies.

Can you verify a caller without storing extra personal data?

Yes, and you generally should. UK GDPR Article 5(1)(c) requires data minimisation, so a well-designed voice agent verifies against data you already hold and confirms a one-time passcode without creating a new store of memorable secrets. Dilr Voice is built to verify with the least data necessary, keeping the verification log itself scoped and access-controlled so the control does not become its own privacy liability.

Service
AI Placement Diagnostic
Service
AI Operating Model
Product
Dilr Voice
Talk to the operators

Verify the caller before the risk lands.

30-min scoping call · No deck · Confidential. We will map where identity risk sits in your call flows, and what verification each action actually needs.

Written by the Dilr.ai engineering team, practitioners who ship enterprise AI in production. Follow us on LinkedIn for shipping notes, or subscribe via the RSS feed.

voice AI caller identity verification enterpriseknowledge-based verification voice AIstep-up authentication voice agentverify caller before sensitive actionvoice ai identity verification redditbest voice AI caller verification 2026Dilr Voice caller verification

Questions this article answers

What is caller identity verification for a voice AI agent?

Caller identity verification is the process a voice AI agent uses to gain confidence that the person on the line is who they claim to be, before it performs any action that depends on that identity. It combines something the caller knows, something they hold, and the risk of the action requested, into a single yes-or-no gate. Dilr Voice treats it as a per-action decision, not a one-time check at the start of a call.

Why does the phone channel need identity verification in 2026?

The phone channel needs verification because it carries the most costly frauds and the weakest default controls. UK Finance recorded £576.4 million in APP fraud in 2025, and telephone-originated cases, though only 17 per cent of the volume, drove 28 per cent of the losses. Unauthorised fraud added a further £703.4 million, and criminals increasingly compromise one-time passcodes through social engineering. A voice agent that authorises actions without verifying is an open door.

How does a voice AI agent verify a caller without voice biometrics?

A voice agent verifies without biometrics by combining independent factors: something the caller knows, such as a memorable detail or a transaction they can describe, and something they possess, such as a registered device that receives a one-time passcode. It reserves the strongest factors for the highest-risk actions, and it never treats caller line identity (the number shown) as proof, because that number can be spoofed. Dilr Voice orchestrates these factors as a scored decision.

When should a voice agent step up verification?

A voice agent should step up verification the moment the requested action crosses a risk threshold, not at a fixed point in the call. Reading general account information may need only a light knowledge check. Moving money, changing contact details, adding a payee or disclosing sensitive data should each trigger a stronger, possession-based factor. The principle is proportionality: verification effort rises with the harm a mistaken action would cause. Dilr Voice binds each intent to a required confidence level.

What happens when a caller fails verification?

When a caller fails verification, the agent must fail safely: it stops short of the sensitive action, does not disclose why the check failed, and routes the caller down a pre-agreed path rather than improvising. Depending on risk, that path is a retry with a different factor, a warm transfer to a trained human, or a hard stop with a callback to a number already on file. Dilr Voice never lets a failed check quietly degrade into a completed action.

What do auditors and regulators expect from voice AI verification?

Auditors and regulators expect a documented, consistent and evidenced verification process, not a best-effort one. They want to see which factors were used, why an action required a step-up, and a tamper-evident log of every decision tied to the call. For payment firms the Financial Conduct Authority and the Payment Systems Regulator set the bar; for everyone else, the Information Commissioner's Office expects you to verify identity using the least data necessary. Dilr Voice produces that audit trail by default.

What is the best way to verify caller identity on voice AI in 2026?

The best way to verify caller identity in 2026 is a layered, risk-based workflow that combines a knowledge factor with a possession factor, steps up by action, fails safely, and logs every decision, rather than any single clever technique. The right platform choice depends on your regulatory exposure and call mix. For most regulated enterprises, a governed platform beats a build-your-own stack, but there are cases where the reverse is true, and an honest answer names them.

Is knowledge-based verification still acceptable on its own?

Knowledge-based verification, the classic security questions, is no longer acceptable on its own for anything sensitive. Answers such as a mother's maiden name or a first school are often discoverable or already breached, so a voice agent that treats them as sufficient for a payment is relying on a factor attackers can research. Dilr Voice permits knowledge checks only as one layer, always paired with a possession factor when the action carries real risk.

Dilr Voice

Put this into production

Dilr Voice runs AI voice agents for inbound and outbound calls: multi-agent handoff, RAG knowledge bases, and per-country compliance in one platform.

Related articles

← Previous
AI voice for gyms and leisure centres: a 2026 guide

One email, once a month. No hype. Just what we learned shipping.