Voice AI

AI Voice Agent Security Questionnaire: A UK Guide

Dilr Voice is an enterprise voice AI platform from DILR.AI that handles live customer calls, the workload a security questionnaire must cover: call audio, transcripts and the telephony carrier. This guide shows UK buyers how to build that questionnaire from the NCSC Cloud Security Principles and CAIQ, and which voice-specific rows to add.

AI Voice Agent Security Questionnaire: A UK Guide DILR VOICE AI Voice Agent Security Questionnaire: A UK Guide 48 % of large UK businesses reviewed cyber risk from immediate suppliers Source: DSIT Cyber Security Breaches Survey 2025/26 dilr.ai/blog

A voice AI vendor reaches procurement, and the CIO's team sends over the standard security questionnaire: two hundred rows written for a generic software service. The vendor answers every row, the pack is filed, and nobody notices that the questionnaire never asked where the call audio goes, who can replay it, which telephony carrier sits in the path, or whether recordings are used to train anything. The pack was complete. The review was not.

The gap is not unusual. The Cyber Security Breaches Survey 2025/2026, published by the Department for Science, Innovation and Technology on 30 April 2026, found that only 15% of UK businesses reviewed the cyber security risks posed by their immediate suppliers, rising to 48% of large businesses. The same survey found that among businesses using, adopting or considering AI, around a quarter (24%) had cyber security practices in place to manage the risks of that AI. A voice agent sits exactly where those two gaps meet: a new supplier, handling live customer conversations, inside an AI system.

This guide is written for the CIO who owns that review. It sets out which standard question sets to start from, the voice-specific rows a generic pack leaves out, and how to read the answers. It deliberately does not repeat what sits elsewhere on this blog: how to read a SOC 2 report or an ISO 27001 certificate is covered in our voice AI vendor security certification guide, the sub-processor chain in our Article 28 sub-processor guide, and the structure of a full tender in our voice AI RFP template.

This guide is shipped by the team behind Dilr Voice, a multi-agent voice AI platform that ships per-country compliance rules and a full audit trail on every call. Or see DATS, our five-stage method for placing AI inside enterprise systems.

What is an AI voice agent security questionnaire?

An AI voice agent security questionnaire is the written set of security and data-handling questions a buyer sends a voice AI vendor before contract, covering hosting, encryption, separation between customers, supply chain, access control and incident alerting. It differs from a generic SaaS questionnaire because a voice agent processes live call audio, transcripts and caller identity data, and depends on a telephony carrier that most standard question sets never name.

In practice, the questionnaire is the document the CIO's team uses to decide whether a vendor can be trusted with a new category of data. For most enterprises that category is new. A chatbot handles typed text. A voice agent handles a recording of a person's voice, a transcript of what they said, the phone number they called from, and often the account details they read out to prove who they are. Each of those can be personal data, and using a voice to identify a caller raises further questions, which is why our guide to voice biometric data security treats identification as a separate risk.

The questionnaire is also where security meets the rest of procurement. Commercial teams score vendors, legal teams negotiate the data processing agreement, and the security review feeds both. If you are building the whole evaluation, the weighted vendor scorecard shows where security answers carry weight against cost, integration and delivery.

Why do generic security questionnaires miss voice AI risks?

Generic security questionnaires miss voice AI risks because they were written for applications that store records, not systems that listen. A standard pack asks about encryption, access control and data centres, but rarely asks where call audio is processed in real time, how long recordings and transcripts persist, which carrier routes the call, or whether conversation data feeds model improvement. The vendor can answer every row truthfully and still leave the voice-specific risks unexamined.

The wider supplier picture explains why this matters. The DSIT survey found that few UK organisations look hard at their suppliers at all, and that the share rises sharply with size. Large businesses are the most diligent group, and even there fewer than half formally review the cyber risks posed by immediate suppliers.

UK businesses reviewing cyber risk from immediate suppliers
12%Micro22%Small30%Medium48%Large15%All businesses
Share of UK businesses, by size, that formally reviewed the cyber security risks posed by their immediate suppliers, 2025/2026. Source: DSIT, Cyber Security Breaches Survey 2025/2026, Figure 3.13

The same survey reports that 11% of businesses required their suppliers to be certified with any standards or accreditations, rising to 41% of large businesses, up from 21% the year before. Large businesses, in other words, are increasingly asking suppliers for certificates. A certificate is useful, but it describes a scope the vendor chose. Whether that scope includes the real-time audio path, the transcription step and the carrier is exactly the kind of question the certification guide teaches you to check.

The wider supply chain is weaker still: the survey found only 6% of businesses reviewed it, and 24% of large ones. For a voice agent, the wider supply chain is not abstract. It is the carrier, the speech recognition step, the language model and the text-to-speech engine, any of which may be a separate company.

Which standard question sets should you start from?

UK buyers should start a voice AI security questionnaire from an established question set rather than writing one from scratch: the National Cyber Security Centre's 14 Cloud Security Principles, which apply to Software-as-a-Service as well as cloud platforms, and the Cloud Security Alliance's CAIQ, a set of yes or no questions mapped to its Cloud Controls Matrix. Both give vendors a familiar structure, then voice-specific rows are added on top.

The NCSC's cloud security principles are the better spine for a UK buyer, because each one is written as a set of goals. For each principle the NCSC describes the security goals a good service should meet, suggestions for how a provider could meet them, and the considerations the buyer still has to make. The guidance is explicit that a yes is not the end of the question: "You should also consider what evidence has been provided to give you enough confidence in the statements being made by the cloud provider."

The Cloud Security Alliance's STAR Level 1 Security Questionnaire (CAIQ v4) is the better tool when you want comparable answers across several vendors. It is a set of yes or no questions a customer can ask a cloud provider to establish how it meets the Cloud Controls Matrix. Providers can publish a completed self-assessment to the CSA's STAR registry at Level 1, which saves time, but a CAIQ is a general cloud question set covering IaaS, PaaS and SaaS, so it will not ask about call audio unless you add those rows.

A sensible approach is to use the NCSC principles as the headings, use CAIQ answers as supporting evidence where a vendor has them, and add the voice rows described below under the principles they belong to. That keeps the questionnaire short enough to be answered properly. Our enterprise voice AI vendor evaluation guide covers what the IT and security function typically asks in the first conversation; the questionnaire turns those questions into a written record.

What voice-specific data should the questionnaire cover?

A voice AI security questionnaire should cover four data types a generic pack overlooks: live call audio, stored recordings, transcripts, and derived data such as summaries, sentiment scores and logs. For each, the vendor should state where it is processed and stored, who can access it and how long it is kept, then answer one further question: whether any call data is used to train or improve a model.

The NCSC already points in this direction. Principle 2 on asset protection and resilience says a buyer should know where its data is and who can access it, and it extends that to derivatives of the data, naming verbose logs and machine learning models. For a voice agent, that sentence is the whole problem. The transcript is a derivative of the audio. The call summary is a derivative of the transcript. A model tuned on calls would be a derivative of all of them.

Turn that into five rows the vendor must answer in writing, one per data type plus the model question:

Data typeThe question to askWhat a good answer names
Live call audioWhere is audio processed in real time, and does any of it leave the hosting region?The hosting region and every processing step in the call path
RecordingsWho can replay a recording, and how is that access logged?Role-based access and an audit record of each playback
TranscriptsHow long are transcripts kept, and can the customer set the period?A configurable retention period and a deletion method
Derived dataAre summaries, scores and logs held to the same retention and access rules?The same controls as the source data, stated explicitly
Model useIs any customer call data used to train or improve a model, yours or a supplier's?A clear yes or no, with the contract clause that fixes it

The retention row deserves its own policy, not just an answer. How long to keep recordings and transcripts under UK GDPR is set out in our voice AI data retention guide, and the questionnaire answer should match the retention schedule you will actually run.

Separation between customers is the other row worth asking explicitly. NCSC Principle 3 asks a provider to explain how one customer is kept apart from another, and, where a SaaS runs on someone else's platform, which separation properties it inherits from the layer beneath. Ask whether dedicated tenancy is available and what it changes.

The same diagnostic logic underpins our AI operating model consulting, which covers governance, RACI and lifecycle, including who owns these answers once the platform is live.

How should the questionnaire treat the telephony carrier?

A voice AI security questionnaire should treat the telephony carrier as a named part of the supply chain, because every call passes through it before the voice agent hears a word. The vendor should state which carriers it supports, whether the carrier account belongs to the vendor or the customer, and what call data the carrier holds. That answer often changes who is responsible for the carrier relationship.

NCSC Principle 8 on supply chain security asks the buyer to understand how its data, and metadata derived from that data, is shared with third party suppliers and their supply chains. In a voice deployment, the carrier is the first of those suppliers. Common carriers in the market include Twilio, Telnyx and Vonage, and the answer to "whose account is it?" varies by vendor and by use case.

That ownership question matters. If the carrier account is the customer's own, the carrier relationship, its terms and its call records sit with the customer, and the vendor connects to it. If the account is the vendor's, the carrier becomes part of the vendor's chain and belongs on its sub-processor list. Either can be acceptable; what is not acceptable is not knowing. The authorisation and flow-down rules for that chain are covered in our sub-processor guide, and the question of who is controller and who is processor in our controller or processor guide.

Ask the same "whose and where" question of every other component in the call path: speech recognition, the language model, text-to-speech, and any analytics service that scores calls after they end. A vendor that can draw that path on one page, with each supplier named, is usually a vendor that has thought about it.

In a voice AI security questionnaire, recording consent, Telephone Preference Service screening and opt-out handling are architecture questions as well as legal ones. The TPS duty sits with the organisation making marketing calls, and the buyer needs to know where in the platform each control runs, whether it is on by default, whether it can be switched off, and what record the platform keeps to prove the control ran on each call.

The duty is clear. For live marketing calls, Ofcom states that the calling company must say who it is and give a contact number, and should not call a number registered with the Telephone Preference Service unless the person has agreed. The ICO's guidance on live marketing calls sets out when numbers on the TPS can be called. Both describe duties on the organisation making the calls, which in a typical voice AI deployment is the buyer rather than the platform vendor.

That is precisely why the questionnaire has to ask about the platform. A buyer cannot evidence compliance with a control it cannot see. The rows to add are short: where is the TPS and CTPS check performed, and against what list; how is an opt-out recognised mid-call and recorded; how is recording consent captured; and what does the audit trail show for each call. Our DNC logic guide explains the screening itself, and the wider legal picture is in our UK and EU voice AI compliance guide.

For context, Dilr Voice ships per-country compliance rules by default, covering recording consent, DNC registry checks, opt-out recognition and permitted calling hours, with a full audit trail on every call. Telephony runs on Twilio or Vobiz, and outbound campaigns require the customer's own Twilio or Vobiz account, which answers the carrier ownership question for outbound directly. The platform is hosted on Google Cloud Platform (GKE), encrypted at rest and in transit, with dedicated tenancy and regional data residency options for enterprise. Those are the facts a questionnaire should then test with evidence rather than accept, whichever vendor you are reviewing.

How do the answers map to UK GDPR Article 32?

The answers in a voice AI security questionnaire map to UK GDPR Article 32, which requires both the controller and the processor to put in place technical and organisational measures appropriate to the risk. The questionnaire is the buyer's record of how it assessed the vendor's measures. A well-built pack lets the CIO show, row by row, why the measures were judged appropriate for call audio and transcripts.

Article 32(1) lists the kinds of measure it has in mind: pseudonymisation and encryption, the ongoing confidentiality, integrity, availability and resilience of systems, the ability to restore access to data after an incident, and a process for regularly testing the measures. Article 32(2) adds that the assessment should account for the risks of accidental or unlawful destruction, loss, alteration or unauthorised disclosure. Each of those maps to a group of questionnaire rows, so the completed pack can form part of your Article 32 record.

Incident alerting is the row most often answered vaguely. NCSC Principle 13 says a provider should alert the customer promptly when it detects attacks against the customer's data or vulnerabilities in the customer's use of the service. Ask how, to whom, and how fast, and ask what audit information the customer can pull without opening a support ticket.

Building a voice AI security questionnaire
01Scope the dataAudio, transcripts, de…02Pick a questionsetNCSC principles, CAIQ03Add voice rowsCarrier, consent, mode…04Evidence eachanswerDocuments, not adjecti…
Start from a recognised question set, add the voice rows, then require evidence for each answer before scoring.

The organisational half of Article 32 is easy to forget in a technical pack. Ask who at the vendor can access customer call data, how that access is approved and reviewed, and how staff are vetted. The NCSC lists personnel security as Principle 6, and the question belongs in the questionnaire rather than in a follow-up call.

What is the best way to run a voice AI security review in 2026?

The best way to run a voice AI security review in 2026 is to send a short questionnaire built on the NCSC principles with voice rows added, require evidence for each answer, and score it against your own risk appetite rather than a vendor's certificate count. A large regulated buyer may still require a named certification as a gate, and for that buyer a vendor with a current, well-scoped report will move faster through review.

That concession matters. A vendor that already maintains a recent SOC 2 Type II report or a completed CAIQ can answer much of a generic pack from documents on the shelf. Whichever vendors are on your shortlist, whether that is PolyAI, Parloa, Cognigy, Vapi, Retell AI or Dilr Voice, ask each one which answers it can evidence from existing documents and which depend on how your own team configures the platform. Different platforms put the security work in different places, and the questionnaire should tell you where.

What the questionnaire cannot do is replace judgement. A buyer who requires a named certification should ask every vendor on the shortlist for the current report or certificate and check its scope, rather than accepting a logo. A buyer who does not require one should still insist on written answers to the voice rows above. If you are comparing platforms more broadly, our best AI voice agent 2026 comparison sets the field side by side, and the enterprise AI voice agents guide explains how the security review fits the full deployment.

Two process points make the review faster without making it weaker. First, send the questionnaire before the pilot, not after it: the pilot will put real call data into the platform, and the review should decide whether that is acceptable. Second, agree which answers become contract terms. A vendor's answer about model training or retention is only as strong as the clause that fixes it, which is why the security review and the data processing agreement should be read together.

If your team is still deciding where voice AI belongs before the questionnaire goes out, an AI placement diagnostic produces a ranked roadmap of where AI belongs and where it does not, and the AI execution office provides embedded delivery from pilot to production.

Can a vendor's SOC 2 report replace the questionnaire?

A vendor's SOC 2 report cannot fully replace a voice AI security questionnaire, because the report covers the controls and systems the vendor chose to put in scope, tested against criteria rather than your questions. It is strong evidence for the rows it covers. The questionnaire still needs to ask about the call audio path, the carrier, retention of transcripts and the use of call data for model training.

Read the report's system description first, then send only the rows it does not answer. That keeps the questionnaire short and the vendor's effort focused. How to read the report itself, including exceptions and bridge letters, is covered in the certification guide linked earlier in this article.

Who answers a voice AI security questionnaire, the vendor or the buyer?

Both the vendor and the buyer answer parts of a voice AI security questionnaire. The vendor answers for its hosting, separation, staff access, supply chain and incident alerting. The buyer answers for configuration it controls: who in its own team can access recordings, which retention period it sets, which compliance rules it enables, and, where the carrier account is the buyer's own, how that account is secured.

The NCSC draws the same line on its cloud security principles page: the principles help you choose a provider, and you will separately need to consider how you configure the service securely. Recording the buyer's own answers in the same document is what turns the questionnaire into a usable record of the deployment, rather than a file that only describes the vendor.

Want to see the controls in a live platform? Try Dilr Voice free, compare the field in our best AI voice agent guide, see our DATS methodology, or read about our approach to placing AI inside enterprise systems. More guides sit in the voice AI category.

Product
Dilr Voice
Service
AI Placement Diagnostic
Guide
Best AI Voice Agent 2026
Talk to the operators

Get the security review right before the pilot.

30-min scoping call · No deck · Confidential. We will tell you whether DATS fits, and where voice AI belongs in your operation.

Written by the Dilr.ai engineering team, practitioners who ship enterprise AI in production. Follow us on LinkedIn for shipping notes, or subscribe via the RSS feed.

ai voice agent security questionnairevoice ai vendor security pack ukvoice ai due diligence questionnairencsc cloud security principles voice aivoice ai security review ukvoice ai redditbest ai voice agent 2026dilr voice

Questions this article answers

What is an AI voice agent security questionnaire?

An AI voice agent security questionnaire is the written set of security and data-handling questions a buyer sends a voice AI vendor before contract, covering hosting, encryption, separation between customers, supply chain, access control and incident alerting. It differs from a generic SaaS questionnaire because a voice agent processes live call audio, transcripts and caller identity data, and depends on a telephony carrier that most standard question sets never name.

Why do generic security questionnaires miss voice AI risks?

Generic security questionnaires miss voice AI risks because they were written for applications that store records, not systems that listen. A standard pack asks about encryption, access control and data centres, but rarely asks where call audio is processed in real time, how long recordings and transcripts persist, which carrier routes the call, or whether conversation data feeds model improvement. The vendor can answer every row truthfully and still leave the voice-specific risks unexamined.

Which standard question sets should you start from?

UK buyers should start a voice AI security questionnaire from an established question set rather than writing one from scratch: the National Cyber Security Centre's 14 Cloud Security Principles, which apply to Software-as-a-Service as well as cloud platforms, and the Cloud Security Alliance's CAIQ, a set of yes or no questions mapped to its Cloud Controls Matrix. Both give vendors a familiar structure, then voice-specific rows are added on top.

What voice-specific data should the questionnaire cover?

A voice AI security questionnaire should cover four data types a generic pack overlooks: live call audio, stored recordings, transcripts, and derived data such as summaries, sentiment scores and logs. For each, the vendor should state where it is processed and stored, who can access it and how long it is kept, then answer one further question: whether any call data is used to train or improve a model.

How should the questionnaire treat the telephony carrier?

A voice AI security questionnaire should treat the telephony carrier as a named part of the supply chain, because every call passes through it before the voice agent hears a word. The vendor should state which carriers it supports, whether the carrier account belongs to the vendor or the customer, and what call data the carrier holds. That answer often changes who is responsible for the carrier relationship.

Where do recording consent and TPS checks sit in the platform?

In a voice AI security questionnaire, recording consent, Telephone Preference Service screening and opt-out handling are architecture questions as well as legal ones. The TPS duty sits with the organisation making marketing calls, and the buyer needs to know where in the platform each control runs, whether it is on by default, whether it can be switched off, and what record the platform keeps to prove the control ran on each call.

How do the answers map to UK GDPR Article 32?

The answers in a voice AI security questionnaire map to UK GDPR Article 32, which requires both the controller and the processor to put in place technical and organisational measures appropriate to the risk. The questionnaire is the buyer's record of how it assessed the vendor's measures. A well-built pack lets the CIO show, row by row, why the measures were judged appropriate for call audio and transcripts.

What is the best way to run a voice AI security review in 2026?

The best way to run a voice AI security review in 2026 is to send a short questionnaire built on the NCSC principles with voice rows added, require evidence for each answer, and score it against your own risk appetite rather than a vendor's certificate count. A large regulated buyer may still require a named certification as a gate, and for that buyer a vendor with a current, well-scoped report will move faster through review.

Dilr Voice

Put this into production

Dilr Voice runs AI voice agents for inbound and outbound calls: multi-agent handoff, RAG knowledge bases, and per-country compliance in one platform.

Related articles

← Previous
Khanmigo Alternative UK: What Actually Changes

One email, once a month. No hype. Just what we learned shipping.