Reusing enterprise call recordings to train a voice AI model is further processing under UK GDPR, governed by the new Article 8A purpose limitation test since February 2026. Dilr Voice explains the compatibility assessment, when you need a fresh lawful basis, and why anonymisation rarely solves it.
DE
Dilr.ai EngineeringEngineering team
Published Jul 28, 2026Updated Jul 28, 2026Read 11 min
Every enterprise voice deployment quietly builds an asset the business rarely planned for: a growing archive of call recordings. They were captured for quality assurance, dispute resolution, complaint handling, or a regulatory record-keeping duty. Then a data science team looks at that archive and asks the obvious question. Could we use these to fine-tune the model, or at least to evaluate it? The commercial logic is strong, and the pressure is rising. In its November 2025 State of AI survey, McKinsey found that 88% of enterprises now use AI, yet only 33% have taken it into production and just 6% capture material value. Reusing your own conversational data is one of the cheapest ways to close that gap.
The problem is that the data was not collected for training. Reusing it is a further purpose, and under UK law a further purpose is not automatically permitted. The Data (Use and Access) Act 2025 rewrote exactly this part of the rulebook, and the new provisions came into force on 5 February 2026. This guide sets out what purpose limitation now requires, how the new compatibility test applies to call audio, when you need a fresh lawful basis, and where anonymisation genuinely helps versus where it only appears to.
This guide is shipped by the team behind Dilr Voice, enterprise voice AI built for regulated deployments. Or see DATS, our five-stage AI consulting system, for help placing the governance around it.
Why the pressure to reuse call data keeps risingShare of enterprises reaching each stage of AI value capture, 2025 to 2026. Source: McKinsey, The State of AI (Nov 2025)
Can you reuse call recordings to train a voice AI model?
Not automatically. Call recordings gathered for quality assurance or dispute resolution are personal data, and using them to train or fine-tune a voice AI model is a new purpose. UK GDPR treats that as further processing, which must either be compatible with the purpose you collected the data for or rest on a fresh lawful basis. Dilr Voice treats this as a governance decision made before any data leaves the recording store, not an afterthought.
The instinct in most engineering teams is that first-party data is fair game. You collected it, you hold it, so surely you can improve your own product with it. That instinct is not wrong so much as incomplete. Owning the data does not extend the purpose it was collected for. A recording captured to resolve a billing dispute was not captured to teach a language model how customers phrase refund requests, even if the second use feels like a natural extension of the first. The law asks you to test that intuition, not to trust it. Our AI placement diagnostic exists precisely to surface these decisions before a deployment commits to them.
What does purpose limitation require after the Data (Use and Access) Act 2025?
Purpose limitation is the principle in UK GDPR Article 5(1)(b): personal data collected for specified, explicit and legitimate purposes must not be further processed in a way incompatible with those purposes. The Data (Use and Access) Act 2025 did not weaken that principle. Instead its section 71 and Schedule 5 replaced the old Article 6(4) with a new Article 8A, "Purpose limitation: further processing", which restates the compatibility factors and adds deemed-compatible purposes in a new Annex 2.
This is the single most important update for anyone planning to reuse recordings in 2026. If your compliance notes still cite Article 6(4), they are describing a provision that no longer exists. Article 8A is now the operative text, in force since 5 February 2026 under the DUAA commencement regulations. The principle is stable, the mechanics have moved, and the new section 71 is where the compatibility analysis now lives.
The regulator has been direct about how this applies to model building. The Information Commissioner's Office, in its response to the consultation series on generative AI, framed the duty in one sentence:
"developers who are reusing personal data for training generative AI must consider whether the purpose of training a model is compatible with their original purpose of collecting that data."
The ICO calls that a compatibility assessment, and it is now the assessment Article 8A demands. The good news for enterprises is that this is easier ground than the web-scraping cases the ICO was mostly examining. You have a direct relationship with the people on the calls, which the ICO itself treats as a factor that makes compatibility more likely.
How do you run the Article 8A compatibility assessment on call audio?
Article 8A(2) sets five factors for deciding whether the new purpose is compatible with the original. Applied to call audio: any link between the collection purpose and training; the context of collection; the nature of the processing, including whether the audio holds special category or criminal-offence data; the possible consequences for the caller; and appropriate safeguards such as encryption or pseudonymisation. Document each factor with evidence, not assertion.
The context factor usually does the heaviest lifting. A caller who left a message with your contact centre reasonably expects that call to be used to serve them, and perhaps to improve the service they receive. Training a narrowly scoped model to understand refund requests sits closer to that expectation than harvesting the same audio to build a general voice clone. The consequences factor is the counterweight: if the recording could feed a model that later makes or influences decisions about that person, the assessment tightens. This is where the design work of a governed AI operating model pays for itself, because the safeguards you can point to are what tips a marginal case toward compatible.
The purpose limitation decision path for call-recording reuseEach gate follows UK GDPR Article 8A as amended by the Data (Use and Access) Act 2025.
Before you run the five-factor test, check the shortcuts. Article 8A(3) deems some further processing compatible without a fresh analysis: where the caller has consented to the new purpose, where the processing is genuine scientific research, archiving in the public interest or statistical work, or where it meets an Annex 2 condition such as detecting crime or meeting a legal obligation. Commercial model training almost never fits these. Building a deployable product is not scientific research, and improving your own service is not a statistical purpose. Most enterprises will fall through to the Article 8A(2) assessment and should plan on that basis.
When do you need a fresh lawful basis to train on recordings?
You need a fresh lawful basis whenever the compatibility assessment fails, and sometimes even when it succeeds. Article 8A(4) is explicit: meeting a compatibility condition does not rescue a lawful basis that is no longer valid for the new purpose. If the original recording rested on consent, that consent almost certainly did not cover training, so you cannot lean on it. You then pick a fresh basis, usually legitimate interests, and evidence it before any training run begins.
Legitimate interests is the realistic route for most first-party training, but it is not a free pass. It requires a documented three-part balancing test weighing your interest, the necessity of the processing, and the caller's rights and expectations. We treat that as its own workstream, and the mechanics are covered in our guide to the legitimate interest balancing test. Consent is cleaner in principle but hard to operate at scale, because you would need specific, informed consent to training from every caller whose audio enters the set. Whichever basis you pick, the caller must be told, which means your call recording consent scripts and privacy information have to name model improvement as a purpose rather than bury it.
The same discipline runs through our AI execution office, where lawful basis is fixed and evidenced before a single training job is scheduled, not reconstructed afterwards for an audit.
There is a documentation layer that many teams skip and regret. A new purpose belongs in your record of processing activities, and a training use of recordings that could affect people will usually need a data protection impact assessment. Neither is bureaucracy for its own sake. They are the artefacts an ICO audit will ask for, and the ones that let you show the compatibility assessment was done, not merely claimed.
Does anonymising or pseudonymising the recordings solve the problem?
Pseudonymisation does not take the recordings outside data protection law. Under UK GDPR Article 4(5), pseudonymised data that can still be linked to a person with additional information remains personal data, so masking names in a transcript or swapping in reference numbers lowers risk without ending the obligation. It is a genuine safeguard, and one of the factors Article 8A(2) rewards, but it is a mitigation, not an exit. Purpose limitation still applies to pseudonymised call data in full.
True anonymisation is a higher bar, and voice makes it harder still. Call audio carries the caller's voice, which is close to biometric, so stripping a name rarely renders a recording genuinely anonymous. Even where you anonymise the training set, a subtle trap remains: a model that memorises its training data can itself amount to personal data. The European Data Protection Board's December 2024 opinion on AI models took the position that a model is not automatically anonymous, and that where memorisation can occur the model itself may fall within the rules. The ICO's anonymisation and pseudonymisation guidance sets a demanding standard for claiming data is anonymous, and the ICO has flagged that this guidance is itself under review because of the DUAA.
This is why the timing of your lawful basis matters so much. If a caller later exercises their right to erasure, deleting the source recording is straightforward, but the model trained on it cannot be un-trained; erasure cascades to the training set, not to the weights already learned. We unpack that limit in our guide to the right to erasure for call data. The un-eraseable weight is the strongest argument for getting purpose and basis right before training, rather than trusting anonymisation to rescue a reuse you were never entitled to make.
What is the best way to structure compliant call-recording reuse for AI in 2026?
The best structure is a documented pipeline that fixes purpose and lawful basis at ingestion, not at training. In practice: name a single specific purpose for the model, run the Article 8A(2) assessment, record the outcome, select and evidence a lawful basis, minimise and pseudonymise the data before it enters the training set, and update your privacy information. Dilr Voice is built around that sequence, because retrofitting it after a model ships costs far more than designing it in.
No single tool makes this decision for you, and honesty matters here. If your reuse is genuinely low-stakes, for example evaluating latency on a small set of already de-identified recordings for internal reporting, a lighter platform such as Vapi, Synthflow or Retell AI with a built-in evaluation dashboard may be all you need, and reaching for a full governance programme would be overkill. Where the reuse touches customer-affecting decisions, regulated records under FCA Consumer Duty expectations, or special category content, the platform that wins is the one that lets you evidence the compatibility assessment, wire lawful basis into the pipeline, and integrate with your systems of record in Salesforce, HubSpot or Twilio without copying raw audio across boundaries. Bland AI, PolyAI and ElevenLabs each solve pieces of the voice problem well; the differentiator for regulated reuse is the governance layer around them, which is what our DATS methodology and a governed operating model are designed to supply.
Is using recordings to evaluate a model different from training on them?
Yes, and the distinction matters. Evaluating a model on held-out recordings is still a further purpose under UK GDPR, so it still needs a compatible purpose or a lawful basis, but the ICO views narrower, more explicit purposes more favourably. Using a small, well-scoped evaluation set involves less processing than open-ended training and is easier to justify under Article 8A. The purpose is still separate from why the calls were recorded, so document it rather than assume evaluation is exempt.
Can legitimate interests be the lawful basis for training on call recordings?
Often, yes. Legitimate interests is the most workable basis for first-party training on your own call recordings, provided you complete and record a three-part balancing test and can show the caller's rights do not override your interest. It is not automatic. Where the audio contains special category data or feeds decisions that significantly affect people, the balance shifts and a different basis or an Article 9 condition may be required. Evidence the assessment before training, not after.
Does the EU AI Act affect reusing call recordings for training?
Indirectly, yes. The EU AI Act's Article 50 transparency duty governs telling people they are dealing with an AI system, which shapes how you collect and disclose recordings. It does not set purpose limitation, which remains UK GDPR for UK callers, but firms operating across both regimes should align disclosure, consent and reuse so a recording is lawful under each. Our guide to voice AI compliance across the UK and EU and our wider compliance library cover that overlap.
Talk to the operators
Reuse your call data without reopening it later.
30-min scoping call · No deck · Confidential. We will map the compatibility assessment, lawful basis and safeguards before your first training run.
Written by the Dilr.ai engineering team, practitioners who ship enterprise AI in production. Follow us on LinkedIn for shipping notes, or subscribe via the RSS feed. This article is general information, not legal advice; take advice on your own processing.
voice AI call recordings model training purpose limitationtraining AI on call recordings GDPRcall recording secondary use UK GDPRvoice AI compliance redditbest voice AI for call data compliance 2026voice AI data protection complianceDilr Voice
Questions this article answers
Can you reuse call recordings to train a voice AI model?
Not automatically. Call recordings gathered for quality assurance or dispute resolution are personal data, and using them to train or fine-tune a voice AI model is a new purpose. UK GDPR treats that as further processing, which must either be compatible with the purpose you collected the data for or rest on a fresh lawful basis. Dilr Voice treats this as a governance decision made before any data leaves the recording store, not an afterthought.
What does purpose limitation require after the Data (Use and Access) Act 2025?
Purpose limitation is the principle in UK GDPR Article 5(1)(b): personal data collected for specified, explicit and legitimate purposes must not be further processed in a way incompatible with those purposes. The Data (Use and Access) Act 2025 did not weaken that principle. Instead its section 71 and Schedule 5 replaced the old Article 6(4) with a new Article 8A, "Purpose limitation: further processing", which restates the compatibility factors and adds deemed-compatible purposes in a new Annex 2.
How do you run the Article 8A compatibility assessment on call audio?
Article 8A(2) sets five factors for deciding whether the new purpose is compatible with the original. Applied to call audio: any link between the collection purpose and training; the context of collection; the nature of the processing, including whether the audio holds special category or criminal-offence data; the possible consequences for the caller; and appropriate safeguards such as encryption or pseudonymisation. Document each factor with evidence, not assertion.
When do you need a fresh lawful basis to train on recordings?
You need a fresh lawful basis whenever the compatibility assessment fails, and sometimes even when it succeeds. Article 8A(4) is explicit: meeting a compatibility condition does not rescue a lawful basis that is no longer valid for the new purpose. If the original recording rested on consent, that consent almost certainly did not cover training, so you cannot lean on it. You then pick a fresh basis, usually legitimate interests, and evidence it before any training run begins.
Does anonymising or pseudonymising the recordings solve the problem?
Pseudonymisation does not take the recordings outside data protection law. Under UK GDPR Article 4(5), pseudonymised data that can still be linked to a person with additional information remains personal data, so masking names in a transcript or swapping in reference numbers lowers risk without ending the obligation. It is a genuine safeguard, and one of the factors Article 8A(2) rewards, but it is a mitigation, not an exit. Purpose limitation still applies to pseudonymised call data in full.
What is the best way to structure compliant call-recording reuse for AI in 2026?
The best structure is a documented pipeline that fixes purpose and lawful basis at ingestion, not at training. In practice: name a single specific purpose for the model, run the Article 8A(2) assessment, record the outcome, select and evidence a lawful basis, minimise and pseudonymise the data before it enters the training set, and update your privacy information. Dilr Voice is built around that sequence, because retrofitting it after a model ships costs far more than designing it in.
Is using recordings to evaluate a model different from training on them?
Yes, and the distinction matters. Evaluating a model on held-out recordings is still a further purpose under UK GDPR, so it still needs a compatible purpose or a lawful basis, but the ICO views narrower, more explicit purposes more favourably. Using a small, well-scoped evaluation set involves less processing than open-ended training and is easier to justify under Article 8A. The purpose is still separate from why the calls were recorded, so document it rather than assume evaluation is exempt.
Can legitimate interests be the lawful basis for training on call recordings?
Often, yes. Legitimate interests is the most workable basis for first-party training on your own call recordings, provided you complete and record a three-part balancing test and can show the caller's rights do not override your interest. It is not automatic. Where the audio contains special category data or feeds decisions that significantly affect people, the balance shifts and a different basis or an Article 9 condition may be required. Evidence the assessment before training, not after.
DE
Dilr.ai Engineering
Engineering team
Compliance
Deploy voice AI without failing an audit
Dilr Voice ships per-country TCPA and GDPR rules, and the UK AI compliance changelog tracks ICO, FCA, and EU AI Act changes as they land.