KYC Onboarding Document Extraction AI: A UK Bank Guide
In short
Dilr Mira is a class of private clinical small language models from DILR.AI that turn clinical documents into schema-valid JSON, with KYC listed only as a research direction. This guide explains what a KYC extraction model may draft for a UK bank, why verification and the onboarding decision stay with the bank, and what to test before relying on it.
DE
Dilr.ai EngineeringEngineering team
Published Oct 9, 2026Read 17 min
A corporate onboarding file at a UK bank can run to a stack of scanned documents: a certificate of incorporation, articles of association, a register of directors, a structure chart, passports and driving licences for the people behind the company, a utility bill or bank statement for an address, and sometimes a source of funds letter. Someone has to read every one of them and key the same names, numbers and dates into the onboarding system. That reading is the part of KYC that looks most like a document extraction problem, which makes it a natural first candidate when a Chief Data and AI Officer is asked where AI could help.
The pressure is real. Fenergo's research put the average cost of a corporate KYC review at a British bank at $2,613 in 2023, up 19% on the year, and found that 39% of banks operating in the UK had lost clients to slow or inefficient onboarding. Where that cost actually sits, and how to diagnose it before buying anything, is the subject of our corporate KYC review cost guide, and this post does not repeat it. Nor does it map every DILR line onto banking, which the AI in UK banking map already does.
This guide holds a narrower question: what a document extraction model may do with a KYC pack under the Money Laundering Regulations 2017, where extraction stops and verification begins, and what such a model would have to prove before a bank relied on it. It is written by the team behind a clinical extraction model whose own product page lists KYC as research, not as a shipped capability, so it is also honest about that status. The measurement method, the verifier workflow and the case for keeping documents inside the estate are set out in our clinical document extraction guide and our health insurance claims extraction guide, and are only summarised here.
This guide is shipped by the team behind Dilr Mira, a class of private clinical small language models that turn lab reports, scans and claim forms into source-grounded, schema-valid JSON on the customer's own hardware. Or see DATS, the five-stage AI consulting system that places models like it inside regulated operations.
What is KYC onboarding document extraction AI?
KYC onboarding document extraction AI is a model that reads the documents a bank collects for customer due diligence, such as incorporation certificates, registers, identity documents and proof of address, and turns them into structured fields that an onboarding system can store. It drafts data from a page. It does not decide whether the document is genuine, whether the person is who they claim to be, or whether the bank should take the customer on.
That boundary matters because the label blurs it. Products sold as "KYC automation" can bundle three different jobs: reading a document, checking that the document and the person are real, and scoring the risk of the relationship. Only the first of those is extraction. A model that is good at the first job tells you nothing about the other two, and a bank that buys it as though it did the third has bought a control gap.
For a Chief Data and AI Officer, the useful framing is the one the bank's model risk function will apply anyway. An extraction model is a component that produces a draft record, and the question is how accurate that draft is, field by field, on the bank's own documents, and what happens to the fields it gets wrong. Our AI operating model work starts from the same place: name the component, name its owner, and name the evidence it leaves.
What do the Money Laundering Regulations require of a KYC document pack?
The Money Laundering Regulations 2017 require a bank, as a relevant person, to identify a corporate customer and verify its name, company number and registered office address, to take reasonable measures to verify its constitution, directors and senior persons, and to identify and take reasonable measures to verify its beneficial owners. Every one of those duties binds the bank, not any tool the bank uses to read the documents.
Regulation 28 of the Money Laundering Regulations 2017 sets those duties out in order. For a body corporate, paragraph (3) requires the bank to obtain and verify the name, the company or registration number and the registered office address, plus the principal place of business where it differs, and to take reasonable measures to determine and verify the law the company is subject to, its constitution and the full names of its board and senior persons. Paragraph (4) covers beneficial owners. Paragraph (10) adds a further duty where someone acts on the customer's behalf: the bank must verify that the person is authorised, identify them and verify their identity.
The regulation also defines the word that matters most for this guide. Under regulation 28(18)(a), to verify means to verify "on the basis of documents or information in either case obtained from a reliable source which is independent of the person whose identity is being verified". Read that against what an extraction model does. A model reading a passport scan the customer uploaded is reading a document the customer supplied. It can capture what the document says. It cannot, on its own, make the document an independent and reliable source.
Two more provisions shape how extraction fits. Regulation 28(9) says a bank does not satisfy its beneficial ownership duty by relying solely on information delivered to the registrar, which is why our corporate KYC guide treats the Companies House register as one input rather than an answer. And regulation 28(16) requires the bank to be able to demonstrate to its supervisor that the extent of its measures is appropriate to the risk. An extraction step is part of those measures, so its accuracy becomes something the bank may have to evidence.
The FCA's own guide restates the shape of the duty. FCG 3.2.4 in the Financial Crime Guide says firms must identify their customers and, where applicable, their beneficial owners, and then verify their identities. Identification and verification are two steps in that sentence, and extraction only ever helps with the first.
Why is extracting a field not the same as verifying it?
Extracting a field means reading what a document says; verifying it means establishing, from a reliable source independent of the customer, that what it says is true. A model can read a passport number perfectly from a forged passport. Authenticity checks, identity matching and the bank's own risk decision are separate steps, and an extraction model performs none of them.
At least one vendor that sells extraction draws this line itself. Google Cloud's Document AI processor list carries extraction parsers such as a US Driver License Parser, which pulls names, document numbers and dates of birth, and a separate Identity Document Proofing processor that predicts whether an ID document is valid using fraud signals. Two products for two jobs is the honest architecture, and it is the one a bank should expect from any supplier.
UK law points the same way. Regulation 28(19) of the Money Laundering Regulations allows information to be regarded as coming from a reliable independent source where it is obtained through an electronic identification process that is secure from fraud and misuse. That is a statement about identity verification services, not about reading text. The UK government's digital verification services trust framework, which the government describes as the set of rules and standards that show what a good digital identity looks like, is the regime under which such services can be independently certified. An extraction model is not one of those services and should never be described as one.
So the honest place for extraction in a KYC flow is narrow and still valuable. It turns the documents a bank has already collected into a draft record, without anyone retyping them. Verification of the document and the person runs alongside it, through specialist tooling and the bank's own checks. The decision to onboard stays with an analyst and, for higher risk cases, with whoever the bank's own policy escalates to. Our AI placement diagnostic is built to find exactly this kind of seam, where a model can take the reading off people without taking any decision with it.
Where extraction sits in a KYC onboarding flowExtraction drafts the record; verification and the onboarding decision stay with independent checks and people.
Which fields in a KYC pack can a model only draft?
A KYC extraction model can draft every field it reads, but it can only ever draft them: company name and number, registered office, directors' names, beneficial owners from a structure chart, and personal details from identity documents. Whether each drafted field is accurate, independent and sufficient is a judgement the bank's analyst makes, and some fields, such as ownership percentages from a chart, need more care than others.
The table below sets each common field against the document it usually comes from, what extraction can contribute, and who confirms it. The right-hand column is the important one: nothing in it is the model.
Field
Usual source document
What extraction contributes
Who confirms it
Company name and number
Certificate of incorporation, register extract
Drafts both fields and links them to the page
Analyst, against an independent source
Registered office and trading address
Certificate, utility bill, lease
Drafts the address as written on each document
Analyst
Constitution and governing law
Articles of association
Locates the governing law clause and the document date
Analyst, under reg 28(3)(b) reasonable measures
Directors and senior persons
Register of directors, board minutes
Drafts the list of names and roles
Analyst
Beneficial owners and percentages
Structure chart, share register
Drafts names and stated holdings as written
Analyst, never the register alone (reg 28(9))
Personal identity details
Passport, driving licence
Drafts name, date of birth, document number, expiry
Identity verification step, then analyst
Authority to act
Board resolution, mandate letter
Drafts the named signatory and the scope stated
Analyst, under reg 28(10)
Source of funds narrative
Letter, accounts
Drafts a summary of what the letter claims
Analyst, escalated under the bank's own policy
Two points follow from the table. The first is that structured documents, such as a certificate of incorporation, put fixed fields in predictable places and should be easier to extract reliably than narrative documents, such as a source of funds letter, where a model is at greater risk of summarising instead of extracting. The second is that every row ends with a person. That is not caution for its own sake. It follows from the regulation, which puts the duty to verify on the bank.
The same diagnostic logic underpins our AI execution office, where a senior team embeds with a client's own staff to run AI in production rather than handing over a model and leaving.
Does Dilr Mira read KYC documents today?
No. Dilr Mira reads clinical documents today. The Dilr Mira models page lists KYC and onboarding as a research direction for the same verifier-gated extraction method, alongside insurance claims, prior authorisation, invoices, lab networks and legal intake, and states plainly that clinical extraction is what is shipped. Mira-Q2 was trained on clinical document types, not on passports or company registers.
That status is stated on the Dilr Mira models page and repeated in the honest line our banking map draws, where Dilr Mira is listed as not applying to the banking build today. The models page says Mira-Q2 is strongest on lab reports, medication lists, discharge summaries, pathology reports, intake forms and progress notes, in English, and that broader document types and languages are in development for Mira-3. None of those is a KYC document.
It is worth being precise about one architectural point, because it is easy to over-read. Mira-Q2 takes the extraction schema as an input, so a customer supplies the fields they want as a schema file. That design is what makes KYC a plausible research direction. It does not mean a bank could point today's model at a KYC schema and expect clinical-grade results, because the model learned its vocabulary and document layouts from clinical material. The published evaluation on the open model card on Hugging Face is a clinical evaluation, and every figure in it should be read that way.
What does carry across is the method. Every Mira-Q2 output is a draft for human review, gated by a deterministic verifier that checks schema validity, source grounding and identifier leakage, and every field is grounded in the source text it came from. A KYC model built on the same method would inherit that shape: drafted fields, each one grounded in its source, nothing autonomous. For a full account of how the class is built, see what Dilr Mira is.
What would a KYC extraction model have to prove before a bank relied on it?
A KYC extraction model would have to prove field-level accuracy on the bank's own document mix, not on a vendor benchmark, with results reported separately for structured documents and narrative ones. It would also have to show that every drafted field traces to its source page, that it leaks no identifiers outside the perimeter, and that its records survive the bank's five-year retention duty.
The strongest argument for testing on the bank's own files comes from Mira-Q2's own published results. The model card reports 100% JSON validity on training-distribution data but 86% on general real physician prose from the MTSamples set, and says the model struggles with document types it never saw, such as operative notes. That gap is published rather than hidden, and the lesson travels: a model's score on documents like its training data is not its score on yours. A bank should not assume that a vendor's KYC demo set predicts results on its own archive of scanned registers, foreign passports and handwritten letters, and should measure the difference before relying on any number. Our Mira-Q2 evaluation scorecard shows how the clinical gap was measured, set by set.
A sensible evaluation plan for a KYC extraction pilot has four parts.
A gold set from the bank's own archive. A sample of closed onboarding files, with the correct values for each field keyed independently by two analysts, so accuracy can be scored per field and per document type.
Separate scores for structured and narrative documents. A certificate of incorporation and a source of funds letter are different problems, and a blended average hides the one that fails.
Grounding and leakage checks. Every drafted field should be grounded in the source text it came from, and no personal identifier should appear anywhere it was not on the page. Mira-Q2's clinical evaluation recorded zero identifier leaks across 782 documents, as our evaluation scorecard sets out; a KYC model would need to show the same property on KYC material, not borrow the clinical result.
A failure route. Every field the model cannot read, or reads with low confidence, should land in front of a person as a gap, not as a guess.
Retention belongs in the plan from the start. Regulation 40 of the Money Laundering Regulations requires a bank to keep a copy of the documents and information it obtained for customer due diligence for five years from the end of the business relationship, and then to delete the personal data unless an exception applies. Once a model drafts fields from those documents, the drafts and the reviewer's corrections sit alongside that record, and in our view a bank should decide deliberately to retain and delete them on the same clock. Designing the audit trail before the pilot is cheaper than reconstructing it for a supervisor afterwards, which is the work our AI operating model engagements set up.
Why does it matter where a KYC extraction model runs?
Where a KYC extraction model runs decides who else sees the documents. A KYC pack holds passports, home addresses and ownership details for named individuals, so a model that sends pages to an external API adds a processor, a transfer question and a third-party dependency to the bank's risk register, while a model that runs inside the bank's own estate does not.
The case for on-premise extraction is made at length in our clinical document extraction guide, and the reasoning carries across sectors: the fewer copies of a sensitive document that exist outside the perimeter, the fewer places it can leak from. In the clinical setting, Mira-Q2 is built for that case. It is open under Apache-2.0, runs on CPU and takes about 2 GB on disk, with a GPU used only for speed, so the whole model fits on hardware the bank already controls.
Governance is the other half. The Bank of England and FCA survey of AI in UK financial services, published in November 2024, found that 75% of firms already use AI but only 34% report a complete understanding of the AI technologies they use, and attributed much of the gap to third-party models. An open model whose weights, evaluation files and scorecard are public is easier to understand than a closed service, which is one reason a model risk function may prefer it even before cost enters the discussion. Keeping an inventory of every model in the onboarding flow is the starting point, as our AI tool inventory guide explains.
What is the best approach to KYC document extraction in 2026?
The best approach to KYC document extraction in 2026 depends on whether documents may leave the bank's estate. If they may, a shipped cloud extraction service paired with a certified identity verification provider is available today. If they may not, an on-premise open model evaluated on the bank's own files is the safer route, accepting that KYC-specific open models are still early.
Five criteria separate a credible approach from a demo:
Extraction and verification are separate components, each with its own owner and its own evidence.
Field-level accuracy on the bank's own files, reported by document type, before any production traffic.
Source grounding, so every drafted field can be traced to the source text it came from.
A perimeter decision made deliberately, with the processor and transfer questions answered in writing.
A retention design that keeps documents, drafts and reviewer actions together for the regulation 40 period.
On the shipped cloud side, several products already cover the identity document part of the pack. Amazon Textract Analyze ID extracts fields such as date of birth and date of expiry from identity documents including U.S. passports and driver's licences. The Azure AI Document Intelligence ID model returns key values from worldwide passports and U.S. driver's licences as structured JSON. Google Document AI, as noted above, pairs extraction parsers with a separate proofing processor. Fenergo, which describes itself as a KYC and client lifecycle management provider, sits at a different layer again: the system that runs the onboarding case.
The honest concession is that today a shipped product wins outright on passports. Dilr Mira does not read KYC documents, Textract and Azure already read passports on the pages cited, with Azure stating worldwide passport coverage, and support for other UK identity documents such as driving licences should be checked vendor by vendor. A bank that is content to send documents to a hyperscaler under its existing outsourcing controls can start with one of them now. None of them, on the pages cited here, claims to read a full corporate pack of articles, registers and structure charts, so that part still needs testing whichever route a bank takes. The on-premise route earns its place only where the perimeter requirement is firm, and even then the bank should treat any KYC-specific open model as something to evaluate, not something to trust on arrival. Our enterprise AI consulting guide covers how to run that choice as a scoped decision rather than a vendor shortlist, and you can book a scoping call if you would rather talk it through with the people who build the models.
Does Mira-Q2's published evaluation say anything about KYC accuracy?
Mira-Q2's published evaluation says nothing direct about KYC accuracy. Its 782 documents are all clinical, spread across four test sets, and its scores describe how well a clinical model reads clinical text. What it does show is a method worth copying for KYC: field-level scoring, published confidence intervals and an honest gap between familiar and unfamiliar document types.
Any vendor that offers a clinical, invoice or generic document benchmark as evidence of KYC performance is making the same transfer error. Ask for results on your own closed files, scored field by field, before any production traffic; an embedded delivery team working inside the bank is one way to run that test.
Does document extraction change how long KYC records must be kept?
Document extraction does not change the retention period: under regulation 40 of the Money Laundering Regulations 2017, a bank keeps customer due diligence documents and information for five years from the end of the business relationship. What extraction changes is what sits alongside that record: the drafted fields, their source text and the reviewer's corrections, which a bank should plan to keep and delete on the same clock.
Planning for that in the data model means deletion at the end of the period removes the drafts as well as the scans. The compliance changelog tracks regulatory changes that could move these rules, and the industries category collects our other sector guides.
Written by the Dilr.ai engineering team, practitioners who ship enterprise AI in production. Follow us on LinkedIn for shipping notes, or subscribe via the RSS feed.
KYC onboarding document extraction AI is a model that reads the documents a bank collects for customer due diligence, such as incorporation certificates, registers, identity documents and proof of address, and turns them into structured fields that an onboarding system can store. It drafts data from a page. It does not decide whether the document is genuine, whether the person is who they claim to be, or whether the bank should take the customer on.
What do the Money Laundering Regulations require of a KYC document pack?
The Money Laundering Regulations 2017 require a bank, as a relevant person, to identify a corporate customer and verify its name, company number and registered office address, to take reasonable measures to verify its constitution, directors and senior persons, and to identify and take reasonable measures to verify its beneficial owners. Every one of those duties binds the bank, not any tool the bank uses to read the documents.
Why is extracting a field not the same as verifying it?
Extracting a field means reading what a document says; verifying it means establishing, from a reliable source independent of the customer, that what it says is true. A model can read a passport number perfectly from a forged passport. Authenticity checks, identity matching and the bank's own risk decision are separate steps, and an extraction model performs none of them.
Which fields in a KYC pack can a model only draft?
A KYC extraction model can draft every field it reads, but it can only ever draft them: company name and number, registered office, directors' names, beneficial owners from a structure chart, and personal details from identity documents. Whether each drafted field is accurate, independent and sufficient is a judgement the bank's analyst makes, and some fields, such as ownership percentages from a chart, need more care than others.
Does Dilr Mira read KYC documents today?
No. Dilr Mira reads clinical documents today. The Dilr Mira models page lists KYC and onboarding as a research direction for the same verifier-gated extraction method, alongside insurance claims, prior authorisation, invoices, lab networks and legal intake, and states plainly that clinical extraction is what is shipped. Mira-Q2 was trained on clinical document types, not on passports or company registers.
What would a KYC extraction model have to prove before a bank relied on it?
A KYC extraction model would have to prove field-level accuracy on the bank's own document mix, not on a vendor benchmark, with results reported separately for structured documents and narrative ones. It would also have to show that every drafted field traces to its source page, that it leaks no identifiers outside the perimeter, and that its records survive the bank's five-year retention duty.
Why does it matter where a KYC extraction model runs?
Where a KYC extraction model runs decides who else sees the documents. A KYC pack holds passports, home addresses and ownership details for named individuals, so a model that sends pages to an external API adds a processor, a transfer question and a third-party dependency to the bank's risk register, while a model that runs inside the bank's own estate does not.
What is the best approach to KYC document extraction in 2026?
The best approach to KYC document extraction in 2026 depends on whether documents may leave the bank's estate. If they may, a shipped cloud extraction service paired with a certified identity verification provider is available today. If they may not, an on-premise open model evaluated on the bank's own files is the safer route, accepting that KYC-specific open models are still early.
DE
Dilr.ai Engineering
Engineering team
Dilr Voice
Voice AI built for your sector
Dilr Voice answers and places calls 24/7 with compliance rules for regulated industries, from clinics and estate agents to financial services.