Industries

Health Insurance Claims Document Extraction: On-Prem Guide

Dilr Mira is a class of private clinical small language models from DILR.AI that turn clinical documents such as lab reports and discharge summaries into schema-valid JSON on your own hardware. This guide shows how a UK insurer can test that on claims files, what remains research, and how to measure accuracy first.

Health Insurance Claims Document Extraction: On-Prem Guide DILR MIRA · INSURANCE Health Insurance Claims Document Extraction: On-Prem Guide 01 Claims file in 02 Clinical pages extracted 03 Verifier gate 04 Handler reviews dilr.ai/blog

A claims director in UK health and protection insurance runs a document business that happens to pay money. A typical income protection or critical illness claim arrives as a bundle: a claim form, a GP report, a consultant letter, discharge notes, test results, and a covering letter that refers to all of them. Before anyone can assess the claim, someone has to read those pages and pull out the diagnosis, the dates, the treating clinician and the functional limits. The volume is not small. Insurers paid out £7.84 billion in individual and group protection claims, a category that includes life cover as well as income protection and critical illness, across 2025, according to figures from the ABI and Group Risk Development published in June 2026, and 258,000 new individual claims were paid in the year.

The obvious answer is to send the medical pages to a general-purpose AI service and let it fill in the claim record. The obvious objection lands immediately: those pages are health data, and the insurer, as controller, answers for where they go. This guide is about doing the extraction anyway, inside the insurer's own estate, and about being honest on how far today's tooling reaches. Its reader is the Claims Director, with the CISO in the room. It covers the insurer-specific legal condition, the documents a clinical model can read today, how to measure accuracy on your own files, and where Dilr Mira fits.

It deliberately cedes three things to neighbouring posts. The general case for on-premise clinical extraction belongs to our clinical document extraction guide. The map of where AI pays across the whole insurance value chain belongs to our UK insurance hub. First notification of loss on the phone belongs to our post on AI voice for insurance claims intake.

This guide is shipped by the team behind Dilr Mira, a class of private clinical small language models that turn scans, lab reports and claim forms into source-grounded JSON on your own hardware. Or see DATS, the consulting route through which regulated teams place models like this inside their estate.

What is health insurance claims document extraction?

Health insurance claims document extraction is the step that turns the medical pages inside a claims file, such as GP reports, consultant letters, discharge summaries and test results, into structured fields a claims handler can act on. For a UK insurer the output is a record of diagnosis, dates, treating clinician, treatment and stated limitations, each traceable to the page it came from, produced before anyone assesses the claim itself.

The distinction that matters is between reading and deciding. Extraction reads; the claims handler decides. A good extraction step shortens the time between a file arriving and a person being able to assess it, and it makes that person's assessment checkable, because every field points back to a source span. It does not decide cover, liability or quantum, and nothing in this guide suggests it should.

The pages differ by product. An income protection claim leans on GP and occupational health reports describing what the claimant can and cannot do. A critical illness claim leans on consultant letters and pathology results confirming a defined condition: the ABI reports that almost two-thirds (65%) of critical illness claims in 2025 were for cancer, so pathology reports and oncology letters are likely to feature heavily in that file. A bodily-injury claim, which typically arises under motor or liability cover rather than health insurance, leans on medico-legal reports, and the same extraction questions apply to its medical pages.

Which documents in a claims file can a clinical model read today?

A clinical extraction model such as Mira-Q2 is strongest today on structured clinical pages: lab reports, medication lists, discharge summaries, pathology reports, intake forms and progress notes, in English. Narrative GP and consultant letters are general physician prose, where measured validity is lower. Non-clinical pages, such as policy schedules, payment details and liability correspondence, sit outside a clinical model's scope, so an insurer should split the file first.

That split is the first design decision, and it is where a claims extraction project is most likely to go wrong. A claims file is not one document type. It is a mixed bundle, and a model that is good at physician prose has no special competence on a loss adjuster's narrative or a bank statement. The honest architecture classifies pages, sends the clinical ones to clinical extraction, and leaves the rest to whatever tooling, or person, already handles them. An AI placement diagnostic is one way to sample real files and see what the mix actually is. The page types matter commercially too: a critical illness file heavy in pathology reports sits close to what Mira-Q2 was trained on, while an income protection file built on narrative GP and occupational health letters sits further away.

The second decision is the schema. A claims team already knows which fields it needs, because its assessment templates and claims system define them. Mira-Q2 treats the schema as input: it is trained to follow the extraction schema it is given, and onboarding needs a schema file and a small set of seed examples, with no code changes. That is different from a model that can follow an unseen schema with no examples at all, which is a capability the live models page lists under Mira-3, a generation that is in development and not available. Writing the schema well is ordinary claims-operations work, and it is often the most valuable artefact of the whole exercise. Our AI operating model work treats it as a claims asset, not an IT one.

The third decision is what counts as done. A field is done when it is present, valid against the schema, and traceable to the page it came from. Anything else goes to a person.

Why keep the extraction inside the insurer's estate?

Health data in a claims file is special category data under UK GDPR Article 9, and a UK insurer processing it to administer a claim typically relies on the insurance condition in Schedule 1 to the Data Protection Act 2018. That condition, and the appropriate policy document it requires, bind the insurer as controller. Keeping extraction inside the insurer's own estate keeps the processing within controls the insurer already owns and documents.

The condition is paragraph 20 of Schedule 1 to the Data Protection Act 2018. It is met where processing is necessary for an insurance purpose, concerns data such as data concerning health, and is necessary for reasons of substantial public interest. The statute defines insurance purpose to include "administering a claim under an insurance contract", which is exactly what a claims team is doing when it reads a GP report. Paragraph 20 also adds a narrower test for some processing that is not used for decisions about the person and concerns someone with no rights or obligations under the contract: there, the processing must be something that can reasonably be carried out without that person's consent. That matters when a file contains health information about someone outside the policy.

Two practical consequences follow. First, paragraph 5 of the same Schedule says a Part 2 condition is met only if the controller has an appropriate policy document in place. Our guide to the appropriate policy document for special category data covers what that document has to contain; here the point is that every new processor or service in the extraction path is something it has to account for. Second, the ICO's guidance on special category data is the public reference point an insurer's DPO will hold the design against.

None of this forbids a cloud service. It does mean that sending medical pages out adds a processor, a transfer question and a new dependency to the insurer's documentation, while running a model on the insurer's own hardware adds none of those. For a CISO, that is usually the deciding argument, and it is why our DATS methodology starts from where the data is allowed to sit rather than from which model is cleverest.

What did the FCA's claims-handling review say about outsourced services?

The FCA's July 2025 claims-handling work found poor practice in the home and travel sector, including weak oversight of outsourced services, insufficient management information and delays in settling claims. It did not examine health or protection insurance. The lesson for a protection insurer is indirect: every external service in the claims path, an extraction service included, is something the insurer has to oversee.

The finding is worth reading in the regulator's own words. Among the poor practices the FCA's press release of 22 July 2025 listed, after noting some good practice in the home and travel sector, was:

"Lack of oversight of outsourced services, resulting in poor customer outcomes, delays in settling claims and high complaint volumes"

The same release named insufficient management information as a cause of failures to identify and resolve claims-handling issues promptly. Scope matters here. Those findings concern home and travel claims, and this guide does not claim the FCA found anything about health or protection claims. The point a protection insurer can take from it, without stretching the finding, is our own reading: an outsourced service in the claims path still has to be overseen by the insurer, and slow claims handling with poor management information drew the regulator's attention.

An on-premise extraction step helps on both counts without making any regulatory promise. It removes one outsourced dependency from the claims path, and because every extracted field is logged against its source page, it produces management information of the kind the review said was insufficient: how many files arrived, how many pages were clinical, how many fields passed the gate, and how many went to a person. It is the kind of telemetry an AI execution office can put in front of a claims board.

How should an insurer measure a model on its own claims files?

An insurer should measure any extraction model on a sample of its own claims files, not on the vendor's benchmark. Build a small gold set of real clinical pages with hand-checked fields, then measure four things: whether the output is valid against the claims schema, whether each field matches the gold value, whether each field is grounded in the source page, and whether any identifier leaks.

The reason is that published benchmarks rarely look like a claims file, and Mira-Q2's do not either. Its published evaluation on Hugging Face covers 782 clinical documents across four test sets, with zero identifier leaks across all 782. None of those documents is an insurance claims file. In our judgement, the closest public proxies for a GP or medico-legal report are the two real-physician sets. On mtsamples, 282 real physician documents across 39 specialties, JSON validity was 85.8%, with a 95% confidence interval of 81.9% to 89.7%. On a separate set of 150 real physician documents that match the training schema, validity was 94.7%.

Read those numbers carefully, because they say less than they appear to. They measure whether the output was valid JSON, not whether each field was correct: field-level accuracy is reported only for the held-out gold set drawn from the training distribution, where field-F1 was 1.000 with a confidence interval of 0.999 to 1.0, and the real-physician sets are unlabelled for field accuracy. So the published evaluation tells an insurer two useful things: the method does not leak identifiers on the documents tested, and on real-world prose roughly one output in seven needs a person before it is even well-formed. Field accuracy on your own files is something you measure, not something you inherit. Our Mira-Q2 evaluation deep-dive explains each test set in full.

What to measureHow to measure it on your filesWhat a failure means
Schema validityValidate every output against the claims schemaThe record cannot be loaded without a person
Field accuracyCompare fields with a hand-checked gold setThe handler would be misinformed
Source groundingCheck each field's cited span contains the valueThe field cannot be audited
Identifier leakageScan outputs for identifiers the schema did not ask forA data protection incident in the making

A gold set of a few hundred pages, drawn across income protection, critical illness and injury files, is enough to see the shape of the error. Re-run it whenever the schema or the model changes.

What does a verifier-gated claims extraction workflow look like?

A verifier-gated claims extraction workflow splits the claims file, sends clinical pages to an on-premise model, checks every output with a deterministic verifier, and routes anything that fails to a person. The verifier checks schema validity, source grounding and identifier leakage before any handler sees the result, so the handler reviews drafts that are already checked rather than raw model output.

The five steps below are the whole design. Nothing about them is specific to one model; they are what makes any model safe enough to put in a claims path.

A verifier-gated claims extraction path
01Claims file inBundle received02Pages splitClinical vs other03Model extractsInsurer schema04Verifier gateSchema, grounding, lea…05Handler reviewsDraft plus source
Every field reaches a handler as a checked draft with its source page attached; failures route to a person, never to an automated decision.

On the live Dilr Mira models page, the deterministic verifier is part of what ships today: it gates every output on schema validity, source grounding and zero identifier leakage, and outputs are drafts for human review, never autonomous decisions. Several governance features shown beside it, including signed receipts, deterministic replay and an externally anchored audit chain, are labelled in development, and an insurer should not plan around them yet. The claims system's own audit trail, which the insurer already runs, remains the record of who decided what.

The same gate pattern runs through our wider work. It is the logic behind the human approval gate we design for voice agents, applied to documents instead of calls.

What is the best way to extract health claims documents in 2026?

The best way to extract health claims documents in 2026 depends on where the data may go. If medical pages may leave the estate, a managed cloud medical-text service is quickest to start. If they may not, a model that runs on the insurer's own hardware, measured on the insurer's own files and gated by a verifier, is the defensible choice for a UK health or protection insurer.

The criteria we would apply, in order, are data location, measured accuracy on your own gold set, auditability of every field, schema fit, and operating effort. Against those criteria the options look different from how vendors present them.

AWS Comprehend Medical is a managed cloud service: fast to trial, but the pages leave the estate, so the appropriate policy document and processor arrangements have to cover it. Google's Healthcare Natural Language API has been deprecated, so it is not a sensible new dependency. John Snow Labs sells an established commercial Healthcare NLP product; for an insurer that wants a large commercial vendor with established support, it may suit better than a younger open model, provided its deployment option passes the same data-location test. General biomedical language models such as Meditron or BioMistral are research models rather than claims extraction tools, and Meditron's own model card recommends against deploying it in medical applications without extensive use-case alignment.

Dilr Mira sits in the narrow space of small, open, on-premise extraction: Mira-Q2 is open on Hugging Face under Apache-2.0, runs on CPU and takes about 2 GB on disk, with GPU only for speed. That makes it cheap to test inside a locked-down environment. It is not the answer for every insurer, and the comparison in our pillar guide sets out where each option wins.

Where does Dilr Mira fit today, and what is still research?

Dilr Mira is a candidate today for the clinical pages inside a claims file, provided an insurer measures it on its own documents first: Mira-Q2 is a shipped clinical extraction model that runs on the insurer's hardware behind a verifier gate. Insurance claims as a workflow is listed on the live models page as a research direction, not a shipped product. Clinical extraction is what ships today.

That distinction should shape any project. The live page lists insurance claims alongside prior authorisation, know-your-customer onboarding, invoices and receipts, lab networks and legal intake as regulated document worlds where the same verifier-gated method is under research. Prior authorisation is its own topic and outside this guide. For an insurer, the practical consequence is that a project today is a scoped engagement: measure Mira-Q2 on the clinical pages of your own files, decide whether the numbers clear your threshold, and build the page split and schema around it.

That engagement runs through DATS, our five-stage AI consulting system, rather than as a standalone product purchase. Our guide to what Dilr Mira is explains the model family and its release history, and the enterprise AI consulting guide explains how a scoped engagement is structured.

The wider context is familiar. McKinsey's State of AI 2025 found that 88% of organisations use AI in at least one function, while only around a third have moved it into production. In our view, in claims the gap between those two numbers is often a data-location argument that was never resolved. Settling where the medical pages may sit, before choosing a model, is how an insurer closes it.

Can Dilr Mira make claims decisions on its own?

No. Dilr Mira is an extraction tool, and its outputs are drafts for human review, never autonomous decisions. In a claims workflow it reads the clinical pages and returns checked, source-grounded fields; a claims handler still decides cover, liability and settlement. Any design that lets extracted fields trigger a decision without a person in between is outside how the model is described and intended to be used.

That is also the safer conduct position. A handler who can see each field's source page can explain a decision to a claimant, a complaints team or the ombudsman, which is harder when the reasoning lives inside a model.

Does the insurer or the model provider carry the special category duty?

The insurer carries it. Under the Data Protection Act 2018, the insurer is the controller processing health data to administer a claim, so the insurance condition in Schedule 1 paragraph 20, and the appropriate policy document it depends on, are the insurer's obligations. A model that runs on the insurer's own hardware does not shift that duty; it simply avoids adding an outside processor to it.

What a supplier can do is make the insurer's job easier to evidence: a model that runs locally, a published evaluation the insurer can rerun, and a gate whose checks are written down. That is what our operating model work is built to document.

Want to test this on your own claims files? See Dilr Mira and Mira-Q2, book an AI placement diagnostic, read our industry AI guides, or learn about Dilr.ai and how we place AI inside regulated systems.

Product
Dilr Mira
Service
AI Solutions
Service
AI Execution Office
Talk to the operators

Read claims files without moving them.

30-min scoping call · No deck · Confidential. We will tell you whether on-premise extraction fits your claims files, and how to measure it first.

Written by the Dilr.ai engineering team, practitioners who ship enterprise AI in production. Follow us on LinkedIn for shipping notes, or subscribe via the RSS feed.

health insurance claims document extractioninsurancebodily injury claims document extractionmedical report extraction insurer on-premprotection claims document ai uklocal llm redditbest clinical nlp model 2026dilr mira

Questions this article answers

What is health insurance claims document extraction?

Health insurance claims document extraction is the step that turns the medical pages inside a claims file, such as GP reports, consultant letters, discharge summaries and test results, into structured fields a claims handler can act on. For a UK insurer the output is a record of diagnosis, dates, treating clinician, treatment and stated limitations, each traceable to the page it came from, produced before anyone assesses the claim itself.

Which documents in a claims file can a clinical model read today?

A clinical extraction model such as Mira-Q2 is strongest today on structured clinical pages: lab reports, medication lists, discharge summaries, pathology reports, intake forms and progress notes, in English. Narrative GP and consultant letters are general physician prose, where measured validity is lower. Non-clinical pages, such as policy schedules, payment details and liability correspondence, sit outside a clinical model's scope, so an insurer should split the file first.

Why keep the extraction inside the insurer's estate?

Health data in a claims file is special category data under UK GDPR Article 9, and a UK insurer processing it to administer a claim typically relies on the insurance condition in Schedule 1 to the Data Protection Act 2018. That condition, and the appropriate policy document it requires, bind the insurer as controller. Keeping extraction inside the insurer's own estate keeps the processing within controls the insurer already owns and documents.

What did the FCA's claims-handling review say about outsourced services?

The FCA's July 2025 claims-handling work found poor practice in the home and travel sector, including weak oversight of outsourced services, insufficient management information and delays in settling claims. It did not examine health or protection insurance. The lesson for a protection insurer is indirect: every external service in the claims path, an extraction service included, is something the insurer has to oversee.

How should an insurer measure a model on its own claims files?

An insurer should measure any extraction model on a sample of its own claims files, not on the vendor's benchmark. Build a small gold set of real clinical pages with hand-checked fields, then measure four things: whether the output is valid against the claims schema, whether each field matches the gold value, whether each field is grounded in the source page, and whether any identifier leaks.

What does a verifier-gated claims extraction workflow look like?

A verifier-gated claims extraction workflow splits the claims file, sends clinical pages to an on-premise model, checks every output with a deterministic verifier, and routes anything that fails to a person. The verifier checks schema validity, source grounding and identifier leakage before any handler sees the result, so the handler reviews drafts that are already checked rather than raw model output.

What is the best way to extract health claims documents in 2026?

The best way to extract health claims documents in 2026 depends on where the data may go. If medical pages may leave the estate, a managed cloud medical-text service is quickest to start. If they may not, a model that runs on the insurer's own hardware, measured on the insurer's own files and gated by a verifier, is the defensible choice for a UK health or protection insurer.

Where does Dilr Mira fit today, and what is still research?

Dilr Mira is a candidate today for the clinical pages inside a claims file, provided an insurer measures it on its own documents first: Mira-Q2 is a shipped clinical extraction model that runs on the insurer's hardware behind a verifier gate. Insurance claims as a workflow is listed on the live models page as a research direction, not a shipped product. Clinical extraction is what ships today.

Dilr Voice

Voice AI built for your sector

Dilr Voice answers and places calls 24/7 with compliance rules for regulated industries, from clinics and estate agents to financial services.

Related articles

← Previous
LangGraph Work Management Layer: Adding a Done Gate

One email, once a month. No hype. Just what we learned shipping.