Strategy

Dilr Mira-Q2 Evaluation: The 782-Document Scorecard

Dilr Mira is a class of private clinical small language models from DILR.AI that extract schema-valid JSON on your own hardware. This is Mira-Q2's published evaluation across 782 documents in four test sets: JSON validity by set, field-F1 on held-out gold, and the honest gap between held-out gold and real physician prose.

Dilr Mira-Q2 Evaluation: The 782-Document Scorecard DILR MIRA Dilr Mira-Q2 Evaluation: The 782-Document Scorecard 782 clinical documents across four public test sets Source: Mira-Q2 scorecard, DILR.AI dilr.ai/blog

A single headline number for a clinical language model, offered with no way to see what was measured, on what, or how often it was wrong, is a claim rather than evidence. For a Head of Data Science or a Chief Data and AI Officer weighing an open model for document extraction inside a regulated organisation, that distinction decides everything. The question that matters before anything goes near a patient record is narrower and harder: on which documents was the model tested, how many, what counted as a pass, and where did it fail.

Dilr Mira is a class of private clinical small language models from DILR.AI that run on the customer's own hardware. Its current release, Mira-Q2, is around 3B parameters, built on Qwen2.5-3B-Instruct, and open on Hugging Face under Apache-2.0. More to the point of this page, its full evaluation is public too: the model, the evaluation files and a complete scorecard sit on the same page, so every figure below can be read in context and rerun rather than taken on trust. This post publishes that scorecard as a reference and explains how to read it.

The scorecard here is the published evaluation of Dilr Mira, the private clinical models from DILR.AI that turn clinical documents into schema-valid JSON on your own hardware. For how Mira fits a wider extraction decision, including deployment and the verifier, see the companion guide on clinical document extraction with open models.

This page stays on the evaluation itself. It does not re-argue why on-premise extraction matters, how the deterministic verifier works, or which model to choose for a given deployment; the open models guide covers all three. For what Dilr Mira is, how the name is distinct from the unrelated Hugging Face model families, and where it sits in the DILR.AI range, read what is Dilr Mira. What follows is the scorecard, the method behind it, and the honest gap the numbers show.

What does Mira-Q2's evaluation actually measure?

Mira-Q2's evaluation measures two different things across 782 clinical documents in four test sets. The first is JSON validity, the share of outputs that are schema-valid structured JSON rather than malformed or free text. The second is field-level F1, how closely the extracted field values match a gold reference. JSON validity is reported for every set and field-F1 for the held-out gold set, and the run recorded zero identifier leaks across all 782 documents.

The four sets are not interchangeable, and reading the scorecard well means knowing what each one probes. One, test_gold, is held out from the training distribution; one, synthetic_v2, tests a different formatting dialect; and two, extraction_relevant and mtsamples, are real physician documents, a different distribution from the constructed data the model was trained on. That split is the whole point of the design: a model can look strong on data shaped like its training set and still stumble on the free-form prose a hospital actually produces. An identifier leak, counted as zero across all 782 documents here, would be a direct patient identifier carried from a source document into the output. The table below is the scorecard in full, each row with its document count and its 95% confidence interval.

Eval setDocumentsWhat it testsJSON validity95% confidence interval
test_gold200Held-out, training distribution100.0%1.0 to 1.0
synthetic_v2150A different formatting dialect100.0%1.0 to 1.0
extraction_relevant150Real physician documents, on-schema94.7%90.7 to 98.0
mtsamples282Real physician documents, 39 specialties85.8%81.9 to 89.7

The documents add up to 782. The two real-physician sets, extraction_relevant and mtsamples, are where the harder test sits, and mtsamples is the broadest of them, spanning 39 specialties. A confidence interval of 1.0 to 1.0 on a set of 200 means no invalid output was observed in that run, not that failure is impossible. The verifier and the deployment choices that surround this model are described in the clinical document extraction guide, while every figure on this page is the live version published on the Mira-Q2 models page.

What is the difference between JSON validity and field-F1 in this scorecard?

JSON validity and field-F1 answer different questions, and Mira-Q2's scorecard keeps them apart on purpose. JSON validity asks whether the output is well-formed, schema-conformant structured data: a yes or no on each document. Field-F1 asks whether the values inside that structure are correct against a gold reference, which needs a labelled answer key. Validity travels across all four sets because it needs no gold labels; field-F1 is published for the held-out gold set alone, where a trusted reference exists.

This matters because the two can move independently. A model can emit perfectly valid JSON that contains the wrong drug name, which is why validity alone is never sufficient for a clinical pipeline, and why the deterministic verifier checks extracted values against the source document before anything reaches a human. Mira-Q2's field-F1 reaches 1.000 on the held-out gold set, at a 95% confidence interval of 0.999 to 1.0. That number is strong, and it is also the narrowest claim on the page: it is measured on data from the same distribution as training, with a gold key, and it says nothing on its own about free-form prose. The other three sets carry no gold key, so field-F1 is marked not applicable for them and only JSON validity is comparable across all four. Read the gold F1 alongside the validity figures for the two real-physician sets, not instead of them.

Why does Mira-Q2 score 100 percent on held-out gold but 85.8 percent on real physician prose?

Because the two sets are not the same kind of document. The held-out gold set is drawn from the same distribution Mira-Q2 was trained on, so a valid output is close to what the model already learned to produce. The mtsamples set is real physician prose across 39 specialties, a different distribution from its training data, where 85.8 percent of outputs were schema-valid. The gap is the distance between a clean test and the messy reality of clinical writing.

DILR.AI publishes that gap rather than reporting only the strongest figure. On the Mira-Q2 scorecard it is stated plainly:

100% on training-distribution data. 86% on general real physician prose. That gap is published, not hidden, and closing it is exactly what the next generation is for.

The 86 percent there is the rounding of the 85.8 percent mtsamples figure. A buyer should read the real-physician figures, 85.8 percent and 94.7 percent, as the realistic range on live documents, and the held-out gold result as a best case, then decide, with the verifier in the loop, whether that clears the bar for a given document type. The chart below shows the two figures side by side so the distance is visible rather than buried.

Mira-Q2 JSON validity: held-out gold versus general physician prose
100%Held-out gold85.8%General prose
Share of Mira-Q2 outputs that were schema-valid JSON, on held-out gold (test_gold, 200 documents) and on general physician prose (mtsamples, 282 documents). Both figures are from the same published Mira-Q2 evaluation. Source: Mira-Q2 scorecard, DILR.AI, 2026. Source: Mira-Q2 scorecard, DILR.AI

How was Mira-Q2 trained, and how did it progress from Mira-Q1?

Mira-Q2 is a QLoRA adapter, 4-bit, rank 16, trained on top of Qwen2.5-3B-Instruct. Its training set is 8,400 examples: 6,400 gold-by-construction, built from real ICD-10 codes, NLM drug names and curated lab reference ranges, plus 2,000 schema variants that teach it to accept different schemas as input. The reported train and evaluation loss are 0.132 and 0.142, an overfit gap of 0.010, which is small within the training distribution, though the real-physician results below show its limits.

The progression across generations is part of the published record, and it is the clearest evidence that the training, not the base model, did the work. Qwen2.5-3B with no training scores 0 percent on this task, inventing its own schema rather than following the one supplied. Mira-Q1, trained on 3,438 examples, reached 98 percent on a 50-example evaluation. Mira-Q2, trained on 8,400 examples, reached 100 percent on a 200-example held-out gold evaluation with field-F1 1.000. The three figures each come from that generation's own evaluation, so the diagram below reads as a progression rather than a single like-for-like run.

How Mira-Q2 reached full validity on held-out gold
01Qwen2.5-3Bzero-shot0%, invents its own sc…02Mira-Q198% on a 50-example ev…03Mira-Q2100% on a 200-example …
The figures are JSON validity, each on that generation's own evaluation, so they are a progression rather than a single like-for-like comparison. Source: Mira-Q2 scorecard, DILR.AI.

The model stays small enough to run where the data lives, about 2 GB on disk at 4-bit, the constraint that makes an open clinical model useful at all. Teams that want the architecture assessed against their own documents can start with an AI operating model review rather than a procurement cycle.

Where does this evaluation fit a real deployment decision?

A scorecard is a starting point for a decision, not the decision itself. The numbers tell you what Mira-Q2 can do on four defined document types; they do not tell you whether your documents look like those sets, what your schema needs, or where a 94.7 percent validity rate is acceptable. That mapping is consulting work, and it is where DATS, the five-stage AI consulting system from DILR.AI, sits.

A short AI placement diagnostic establishes which of your document flows fit a small on-premise model and which do not, before any model is deployed. The same logic runs through the AI execution office, where a DILR.AI team sits inside the delivery rather than handing over a report, which is the pattern that suits a regulated extraction rollout where the schema and the risk tolerance are still being settled. A private or schema-matched deployment is scoped through DATS rather than bought off the shelf, so the route in runs through a consulting engagement, not a checkout.

How do you reproduce Mira-Q2's evaluation yourself?

You reproduce it from the public artefacts. The model, the evaluation files and the full scorecard are published together as dilr/Mira-Q2 on Hugging Face under Apache-2.0. The model runs on CPU, takes about 2 GB on disk at 4-bit quantisation, and uses a GPU only for speed, so a validation run does not need specialist hardware and can be done inside your own network with no data leaving it.

Three things are worth knowing before you rerun it. The base is Qwen2.5-3B-Instruct loaded via Unsloth, which the card specifies, so load it the same way rather than through a standard adapter call. The evaluation files and the full scorecard are published alongside the model, so the figures can be checked against the published evaluation rather than taken on trust. The zero-leak figure is a measured result across all 782 documents, not a design promise. The model itself lives at dilr/Mira-Q2 on Hugging Face. For teams standing this up as part of a wider programme, the named enterprise AI solutions, the DILR.AI deployment approach and the DATS consulting method describe how a validated model moves into production.

Is Mira-Q2 a medical device, and can it make clinical decisions?

No. Mira-Q2 is a document extraction tool, and its outputs are structured drafts for human review, not autonomous clinical decisions. It reads a document and returns schema-valid fields grounded in that document; a clinician or a reviewer remains responsible for what is done with them. It is not marketed or validated as software as a medical device under the MHRA framework, and nothing in the scorecard should be read as a clinical accuracy claim about patient care.

That boundary is deliberate and it shapes how the numbers should be used. A 100 percent validity figure on held-out gold means the output was well-formed structured data, not that a diagnosis was correct. The verifier exists precisely so that an extracted value that the model invented rather than read is flagged before a person sees it. Where that review fits is part of the AI operating model a team puts around the tool. Used this way, the model sits under the human decision rather than replacing it, which is also the compliance posture that lets health data stay inside the organisation, since under UK GDPR Article 9 health data is special category and the safest handling is for it never to leave the building in the first place.

What is the best open clinical extraction model to evaluate in 2026?

The best model to evaluate depends on the document and the constraint, and an honest answer names more than one. For teams that need source-grounded, schema-valid extraction on their own hardware, with a published evaluation they can rerun, Mira-Q2 is a strong candidate where the schema is known and the documents resemble structured clinical records. For broad clinical reasoning across general medical text, larger biomedical models such as Meditron, BioMistral or OpenBioLLM are built for a different job.

The concession is real: a 3B extraction model is not the right tool for open-ended clinical reasoning, and on the most free-form physician prose its validity drops into the mid-eighties, which the scorecard shows rather than hides. The decision should turn on three things: whether your documents match the sets the model was tested on, whether a small on-premise footprint matters more than raw breadth, and whether you can rerun the vendor's evaluation at all. On the last point, a model whose evaluation files are public can be checked, which a headline figure on its own cannot. The AI for healthcare and AI for pharma guides set out where this kind of extraction pays first, and the AI for insurance guide covers the claims documents that share the pattern.

Is Mira-3 available yet?

No. Only Mira-Q2 is live and open on Hugging Face. Mira-3 is described by DILR.AI as in development: a multilingual model with zero-shot schema support and a trust layer. It is a roadmap, not a release, and nothing in this scorecard belongs to it. Any evaluation you run today should be run against Mira-Q2, the current published model, and judged on the figures in the table above.

Where do I find the exact model and evaluation files?

The exact model is dilr/Mira-Q2 on Hugging Face, published under Apache-2.0 with the evaluation files and the full scorecard alongside it. The name matters because the bare term resolves to unrelated model families and a listed pharmaceutical company, so use the dilr/Mira-Q2 namespace and the exact URL rather than a plain search. The Mira-Q2 models page carries the same figures in context.

To go further: read what Dilr Mira is, see the five-stage DATS method, review the strategy writing from the DILR.AI team, or read about DILR.AI and the people behind the work.

Product
Dilr Mira
Service
AI Placement Diagnostic
Service
AI Operating Model
Talk to the operators

Put a private clinical model where the evidence holds.

30-min scoping call · No deck · Confidential. We will tell you whether a small on-premise model fits your documents, and where it does not.

Written by the Dilr.ai engineering team, practitioners who ship enterprise AI in production. Follow us on LinkedIn for shipping notes, or subscribe via the RSS feed.

dilr mira-q2 evaluationmira-q2 scorecardclinical llm evaluationschema-valid json extractionopen clinical language modellocal llm redditbest clinical nlp model 2026dilr mira

Questions this article answers

What does Mira-Q2's evaluation actually measure?

Mira-Q2's evaluation measures two different things across 782 clinical documents in four test sets. The first is JSON validity, the share of outputs that are schema-valid structured JSON rather than malformed or free text. The second is field-level F1, how closely the extracted field values match a gold reference. JSON validity is reported for every set and field-F1 for the held-out gold set, and the run recorded zero identifier leaks across all 782 documents.

What is the difference between JSON validity and field-F1 in this scorecard?

JSON validity and field-F1 answer different questions, and Mira-Q2's scorecard keeps them apart on purpose. JSON validity asks whether the output is well-formed, schema-conformant structured data: a yes or no on each document. Field-F1 asks whether the values inside that structure are correct against a gold reference, which needs a labelled answer key. Validity travels across all four sets because it needs no gold labels; field-F1 is published for the held-out gold set alone, where a trusted reference exists.

Why does Mira-Q2 score 100 percent on held-out gold but 85.8 percent on real physician prose?

Because the two sets are not the same kind of document. The held-out gold set is drawn from the same distribution Mira-Q2 was trained on, so a valid output is close to what the model already learned to produce. The mtsamples set is real physician prose across 39 specialties, a different distribution from its training data, where 85.8 percent of outputs were schema-valid. The gap is the distance between a clean test and the messy reality of clinical writing.

How was Mira-Q2 trained, and how did it progress from Mira-Q1?

Mira-Q2 is a QLoRA adapter, 4-bit, rank 16, trained on top of Qwen2.5-3B-Instruct. Its training set is 8,400 examples: 6,400 gold-by-construction, built from real ICD-10 codes, NLM drug names and curated lab reference ranges, plus 2,000 schema variants that teach it to accept different schemas as input. The reported train and evaluation loss are 0.132 and 0.142, an overfit gap of 0.010, which is small within the training distribution, though the real-physician results below show its limits.

Where does this evaluation fit a real deployment decision?

A scorecard is a starting point for a decision, not the decision itself. The numbers tell you what Mira-Q2 can do on four defined document types; they do not tell you whether your documents look like those sets, what your schema needs, or where a 94.7 percent validity rate is acceptable. That mapping is consulting work, and it is where DATS, the five-stage AI consulting system from DILR.AI, sits.

How do you reproduce Mira-Q2's evaluation yourself?

You reproduce it from the public artefacts. The model, the evaluation files and the full scorecard are published together as dilr/Mira-Q2 on Hugging Face under Apache-2.0. The model runs on CPU, takes about 2 GB on disk at 4-bit quantisation, and uses a GPU only for speed, so a validation run does not need specialist hardware and can be done inside your own network with no data leaving it.

Is Mira-Q2 a medical device, and can it make clinical decisions?

No. Mira-Q2 is a document extraction tool, and its outputs are structured drafts for human review, not autonomous clinical decisions. It reads a document and returns schema-valid fields grounded in that document; a clinician or a reviewer remains responsible for what is done with them. It is not marketed or validated as software as a medical device under the MHRA framework, and nothing in the scorecard should be read as a clinical accuracy claim about patient care.

What is the best open clinical extraction model to evaluate in 2026?

The best model to evaluate depends on the document and the constraint, and an honest answer names more than one. For teams that need source-grounded, schema-valid extraction on their own hardware, with a published evaluation they can rerun, Mira-Q2 is a strong candidate where the schema is known and the documents resemble structured clinical records. For broad clinical reasoning across general medical text, larger biomedical models such as Meditron, BioMistral or OpenBioLLM are built for a different job.

AI consulting (DATS)

Place AI where the P&L moves

The DATS system runs from a fixed-fee placement diagnostic through to embedded delivery, so AI reaches production instead of staying a pilot.

Related articles

← Previous
AI for Wealth and PE in the UK: Where It Pays in 2026

One email, once a month. No hype. Just what we learned shipping.