Module 6 · 85 min
Evaluating Healthcare AI
From Vendor Claims to Evidence-Based Decisions. How to judge whether a healthcare AI claim is credible: which evidence to demand at each layer, what the headline metrics do and do not tell you, whether a result transfers to your setting, the red flags in an impressive-looking deck, the evidence ladder from validation to real-world impact, and the questions to ask a vendor or data-science team. Technical methodology is available as optional deep dives.
- Separate the three claims an AI proposal usually blends together: model performance, workflow performance, and patient or operational impact.
- Interpret sensitivity, specificity, PPV, NPV, AUC and calibration in plain language, using absolute counts and alert workload rather than formulas.
- State precisely what each headline metric cannot tell you, and convert that gap into a question for the vendor or data-science team.
- Judge whether evidence generated elsewhere is likely to hold for your patients, your case mix and your workflow, and treat transportability as an empirical question rather than a permanent badge.
- Recognise the recurring red flags in impressive-looking evidence packs, including subgroup claims resting on a handful of events.
- Place a piece of evidence on the ladder from retrospective validation to post-deployment monitoring, and say which claim that rung can and cannot support.
- Run a structured evidence review of a realistic package and produce a defensible go / limited pilot / more evidence first / no-go recommendation, naming the evidence that would change it.
Continuing from Module 5
Module 5 argued that a grounded, tool-using system has to be measured in layers rather than by a single accuracy figure. This module takes the leader's seat at the other end of that work: a pack lands on the table, a decision is due, and the question is whether the evidence in front of you supports the claim being made.
Building on Essentials
You already know how to challenge a claim: Essentials taught claim, evidence, fit and impact, the comparator question, and how to spot a vanity metric. That framework is not repeated here. This module makes the same judgement quantitative and deployment-specific — and note that its ladder is a different ladder: Essentials ranked the strength of evidence behind a claim, whereas here the rungs are the stages a system passes through, from retrospective validation to prospective silent mode, supervised live use, comparative impact evaluation and post-deployment monitoring.
What evidence should I demand?
Evaluation does not begin with a metric. It begins with a claim, and Module 3 insisted that the claim be a precise intended use: which patients, at which moment, for which decision, with which action available.
Most disputes about healthcare AI are not arithmetic disputes. The numbers on the slide are often calculated correctly. What goes wrong is the distance between the number and the claim: a ranking statistic used to argue that a risk percentage can be trusted; a retrospective analysis used to argue that a ward will run better; an average across thousands of patients used to argue that a small specialty is safely covered.
So the habit worth building is short enough to use in a meeting. For every number you are shown, ask: what exactly was measured, in whom, and at what operating point? What claim does that measurement support? What claim is actually being made?
1. Model performance
Does the model rank or estimate risk well on data?
- What counts as evidence
- Retrospective analyses reporting discrimination, calibration, and threshold-specific figures such as sensitivity, PPV and alert volume.
- What it still cannot show
- This is the easiest evidence to produce and the weakest kind of proof. It says nothing about what clinicians do with an output, or whether anything changes for a patient.
2. Workflow performance
Does the system work in the actual service, with real staff and real timing?
- What counts as evidence
- Silent-mode runs, live pilots, alert volumes per shift, data completeness and latency, response and override behaviour, usability and time cost.
- What it still cannot show
- Strong workflow evidence still does not establish patient benefit. A well-run alert that no one can act on in time changes nothing except workload.
3. Patient and operational impact
Do outcomes, safety or capacity actually improve — and compared with what?
- What counts as evidence
- Comparative evaluation against current practice: before-and-after with controls, stepped-wedge, or randomised designs, with pre-specified outcomes.
- What it still cannot show
- This is the evidence most often claimed and least often supplied. Where it is absent, the honest phrasing is that impact is expected, not demonstrated.
One metric is never enough — and not every system uses the same ones
Prediction models
Risk, diagnosis and prognosis models that output a probability for an event that either did or did not happen. These carry the familiar metric set — sensitivity, specificity, PPV, NPV, AUC, calibration — and they are the focus of this module.
Generative systems
Ambient documentation, drafting and summarisation, where there is rarely one correct output. Judge them on fidelity to the source, omission and fabrication rates, clinically significant error rates, edit burden and time. Sensitivity and AUC generally do not apply.
Retrieval and agent systems
Grounded assistants and bounded workflow agents, measured in layers as Module 5 described: retrieval quality, then answer groundedness and citation support, then action correctness and approval behaviour. One end-to-end accuracy figure hides which layer failed.
The study-design half of this module — local validity, real-world evidence, the questions to ask — applies to all three. The metric vocabulary in the next step applies mainly to the first, which is also where most vendor claims live.
Sources & evidence · 8 sources
This module cites scholarly literature.
Content reviewed: September 2026. Publication dates of the individual sources are shown in each citation.
Collins GS, Moons KGM, Dhiman P, et al. TRIPOD+AI statement: updated guidance for reporting clinical prediction models that use regression or machine learning methods. BMJ. 2024;385:e078378. doi:10.1136/bmj-2023-078378. Published 16 April 2024. (A correction was published on 18 April 2024: BMJ. 2024;385:q902, doi:10.1136/bmj.q902; the substantive guidance is unchanged.)
The reporting standard for prediction-model studies, including AI methods. It specifies what must be reported so a study can be appraised. It is not a quality or safety standard, and adherence says nothing about whether the model performs well.
Open sourceMoons KGM, Damen JAA, Kaul T, et al. PROBAST+AI: an updated quality, risk of bias, and applicability assessment tool for prediction models using regression or artificial intelligence methods. BMJ. 2025;388:e082505. doi:10.1136/bmj-2024-082505
A structured tool for appraising risk of bias and applicability in prediction-model studies. It supports a judgement about a study; it does not certify a model or license deployment.
Open sourceVan Calster B, Collins GS, Vickers AJ, et al. Evaluation of performance measures in predictive artificial intelligence models to support medical decisions: overview and guidance. Lancet Digital Health. 2025;7(12):100916. doi:10.1016/j.landig.2025.100916. PMID 41391983.
A contemporary overview of which performance measures answer which question, and how they are misused. Methodological guidance from an expert group, not a regulation or a mandatory standard.
Open sourceVan Calster B, McLernon DJ, van Smeden M, Wynants L, Steyerberg EW. Calibration: the Achilles heel of predictive analytics. BMC Medicine. 2019;17:230. doi:10.1186/s12916-019-1466-7. PMID 31842878.
The standard reference for why calibration matters and how discrimination can look strong while probabilities are wrong. Background for the calibration deep dive; a methodological review rather than evidence about any particular system.
Open sourceVickers AJ, Elkin EB. Decision curve analysis: a novel method for evaluating prediction models. Medical Decision Making. 2006;26(6):565-574. doi:10.1177/0272989X06295361. PMID 17099194.
The original description of net benefit and decision curves, referenced by the optional threshold deep dive. It compares strategies on evaluation data; it does not demonstrate that deploying a model improves outcomes.
Open sourceRiley RD, Debray TPA, Collins GS, et al. Minimum sample size for external validation of a clinical prediction model with a binary outcome. Statistics in Medicine. 2021;40(19):4230-4251. doi:10.1002/sim.9025. PMID 34031906.
Sizing for external validation driven by the precision required, given the expected event proportion. Referenced by the optional validation deep dive; useful for retiring fixed event-count folklore.
Open sourceVasey B, Nagendran M, Campbell B, et al. Reporting guideline for the early-stage clinical evaluation of decision support systems driven by artificial intelligence: DECIDE-AI. Nature Medicine. 2022;28:924-933. doi:10.1038/s41591-022-01772-9. (A publisher correction has been issued; cite the current record.)
A reporting frame for early live evaluation, covering human factors, safety and implementation — the third rung of the evidence ladder. It structures how such a study is described; it is not evidence that a system is safe or effective.
Open sourceSounderajah V, Guni A, Liu X, et al. The STARD-AI reporting guideline for diagnostic accuracy studies using artificial intelligence. Nature Medicine. 2025;31(10):3283-3289. doi:10.1038/s41591-025-03953-8. Published online 15 September 2025. (An Author Correction, doi:10.1038/s41591-026-04570-9, corrected an author name; the methodological content is unchanged.)
The AI-specific extension of STARD, including reporting of setting, population, reference standard and operating point — the details that determine whether an accuracy estimate travels. A reporting guideline, not an evidence standard.
Open source
Applied practice: Model Evaluation Lab
Review a hypothetical vendor evidence pack as a decision-maker: read the headline metrics, find what is missing, choose the next evidence step and make a pilot recommendation.
Capstone Board Pack
Add what you just learned to your own strategy document while it is fresh.