Module 6 · 85 min
Evaluating Healthcare AI
From Vendor Claims to Evidence-Based Decisions. How to judge whether a healthcare AI claim is credible: which evidence to demand at each layer, what the headline metrics do and do not tell you, whether a result transfers to your setting, the red flags in an impressive-looking deck, the evidence ladder from validation to real-world impact, and the questions to ask a vendor or data-science team. Technical methodology is available as optional deep dives.
- Separate the three claims an AI proposal usually blends together: model performance, workflow performance, and patient or operational impact.
- Interpret sensitivity, specificity, PPV, NPV, AUC and calibration in plain language, using absolute counts and alert workload rather than formulas.
- State precisely what each headline metric cannot tell you, and convert that gap into a question for the vendor or data-science team.
- Judge whether evidence generated elsewhere is likely to hold for your patients, your case mix and your workflow, and treat transportability as an empirical question rather than a permanent badge.
- Recognise the recurring red flags in impressive-looking evidence packs, including subgroup claims resting on a handful of events.
- Place a piece of evidence on the ladder from retrospective validation to post-deployment monitoring, and say which claim that rung can and cannot support.
- Run a structured evidence review of a realistic package and produce a defensible go / limited pilot / more evidence first / no-go recommendation, naming the evidence that would change it.
The questions to ask the vendor or data-science team
Ten questions, in the order that usually works. The first three settle what is being claimed; the middle four test whether the evidence supports it here; the last three decide what happens next.
About this aid
An educational review aid for structuring a conversation with a vendor, a data-science team or a project sponsor. It is not a validated instrument and has no regulatory or clinical standing; it does not replace clinical safety assessment, information-governance review or regulatory conformity assessment.
1. What exactly is the intended use, and which decision changes?
- Why it matters
- Everything downstream depends on this. A system without a named decision has no evaluable claim.
- Weak answer
- A description of the technology instead of a decision, or a list of possible uses.
2. In which population was this evaluated, and how does it compare with ours?
- Why it matters
- Case mix, event rate and data capture determine whether the numbers travel.
- Weak answer
- Only a total sample size, with no description of the patients or the setting.
3. What is the operating threshold, and what alert or workload volume does it produce here?
- Why it matters
- Without an operating point there is no workload plan, no PPV and no staffing conversation.
- Weak answer
- “The threshold is configurable”, offered as an answer rather than as a next step.
4. What are the discrimination and calibration results — not just the headline figure?
- Why it matters
- If a displayed risk drives the decision, calibration is as important as ranking.
- Weak answer
- AUC alone, or calibration described as good without a plot or a table.
5. What is the PPV, and what false-positive burden falls on the responding service?
- Why it matters
- This is the number the service experiences daily, and it is a common operational failure point.
- Weak answer
- PPV from a higher-risk population than the one you plan to run on.
6. Has it been evaluated externally or locally, without tuning on the evaluation data?
- Why it matters
- Tuning on the data used to report the result reproduces the developer's optimism.
- Weak answer
- “We recalibrated to your site and report performance on the same records.”
7. What do subgroup results look like, and how many events sit behind each one?
- Why it matters
- Averages hide the groups where performance is unknown, and small groups need honesty rather than reassurance.
- Weak answer
- A significance test offered in place of event counts and intervals.
8. Has the system run prospectively on live data, and what changed when it did?
- Why it matters
- Live data completeness, timing or alert volume can differ materially from retrospective estimates.
- Weak answer
- No silent-mode run, or one whose alert volume is not reported.
9. What comparative evidence supports the claimed workflow or patient benefit?
- Why it matters
- This separates a demonstrated impact from a modelled expectation.
- Weak answer
- Projected benefits derived from retrospective accuracy, presented as results.
10. Is there a named plan for monitoring this after go-live, and who is accountable for reviewing it?
- Why it matters
- A deployment with no monitoring plan cannot detect its own failure. What exactly gets watched, who owns it and what triggers a stop is Module 7's territory; this decision only needs to know a plan exists.
- Weak answer
- Monitoring described as available in the platform, with no plan or owner confirmed.
Used well, this is not an interrogation. Good suppliers answer most of it quickly and tell you plainly which rungs they have not reached — and that answer, given openly, is a better signal of quality than any single number in the pack.
Sources & evidence · 8 sources
This module cites scholarly literature.
Content reviewed: September 2026. Publication dates of the individual sources are shown in each citation.
Collins GS, Moons KGM, Dhiman P, et al. TRIPOD+AI statement: updated guidance for reporting clinical prediction models that use regression or machine learning methods. BMJ. 2024;385:e078378. doi:10.1136/bmj-2023-078378. Published 16 April 2024. (A correction was published on 18 April 2024: BMJ. 2024;385:q902, doi:10.1136/bmj.q902; the substantive guidance is unchanged.)
The reporting standard for prediction-model studies, including AI methods. It specifies what must be reported so a study can be appraised. It is not a quality or safety standard, and adherence says nothing about whether the model performs well.
Open sourceMoons KGM, Damen JAA, Kaul T, et al. PROBAST+AI: an updated quality, risk of bias, and applicability assessment tool for prediction models using regression or artificial intelligence methods. BMJ. 2025;388:e082505. doi:10.1136/bmj-2024-082505
A structured tool for appraising risk of bias and applicability in prediction-model studies. It supports a judgement about a study; it does not certify a model or license deployment.
Open sourceVan Calster B, Collins GS, Vickers AJ, et al. Evaluation of performance measures in predictive artificial intelligence models to support medical decisions: overview and guidance. Lancet Digital Health. 2025;7(12):100916. doi:10.1016/j.landig.2025.100916. PMID 41391983.
A contemporary overview of which performance measures answer which question, and how they are misused. Methodological guidance from an expert group, not a regulation or a mandatory standard.
Open sourceVan Calster B, McLernon DJ, van Smeden M, Wynants L, Steyerberg EW. Calibration: the Achilles heel of predictive analytics. BMC Medicine. 2019;17:230. doi:10.1186/s12916-019-1466-7. PMID 31842878.
The standard reference for why calibration matters and how discrimination can look strong while probabilities are wrong. Background for the calibration deep dive; a methodological review rather than evidence about any particular system.
Open sourceVickers AJ, Elkin EB. Decision curve analysis: a novel method for evaluating prediction models. Medical Decision Making. 2006;26(6):565-574. doi:10.1177/0272989X06295361. PMID 17099194.
The original description of net benefit and decision curves, referenced by the optional threshold deep dive. It compares strategies on evaluation data; it does not demonstrate that deploying a model improves outcomes.
Open sourceRiley RD, Debray TPA, Collins GS, et al. Minimum sample size for external validation of a clinical prediction model with a binary outcome. Statistics in Medicine. 2021;40(19):4230-4251. doi:10.1002/sim.9025. PMID 34031906.
Sizing for external validation driven by the precision required, given the expected event proportion. Referenced by the optional validation deep dive; useful for retiring fixed event-count folklore.
Open sourceVasey B, Nagendran M, Campbell B, et al. Reporting guideline for the early-stage clinical evaluation of decision support systems driven by artificial intelligence: DECIDE-AI. Nature Medicine. 2022;28:924-933. doi:10.1038/s41591-022-01772-9. (A publisher correction has been issued; cite the current record.)
A reporting frame for early live evaluation, covering human factors, safety and implementation — the third rung of the evidence ladder. It structures how such a study is described; it is not evidence that a system is safe or effective.
Open sourceSounderajah V, Guni A, Liu X, et al. The STARD-AI reporting guideline for diagnostic accuracy studies using artificial intelligence. Nature Medicine. 2025;31(10):3283-3289. doi:10.1038/s41591-025-03953-8. Published online 15 September 2025. (An Author Correction, doi:10.1038/s41591-026-04570-9, corrected an author name; the methodological content is unchanged.)
The AI-specific extension of STARD, including reporting of setting, population, reference standard and operating point — the details that determine whether an accuracy estimate travels. A reporting guideline, not an evidence standard.
Open source
Applied practice: Model Evaluation Lab
Review a hypothetical vendor evidence pack as a decision-maker: read the headline metrics, find what is missing, choose the next evidence step and make a pilot recommendation.
Capstone Board Pack
Add what you just learned to your own strategy document while it is fresh.