Healthcare AI Learning
Course overview

Module 6 · 85 min

Evaluating Healthcare AI

From Vendor Claims to Evidence-Based Decisions. How to judge whether a healthcare AI claim is credible: which evidence to demand at each layer, what the headline metrics do and do not tell you, whether a result transfers to your setting, the red flags in an impressive-looking deck, the evidence ladder from validation to real-world impact, and the questions to ask a vendor or data-science team. Technical methodology is available as optional deep dives.

Lesson progress
0 / 9 steps
Not started
Learning objectives
  • Separate the three claims an AI proposal usually blends together: model performance, workflow performance, and patient or operational impact.
  • Interpret sensitivity, specificity, PPV, NPV, AUC and calibration in plain language, using absolute counts and alert workload rather than formulas.
  • State precisely what each headline metric cannot tell you, and convert that gap into a question for the vendor or data-science team.
  • Judge whether evidence generated elsewhere is likely to hold for your patients, your case mix and your workflow, and treat transportability as an empirical question rather than a permanent badge.
  • Recognise the recurring red flags in impressive-looking evidence packs, including subgroup claims resting on a handful of events.
  • Place a piece of evidence on the ladder from retrospective validation to post-deployment monitoring, and say which claim that rung can and cannot support.
  • Run a structured evidence review of a realistic package and produce a defensible go / limited pilot / more evidence first / no-go recommendation, naming the evidence that would change it.

How to spot suspiciously impressive evidence

Before the red flags, it is worth knowing what a genuinely good pack looks like — partly so you can recognise one, and partly so you can tell a vendor exactly what would get a yes. Strong evidence is not the same as flawless evidence; it is evidence whose claim is no larger than what was shown.

The claim matches the rung

The pack says what its strongest study design can support and no more — 'externally validated on two health systems' rather than 'proven to save lives'.

A named decision and a named owner

The evidence is framed around a decision that changes, the person who makes it and the action they take — not around a metric in isolation.

Performance reported at the operating point you would use

Sensitivity, specificity, positive predictive value and expected alert volume at the actual threshold and local prevalence, not only a summary curve.

Subgroups reported without being asked

Results broken down by the groups that matter locally, with sample sizes shown, including where the pack is underpowered to say anything.

A comparator that reflects current practice

The relevant question is incremental value over what the service does today, including any existing score or rule.

Balancing measures and workload reported

The pack states what got worse or busier as well as what improved — a credible sign that the evaluation was not designed only to succeed.

A pre-specified analysis and disclosed limitations

Outcomes and thresholds fixed in advance, deviations reported, and limitations named by the vendor before you find them.

A monitoring plan already drafted

The vendor expects performance and usage to be watched after go-live, and has named the signals to report, the review cadence and the owner who reviews them. (What safety signals, stop rules and operational pause logic that owner needs is Module 7's territory, not the evidence pack's.)

A credible positive evidence package (fictional)

A vendor proposes a deterioration score for adult medical wards. The pack contains: external validation at three health systems that supplied no development data, with site-by-site discrimination and calibration reported rather than pooled; performance at the intended operating point, with expected alerts per ward per day at local prevalence; subgroups by age, sex, ethnicity and admission route, with sample sizes shown and two subgroups declared underpowered; a comparator against the early-warning score the wards use today, showing modest incremental sensitivity at equal alert volume; and a prospective silent-mode period at one site describing who would have acted, when, and what the workload would have been.

The plausible benefit signal is a shorter time to review for patients the current score misses. The pack states what is not shown: no effect on mortality or length of stay has been demonstrated, and no site has yet run it live with clinicians acting on it.

The right conclusion is not that benefit is proven. It is that this evidence is good enough to justify the next stage — a supervised live pilot with pre-specified endpoints — and that the remaining uncertainty is named precisely enough to design that pilot around it.

Very few evidence packs are dishonest. Most are selectively assembled by people who genuinely believe in the product, using figures that are individually correct. The red flags below are about what is missing or mismatched, not about arithmetic errors.

Treat each one as a prompt to ask, not as a verdict. Several are entirely reasonable at an early stage — provided the claim being made is correspondingly modest.

Only development or internal results

The model has never been shown to work anywhere it was not built. It may still be promising; it is not yet transportable.

Ask: Where has this run without refitting, and on whose patients?

AUC quoted as accuracy or as benefit

A ranking statistic is being asked to stand in for correctness, calibration and clinical value at once.

Ask: What is the performance at the threshold we would use, and what are the calibration results?

Absolute risk shown, no calibration reported

Clinicians will read a percentage and act on it, with no evidence that the percentage corresponds to reality.

Ask: Among patients told 20%, what proportion actually had the outcome?

No threshold, no alert numbers

Without an operating point there is no workload, no PPV and no way to plan the responding service.

Ask: At the proposed threshold, how many alerts per day, and who answers them?

Subgroup claims resting on tiny event counts

A percentage computed from a handful of events is compatible with almost any true value.

Ask: How many events sit behind each subgroup figure, and how wide are the intervals?

“No significant difference between groups” used as reassurance

Non-significance in a small comparison usually reflects low power, not equivalence.

Ask: What difference would this study have been able to detect?

Huge sample, questionable data construction

Size buys precision, not validity. Leakage or a badly chosen index time makes a very precise wrong answer.

Ask: What information was available at the moment of prediction, and how was the outcome defined?

Cross-study comparison of headline numbers

Comparing a vendor's AUC with a competitor's from a different cohort compares populations, not products.

Ask: Has anyone compared these on the same data, at the same operating point?

Outcome or reference standard mismatched to intended use

A model predicting a coding-derived proxy may be measured against something the clinical decision does not turn on.

Ask: How was the outcome ascertained, and is it the event we care about?

“Improved outcomes” supported only by retrospective accuracy

This is the most consequential category error in the field: a measurement of prediction presented as a measurement of impact.

Ask: What comparative evidence exists that care changed and patients did better?

Review the deck

Six slides from a fictional vendor pack. For each, name the dominant problem. Attempt all six to continue; correctness is not required.

0 / 6 correct · 0 of 6 attempted

1. Slide 4

“AUC 0.94 — the model is correct in 94% of cases.”

2. Slide 7

“Developed and validated on 180,000 encounters from our partner network, with a held-out test set.”

3. Slide 11

“Sensitivity 95% for sepsis onset. Deployment requires no additional staffing.”

4. Slide 14

“In our retrospective analysis the model identified 82% of deteriorations, so earlier intervention would have prevented an estimated 30 ICU admissions.”

5. Slide 18

“Performance is consistent across age, sex and ethnicity (all differences p > 0.05).”

6. Slide 21

“Our AUC of 0.88 exceeds the 0.81 reported for the leading competitor.”

A pack with red flags is not necessarily a bad product. It is a product whose claims currently run ahead of its evidence — which is a negotiable position, and usually the right starting point for a limited, well-instrumented pilot.

Attempt all 6 items to continue — 0 done so far. Answers do not have to be correct.
Sources & evidence · 8 sources

This module cites scholarly literature.

Content reviewed: September 2026. Publication dates of the individual sources are shown in each citation.

  • Collins GS, Moons KGM, Dhiman P, et al. TRIPOD+AI statement: updated guidance for reporting clinical prediction models that use regression or machine learning methods. BMJ. 2024;385:e078378. doi:10.1136/bmj-2023-078378. Published 16 April 2024. (A correction was published on 18 April 2024: BMJ. 2024;385:q902, doi:10.1136/bmj.q902; the substantive guidance is unchanged.)

    The reporting standard for prediction-model studies, including AI methods. It specifies what must be reported so a study can be appraised. It is not a quality or safety standard, and adherence says nothing about whether the model performs well.

    Open source
  • Moons KGM, Damen JAA, Kaul T, et al. PROBAST+AI: an updated quality, risk of bias, and applicability assessment tool for prediction models using regression or artificial intelligence methods. BMJ. 2025;388:e082505. doi:10.1136/bmj-2024-082505

    A structured tool for appraising risk of bias and applicability in prediction-model studies. It supports a judgement about a study; it does not certify a model or license deployment.

    Open source
  • Van Calster B, Collins GS, Vickers AJ, et al. Evaluation of performance measures in predictive artificial intelligence models to support medical decisions: overview and guidance. Lancet Digital Health. 2025;7(12):100916. doi:10.1016/j.landig.2025.100916. PMID 41391983.

    A contemporary overview of which performance measures answer which question, and how they are misused. Methodological guidance from an expert group, not a regulation or a mandatory standard.

    Open source
  • Van Calster B, McLernon DJ, van Smeden M, Wynants L, Steyerberg EW. Calibration: the Achilles heel of predictive analytics. BMC Medicine. 2019;17:230. doi:10.1186/s12916-019-1466-7. PMID 31842878.

    The standard reference for why calibration matters and how discrimination can look strong while probabilities are wrong. Background for the calibration deep dive; a methodological review rather than evidence about any particular system.

    Open source
  • Vickers AJ, Elkin EB. Decision curve analysis: a novel method for evaluating prediction models. Medical Decision Making. 2006;26(6):565-574. doi:10.1177/0272989X06295361. PMID 17099194.

    The original description of net benefit and decision curves, referenced by the optional threshold deep dive. It compares strategies on evaluation data; it does not demonstrate that deploying a model improves outcomes.

    Open source
  • Riley RD, Debray TPA, Collins GS, et al. Minimum sample size for external validation of a clinical prediction model with a binary outcome. Statistics in Medicine. 2021;40(19):4230-4251. doi:10.1002/sim.9025. PMID 34031906.

    Sizing for external validation driven by the precision required, given the expected event proportion. Referenced by the optional validation deep dive; useful for retiring fixed event-count folklore.

    Open source
  • Vasey B, Nagendran M, Campbell B, et al. Reporting guideline for the early-stage clinical evaluation of decision support systems driven by artificial intelligence: DECIDE-AI. Nature Medicine. 2022;28:924-933. doi:10.1038/s41591-022-01772-9. (A publisher correction has been issued; cite the current record.)

    A reporting frame for early live evaluation, covering human factors, safety and implementation — the third rung of the evidence ladder. It structures how such a study is described; it is not evidence that a system is safe or effective.

    Open source
  • Sounderajah V, Guni A, Liu X, et al. The STARD-AI reporting guideline for diagnostic accuracy studies using artificial intelligence. Nature Medicine. 2025;31(10):3283-3289. doi:10.1038/s41591-025-03953-8. Published online 15 September 2025. (An Author Correction, doi:10.1038/s41591-026-04570-9, corrected an author name; the methodological content is unchanged.)

    The AI-specific extension of STARD, including reporting of setting, population, reference standard and operating point — the details that determine whether an accuracy estimate travels. A reporting guideline, not an evidence standard.

    Open source

Applied practice: Model Evaluation Lab

Review a hypothetical vendor evidence pack as a decision-maker: read the headline metrics, find what is missing, choose the next evidence step and make a pilot recommendation.

Open lab

Capstone Board Pack

Add what you just learned to your own strategy document while it is fresh.

Open capstone