Module 3 · 85 min
Clinical AI: From Prediction to Clinical Utility
Diagnosis, prognosis, risk stratification, monitoring and treatment response; precise intended use, the chain from prediction to decision to action to outcome, threshold trade-offs, causal versus predictive claims, and the clinical evidence that actually counts.
- Distinguish diagnosis, prognosis, risk stratification, monitoring and treatment-response tasks, and say why the distinction changes the evidence you need.
- Write a precise intended-use statement: population, index moment, output, user, decision and care setting.
- Trace the full chain from prediction to decision to action to outcome, and locate the weakest link.
- Judge whether a prediction is actionable at all, given the capacity, timing and available intervention.
- Reason qualitatively about operating thresholds as an explicit trade of one kind of harm against another.
- Separate a prediction claim from a treatment-effect claim, and recognise when a predictive model is being used as if it were causal.
- Ask for the right level of clinical evidence — internal, external, prospective, live and comparative — and know what each does and does not establish.
- Apply a reusable ten-item executive review to any proposed clinical AI system.
Building on Essentials
Essentials established that an AI output is not value in itself: someone has to act on it for anything to change. That point is not re-argued here. This module formalises it — intended use, the five clinical tasks, the chain from prediction to decision to action to outcome, thresholds and harm, actionability under real capacity, and the line between a predictive and a causal claim.
Why it matters
Continuing from Module 2: that module asked whether the data can support a model at all. This one asks whether the resulting prediction changes anything in clinical practice. Both checks are required — a model built on impeccable data can still be clinically useless, and a well-specified clinical purpose can still be undone by the data problems Module 2 described.
Several failure modes mark out the gap. A deterioration model can fire mainly on patients the ward already knows are sick. A readmission model can identify far more eligible patients than a transitional-care service has capacity to serve. A sepsis alert can arrive no earlier than the clinical suspicion it was meant to precede. In examples like these the modelling can be technically sound while the clinical benefit is small, absent or offset by added workload and alert burden.
Leaders are often asked to judge a proposed change in clinical work, not only a model metric. The questions that follow — what task is this, who acts, when, with what capacity, and what would have happened anyway — are the ones that separate a system that changes care from a dashboard that decorates it. Module 6 will make the evaluation metrics formal; here the discipline is conceptual and it comes first.
Sources & evidence · 8 sources
This module cites scholarly literature, other cited references.
Content reviewed: September 2026. Publication dates of the individual sources are shown in each citation.
Wong A, Otles E, Donnelly JP, et al. External validation of a widely implemented proprietary sepsis prediction model in hospitalized patients. JAMA Intern Med 2021;181(8):1065–1070.
The reference case for the gap between vendor-reported performance, external performance and real alert burden, for the model version and setting studied. A published correction to this article exists; consult the current article record.
Open sourceCaruana R, Koch P, Lou Y, Sturm M, Gehrke J, Elhadad N. Intelligible Models for HealthCare: Predicting Pneumonia Risk and Hospital 30-day Readmission. KDD 2015:1721–1730.
The asthma–pneumonia example: a genuinely predictive pattern that is causally inverted, and dangerous if acted on.
Open sourceThe Predictive Approaches to Treatment effect Heterogeneity (PATH) Statement. Annals of Internal Medicine, 2020. PMID 31711134.
Guidance on using risk prediction within randomised trials to examine heterogeneous treatment effects. Conservative citation: title, journal, year and PMID only.
Open sourceShah NH, Milstein A, Bagley SC. Making machine learning models clinically useful. JAMA 2019;322(14):1351–1352.
The argument that utility, not accuracy, is the unit of value — and that actionability must be assessed before development.
Sendak MP, Ratliff W, Sarro D, et al. Real-world integration of a sepsis deep learning technology into routine clinical care: implementation study. JMIR Med Inform 2020;8(7):e15182.
What integrating a model into routine care actually involves, from workflow to governance.
Open sourceLenert MC, Matheny ME, Walsh CG. Prognostic models will be victims of their own success, unless…. JAMIA 2019;26(12):1645–1650.
The treatment paradox: once a model changes care, the data it generates no longer reflects the untreated world.
Obermeyer Z, Powers B, Vogeli C, Mullainathan S. Dissecting racial bias in an algorithm used to manage the health of populations. Science 2019;366(6464):447–453.
Risk stratification on a cost proxy: how the choice of label reallocates access to care.
Open sourceSperrin M, Martin GP, Pate A, Van Staa T, Peek N, Buchan I. Using marginal structural models to adjust for treatment drop-in when developing clinical prediction models. Statistics in Medicine 2018;37(28):4142–4154.
A worked approach to the interaction between prediction models and the treatments they trigger.
Applied practice: Diagnostic Metrics Playground
Move prevalence, sensitivity and specificity and watch PPV, NPV and the absolute number of false alarms move with them, in a cohort of 1,000 patients.
Capstone Board Pack
Add what you just learned to your own strategy document while it is fresh.