Healthcare AI Learning
Course overview

Module 3 · 85 min

Clinical AI: From Prediction to Clinical Utility

Diagnosis, prognosis, risk stratification, monitoring and treatment response; precise intended use, the chain from prediction to decision to action to outcome, threshold trade-offs, causal versus predictive claims, and the clinical evidence that actually counts.

Lesson progress
0 / 12 steps
Not started
Learning objectives
  • Distinguish diagnosis, prognosis, risk stratification, monitoring and treatment-response tasks, and say why the distinction changes the evidence you need.
  • Write a precise intended-use statement: population, index moment, output, user, decision and care setting.
  • Trace the full chain from prediction to decision to action to outcome, and locate the weakest link.
  • Judge whether a prediction is actionable at all, given the capacity, timing and available intervention.
  • Reason qualitatively about operating thresholds as an explicit trade of one kind of harm against another.
  • Separate a prediction claim from a treatment-effect claim, and recognise when a predictive model is being used as if it were causal.
  • Ask for the right level of clinical evidence — internal, external, prospective, live and comparative — and know what each does and does not establish.
  • Apply a reusable ten-item executive review to any proposed clinical AI system.

Executive Clinical AI Review Checklist

Ten questions to take into any clinical AI proposal, in any specialty. They are deliberately answerable by a non-technical executive, and they are ordered so that the earlier questions define whether the later evidence is even relevant — later safety and governance findings still matter in their own right.

Educational governance checklist — not a validated clinical, regulatory or procurement instrument.

This is a review checklist, not a scored exercise. For each question, think what a strong answer would sound like, then reveal the guidance.

  1. 1. What clinical task is this — diagnosis, prognosis, risk stratification, monitoring or treatment response?

  2. 2. Can you read me the intended-use statement: population, index moment, output, horizon, user, action, setting and exclusions?

  3. 3. Who acts on the output, at what moment, and what exactly do they do differently?

  4. 4. Does the action exist, and is there capacity to deliver it at the volume the threshold implies — including nights and weekends?

  5. 5. What is the incremental value over current practice and over the simplest reasonable baseline?

  6. 6. Where has this been validated — internally, externally, prospectively, live — and in what population?

  7. 7. How was the threshold chosen, who owns it, and what are the two harms it trades off?

  8. 8. Is any causal claim being made, and what evidence supports it?

  9. 9. How will this fail, how would we notice, and who is accountable when it does?

  10. 10. What outcomes will we measure after deployment, and what would make us stop?

Sources & evidence · 8 sources

This module cites scholarly literature, other cited references.

Content reviewed: September 2026. Publication dates of the individual sources are shown in each citation.

  • Wong A, Otles E, Donnelly JP, et al. External validation of a widely implemented proprietary sepsis prediction model in hospitalized patients. JAMA Intern Med 2021;181(8):1065–1070.

    The reference case for the gap between vendor-reported performance, external performance and real alert burden, for the model version and setting studied. A published correction to this article exists; consult the current article record.

    Open source
  • Caruana R, Koch P, Lou Y, Sturm M, Gehrke J, Elhadad N. Intelligible Models for HealthCare: Predicting Pneumonia Risk and Hospital 30-day Readmission. KDD 2015:1721–1730.

    The asthma–pneumonia example: a genuinely predictive pattern that is causally inverted, and dangerous if acted on.

    Open source
  • The Predictive Approaches to Treatment effect Heterogeneity (PATH) Statement. Annals of Internal Medicine, 2020. PMID 31711134.

    Guidance on using risk prediction within randomised trials to examine heterogeneous treatment effects. Conservative citation: title, journal, year and PMID only.

    Open source
  • Shah NH, Milstein A, Bagley SC. Making machine learning models clinically useful. JAMA 2019;322(14):1351–1352.

    The argument that utility, not accuracy, is the unit of value — and that actionability must be assessed before development.

  • Sendak MP, Ratliff W, Sarro D, et al. Real-world integration of a sepsis deep learning technology into routine clinical care: implementation study. JMIR Med Inform 2020;8(7):e15182.

    What integrating a model into routine care actually involves, from workflow to governance.

    Open source
  • Lenert MC, Matheny ME, Walsh CG. Prognostic models will be victims of their own success, unless…. JAMIA 2019;26(12):1645–1650.

    The treatment paradox: once a model changes care, the data it generates no longer reflects the untreated world.

  • Obermeyer Z, Powers B, Vogeli C, Mullainathan S. Dissecting racial bias in an algorithm used to manage the health of populations. Science 2019;366(6464):447–453.

    Risk stratification on a cost proxy: how the choice of label reallocates access to care.

    Open source
  • Sperrin M, Martin GP, Pate A, Van Staa T, Peek N, Buchan I. Using marginal structural models to adjust for treatment drop-in when developing clinical prediction models. Statistics in Medicine 2018;37(28):4142–4154.

    A worked approach to the interaction between prediction models and the treatments they trigger.

Applied practice: Diagnostic Metrics Playground

Move prevalence, sensitivity and specificity and watch PPV, NPV and the absolute number of false alarms move with them, in a cohort of 1,000 patients.

Open lab

Capstone Board Pack

Add what you just learned to your own strategy document while it is fresh.

Open capstone