Healthcare AI Learning
Back to Cases

Lab · Module 6

Model Evaluation Lab

A vendor has sent an evidence pack for a sepsis-risk model. The clinical director wants your recommendation this week. Your job is not to redo the statistics — it is to work out what the pack does and does not establish, and what you would need next.

Situation

You lead digital transformation at a 600-bed hospital group. A vendor is offering an early-warning model for sepsis on general wards, priced as a three-year subscription with an integration cost in the first year.

The pack is professionally produced and the headline number is genuinely good. The clinical director is enthusiastic; the ward matron is sceptical because the current manual escalation protocol already generates more alerts than the outreach team can answer.

You have five items of evidence in front of you and a decision meeting on Thursday.

Fictional vendor and fictional figures, constructed for teaching. No real product, study or dataset is described.

The evidence in front of you

Headline discrimination

AUC 0.91 for sepsis within 24 hours, reported on a held-out split of the vendor's development data.

Same institutions, same period, same care processes as training. It tells you the model can rank patients on that data. It does not tell you how it behaves at a usable threshold, whether its probabilities are trustworthy, or whether it transfers to your wards.

Calibration

Not reported. The pack shows no calibration plot and no observed-versus-expected table.

Without calibration you cannot say what a stated risk of 20% means in practice, which makes any threshold policy — and any conversation with clinicians about what the number means — guesswork.

Threshold-specific performance

Sensitivity 88% quoted; specificity, PPV and alert volume at that operating point are absent.

Sepsis on a general ward is uncommon, so at a fixed sensitivity the share of alerts that turn out to be real depends heavily on local prevalence. Without the operating point you cannot estimate the workload the model creates.

Comparator

Compared against 'no early warning system'. Your wards already run a structured manual escalation protocol.

The relevant question is incremental value over current practice, not value over nothing. A model that matches your existing protocol adds cost without adding safety.

Prospective and local evidence

No prospective deployment, no external site, no evidence about what clinicians did after an alert.

Every result in the pack is retrospective. Nothing shows the chain from prediction to decision to action to patient outcome under real conditions.

Your decisions

Decision 1. What is the most consequential gap in this pack?

Several things are missing. Which one most limits your ability to decide?

Decision 2. What is the single most useful next evidence step?

You can ask for one thing before Thursday. Choose the one that most reduces uncertainty.

Decision 3. What do you recommend on Thursday?

You can change any answer until you confirm.