Guide 06 — AI Evaluation in Healthcare
From model metrics to safety, workflow value and real-world performance.
Four distinct evaluation families, the lifecycle from offline test set to production monitoring, and the pitfalls in between.
A4 sheet — scroll sideways or pinch to zoom
Visual Guide 06 · Evaluation / Foundation / Builder
AI Evaluation in Healthcare
From model metrics to safety, workflow value and real-world performance.
Healthcare AI Learning
A
Predictive / classification models
- Sensitivity · specificity
- PPV / NPV (prevalence-dependent)
- AUROC / AUPRC where appropriate
- Calibration
- Threshold-dependent errors / confusion matrix
Choose metrics from the intended use and the consequences of false positives and false negatives.
B
Generative AI
- Factuality / correctness
- Grounding to supplied sources
- Relevance · completeness
- Robustness to phrasing and context
- Consistency where it matters
- Harmful or unsafe output rate
Grounding does not automatically equal truth: retrieval can be wrong or incomplete.
C
Tools / agents / workflows
- Task success
- Tool-call correctness
- Permission adherence
- Step completion
- Invalid or harmful action rate
- Escalation behaviour
- Latency · cost / resource use
A fluent final answer can hide a failed process or tool call.
D
Real-world clinical & operational value
- Adoption and use
- Time saved · error / rework rate
- Safety incidents
- Workflow burden
- Clinician and patient experience
- Subgroup performance / equity
- Patient or clinical outcomes where relevant
- Financial and operational impact
Good model performance ≠ clinical value.
Evaluation lifecycle
Offline test set
Silent / shadow evaluation where feasible
Controlled pilot
Production monitoring
Incident review, revalidation, change control
Pitfalls
- One aggregate accuracy number
- AUROC without the operating threshold or use case
- Judging generative output only by “looks good”
- No subgroup or error analysis
- No workflow or human-factors evaluation
- Launch without monitoring for drift and change
Evaluation is a system: technical quality + safety + people + workflow + impact.
Related learning
Guides summarise the same primary sources cited in the related courses — official EU legal texts, standards bodies and peer-reviewed literature. Educational summaries only, not legal or clinical advice; regulatory dates were checked between 25 August and 10 September 2026.
Updated Sep 2026