Module 7 · 70 min
Safety, Human Factors & Responsible AI
Designing for Harm, Not Just Accuracy. The safety-case chain from failure mode to hazard, harm, control, monitoring signal, stop rule and owner: human factors and meaningful oversight, differential harm, explainability claims, security read as patient safety, and an applied safety review. Failure modes themselves are recalled from earlier modules rather than re-taught.
- Construct a simple socio-technical safety argument for a healthcare AI use case, running from failure mode through hazard and harm to control, monitoring signal and residual risk.
- Distinguish a model failure from a workflow, human-factors or organisational failure, and recognise that a technically correct output can still contribute to harm.
- Translate a failure mode you already know — predictive, generative, retrieval or agentic — into the hazard and harm it could create in a specific workflow.
- Recognise automation bias, alert fatigue and out-of-the-loop effects, and judge whether a proposed human oversight step is a real control or a formality.
- Identify where unfairness can enter — representation, proxies, labels, access, workflow allocation and differential error burden — and specify how differential harm would be evaluated.
- Judge explainability and uncertainty claims without mistaking a plausible explanation for evidence of correctness, causality or safety.
- Read security weaknesses as routes to patient harm rather than as IT concerns, and say what the safety case has to state about them.
- Organise controls as prevention, detection, mitigation and recovery, and map each control to a specific failure mode rather than a general reassurance.
- Specify a monitoring set with named owners, covering safety events, override behaviour, subgroup signals, data and model change, workflow impact and action errors.
- Write pre-agreed stop and rollback rules that are operationally executable by the people on shift.
Continuing from Module 6
Module 6 asked whether the evidence supports the claim. This module asks a different question: given that the system will sometimes be wrong, and that people and workflows will respond to it in ways nobody designed, what could go wrong here, who would be harmed, and what would catch it in time?
Building on Essentials
Essentials introduced the failure modes: models that are confidently wrong, that drift, that work less well for some groups, that get trusted more than they deserve. Those are taken as known. Practitioner work starts one step later — designing the safety case that turns each failure mode into a hazard, a harm, a control, a monitoring signal, a stop rule and a named owner.
Safety is a system property
Safety work has a reputation as the department of no. Treat it here as the opposite: it is what makes an approvable, defensible scaling decision possible. A team that can state how its system fails, who would be harmed, what catches it and when it stops gives approvers, insurers and clinical councils something they can actually say yes to. Teams that cannot are often held at pilot indefinitely — not because the model was worse, but because nobody could answer the question. The chain in this module is how a system that matters gets a decision rather than a delay.
A model has an error rate. A deployed system has consequences. The distance between those two statements is where most healthcare AI safety work actually lives, and it is not closed by improving the model.
Consider two hospitals running the same sepsis model at the same threshold with identical measured performance. In one, alerts arrive on a shared screen no one owns, during the busiest hour, with no defined response. In the other, they route to a named outreach nurse with protected time and an agreed escalation path. Same model, same accuracy, materially different safety profile — because risk is produced by the interaction of model, data, workflow, users, incentives, environment and the actions the output triggers.
This is why safety cannot be delegated to the vendor or inferred from a validation table. Many important hazards in a deployment are created or amplified locally, by decisions your organisation makes about who sees what, when, with how much time, and with what authority to act.
It also means safety work is not a document produced once. Workflows change, populations change, staffing changes, the model gets updated, and the environment around it moves. A safety argument that was sound in March can be wrong by November without a single line of code changing.
Throughout this module we use one chain: failure mode → hazard → harm → control → monitoring signal → stop rule → owner. Treat it as this course's practical safety-review heuristic. It is informed by lifecycle risk-management guidance — WHO's governance guidance, the NIST AI Risk Management Framework and the FUTURE-AI consensus guideline — but none of those prescribes this six-column format, and completing the chain is not itself evidence that a system is safe or compliant.
Deep dive: risk grids and why their numbers misleadOptional. How to use a likelihood × severity grid without treating its product as a probability.
A methodological caution rather than a measured finding. Most organisations score risk on a qualitative grid — likelihood against severity, red/amber/green. That is a useful way to structure prioritisation and to surface disagreement, and a poor way to produce a number: the ratings are ordinal judgements, usually made without data on how often the failure occurs in your setting, so their product should not be presented as an empirical probability. Use the grid to order the conversation, and do not let a green cell substitute for a control.
One consequence is worth stating plainly, because it contradicts how most business cases are written: evaluation evidence is necessary but not sufficient for safe use. A model can be well-calibrated, externally validated and prospectively tested, and still be unsafe in your hospital because the alert lands where no one can act on it, or because the people receiving it have learned to dismiss it.
The reverse also holds. Good local design does not rescue a model that is wrong for the population, and a strong safety case cannot be built on evidence that does not exist. Module 6 and Module 7 are two halves of one argument.
A safety case has four legitimate outcomes, and only one of them is a stop: proceed; proceed with constraints (a narrower population, a lower autonomy level, a supervised phase); redesign or gather more evidence before deciding; or do not proceed. The purpose of the work is not to find reasons to stop. It is to reach a decision you can defend afterwards — including the decision to go ahead.
Sources & evidence · 13 sources
This module cites public or consensus guidance, scholarly literature.
Content reviewed: September 2026. Publication dates of the individual sources are shown in each citation.
World Health Organization. Ethics and governance of artificial intelligence for health. Geneva: WHO; 2021. ISBN 9789240029200.
WHO guidance setting out ethical principles and governance considerations for health AI, including autonomy, safety, transparency, accountability, equity and responsiveness. A normative guidance document, not a regulation or a technical safety standard.
Open sourceWorld Health Organization. Ethics and governance of artificial intelligence for health: Guidance on large multi-modal models. Geneva: WHO; 2024. ISBN 9789240084759.
Extends the 2021 guidance to large multi-modal and generative models, covering uses, risks and governance across the value chain. Guidance rather than evidence about the performance or safety of any specific system.
Open sourceNational Institute of Standards and Technology. Artificial Intelligence Risk Management Framework (AI RMF 1.0). NIST AI 100-1. Gaithersburg, MD: NIST; 2023. doi:10.6028/NIST.AI.100-1
A voluntary framework organising AI risk work into Govern, Map, Measure and Manage. Useful as a structure for a safety case and for assigning ownership; it is not sector-specific, not mandatory, and does not certify any system.
Open sourceAutio C, Schwartz R, Dunietz J, et al. Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile. NIST AI 600-1. Gaithersburg, MD: NIST; 2024. doi:10.6028/NIST.AI.600-1
A companion profile enumerating risks characteristic of generative AI — including confabulation, information security and data privacy — with suggested actions. A cross-sector risk catalogue, not clinical guidance.
Open sourceGoddard K, Roudsari A, Wyatt JC. Automation bias: a systematic review of frequency, effect mediators, and mitigators. J Am Med Inform Assoc. 2012;19(1):121-127. doi:10.1136/amiajnl-2011-000089
The standard systematic review of automation bias in clinical decision support, covering how often incorrect advice is accepted and which factors moderate it. The included studies vary in design and setting, so treat it as establishing the phenomenon rather than a rate that transfers to your deployment.
Open sourceObermeyer Z, Powers B, Vogeli C, Mullainathan S. Dissecting racial bias in an algorithm used to manage the health of populations. Science. 2019;366(6464):447-453. doi:10.1126/science.aax2342
An empirical dissection of one widely used population-health algorithm whose cost-based target understated need for Black patients at equal illness. Evidence about that algorithm and that proxy choice; it does not establish a general prevalence of bias across healthcare AI.
Open sourceLekadir K, Frangi AF, Porras AR, et al. FUTURE-AI: international consensus guideline for trustworthy and deployable artificial intelligence in healthcare. BMJ. 2025;388:e081554. doi:10.1136/bmj-2024-081554
An international consensus guideline built on six principles — fairness, universality, traceability, usability, robustness and explainability — with 30 lifecycle best practices. Supports the lifecycle framing of safety work; following it is not proof that any specific system is safe, and it does not prescribe the six-column chain used in this module.
Open sourceRay CE, Wilson GM, Hughes AM, et al. Alert fatigue measurement in clinical decision support: a systematic review. J Am Med Inform Assoc. 2026;33(8):1523-1531. doi:10.1093/jamia/ocag064
A review of 22 systematic reviews finding that alert-fatigue measurement is poorly standardised; alert quantity, override and acceptance rates dominate. The authors recommend a significant sustained decrease in appropriate alert response from an established baseline as a more defensible operational measure — which is why override rate alone is treated here as a signal, not a diagnosis.
Open sourceGhassemi M, Oakden-Rayner L, Beam AL. The false hope of current approaches to explainable artificial intelligence in health care. Lancet Digit Health. 2021;3(11):e745-e750. doi:10.1016/S2589-7500(21)00208-9
Argues that current post-hoc explanation methods do not reliably convey why an individual prediction was made and should not substitute for rigorous validation. Supports the boundary drawn in step 5; it is an argued position paper rather than an empirical result about a specific tool.
Open sourceKleinberg J, Mullainathan S, Raghavan M. Inherent Trade-Offs in the Fair Determination of Risk Scores. ITCS 2017. doi:10.4230/LIPIcs.ITCS.2017.43
The formal result behind the fairness trade-off taught in step 4: calibration within groups and balance in error rates cannot generally be satisfied simultaneously when base rates differ, except in special cases such as perfect prediction or equal base rates. A mathematical result, not a statement about any healthcare system.
Open sourceChouldechova A. Fair prediction with disparate impact: a study of bias in recidivism prediction instruments. Big Data. 2017;5(2):153-163. doi:10.1089/big.2016.0047
Complementary evidence for the same incompatibility, showing how a calibrated instrument can still produce unequal error rates across groups when prevalence differs. Developed in criminal justice, not healthcare; the formal argument transfers, the setting does not.
Open sourceGreshake K, Abdelnabi S, Mishra S, et al. Not what you've signed up for: compromising real-world LLM-integrated applications with indirect prompt injection. AISec '23. doi:10.1145/3605764.3623985
Reused from Module 5. Demonstrates that content an application retrieves can act as instruction, which is the mechanism behind the security hazards in step 6. It shows feasibility in general LLM applications; it does not establish incidence in clinical deployments.
Open sourceDebenedetti E, Zhang JZ, Balunović M, et al. AgentDojo: a dynamic environment to evaluate prompt injection attacks and defenses for LLM agents. NeurIPS 2024 Datasets and Benchmarks. arXiv:2406.13352
Reused from Module 5. A benchmark showing that current defences reduce but do not eliminate successful injection against tool-using agents, which is why containment and reversibility, rather than filtering alone, carry the safety argument here.
Open source
Applied practice: AI Safety Incident Lab
Work a near miss in an AI-drafted discharge summary through the safety-case chain: failure mode, hazard, harm, control, monitoring signal, stop rule and owner.
Capstone Board Pack
Add what you just learned to your own strategy document while it is fresh.