Module 7 · 70 min
Safety, Human Factors & Responsible AI
Designing for Harm, Not Just Accuracy. The safety-case chain from failure mode to hazard, harm, control, monitoring signal, stop rule and owner: human factors and meaningful oversight, differential harm, explainability claims, security read as patient safety, and an applied safety review. Failure modes themselves are recalled from earlier modules rather than re-taught.
- Construct a simple socio-technical safety argument for a healthcare AI use case, running from failure mode through hazard and harm to control, monitoring signal and residual risk.
- Distinguish a model failure from a workflow, human-factors or organisational failure, and recognise that a technically correct output can still contribute to harm.
- Translate a failure mode you already know — predictive, generative, retrieval or agentic — into the hazard and harm it could create in a specific workflow.
- Recognise automation bias, alert fatigue and out-of-the-loop effects, and judge whether a proposed human oversight step is a real control or a formality.
- Identify where unfairness can enter — representation, proxies, labels, access, workflow allocation and differential error burden — and specify how differential harm would be evaluated.
- Judge explainability and uncertainty claims without mistaking a plausible explanation for evidence of correctness, causality or safety.
- Read security weaknesses as routes to patient harm rather than as IT concerns, and say what the safety case has to state about them.
- Organise controls as prevention, detection, mitigation and recovery, and map each control to a specific failure mode rather than a general reassurance.
- Specify a monitoring set with named owners, covering safety events, override behaviour, subgroup signals, data and model change, workflow impact and action errors.
- Write pre-agreed stop and rollback rules that are operationally executable by the people on shift.
Human factors & automation bias
Once a system is in use, the relevant unit of analysis is not the model or the clinician but the pair. Human factors is the study of how that pairing behaves under real conditions: interruptions, time pressure, fatigue, shift handover, competing tasks and a screen that has to be dealt with before the next patient.
The systematic review evidence on automation bias in clinical decision support is now old enough to be sobering. Goddard and colleagues (2012) found that clinicians accept incorrect automated advice at measurable rates, and that the effect is moderated by factors including experience, task difficulty, confidence and workload. Treat it as a design constraint, not a training problem.
Previously, in Module 3
Module 3 introduced human oversight as a condition of clinical utility. Here the question is sharper: under what conditions is that oversight step a control at all, and how would you tell if it had quietly stopped being one?
Automation bias (commission)
Acting on an incorrect recommendation the user would not have arrived at unaided.
Accepting a drafted dose or a suggested code because it appeared in the field, particularly when the system is usually right.
Omission and complacency
Failing to notice a problem the system did not flag, because the absence of a flag reads as reassurance.
A deteriorating patient whose score stayed below threshold receives less scrutiny than they would have before the system existed.
Alert fatigue
High alert burden with low relevance can erode appropriate responses over time, including responses to the alerts that matter.
A ward that dismisses most alerts within seconds is a warning signal rather than a diagnosis. Measurement of alert fatigue is poorly standardised (Ray et al., 2026); override or dismissal rate alone does not establish it, and a sustained fall in appropriate response from an established baseline, read alongside alert burden and relevance, is the stronger signal.
Out-of-the-loop performance
Supervising an automated process degrades situational awareness relative to doing the task.
A clinician who edits summaries rather than composing them may hold a thinner picture of the admission when a question arises.
Workload and time compression
Efficiency gains are frequently reabsorbed as throughput rather than as review time.
If the review step depends on time that the deployment's own business case has already spent, the control is unfunded.
Deskilling (plausible, less well quantified)
Reduced practice of a skill over time may reduce competence at it, including the competence needed to catch the system's errors.
A concern worth monitoring rather than a settled finding; treat it as a reason to preserve unaided practice, not as an established rate.
Six conditions for oversight to be a control
Time
Is there realistically enough time in this workflow to review this output properly, and is that time funded and protected?
Information
Can the reviewer see what the output was based on — the source passage, the inputs, the provenance — without leaving the task?
Authority
Can this person override or refuse without escalation, career risk or a workaround? Is the override path faster than compliance?
Competence
Does the reviewer have the expertise to detect this specific class of error, including plausible-looking ones and omissions?
Escalation path
When the reviewer disagrees or is unsure, is there a named, available route that resolves it within the clinical timeframe?
Feedback
Does the reviewer ever learn whether their overrides and acceptances were right? Without feedback, calibration cannot develop.
Exercise: real control or formality?
Six claimed oversight arrangements. Using the six conditions — time, information, authority, competence, escalation path and feedback — judge whether each is a control the safety case can rely on, or a formality. Attempt all six to continue.
1. Sign-off at the end of the list
Doctors approve drafted summaries in a batch at the end of the discharge round, with the source record on another screen.
2. Protected review slot, source alongside
Ten minutes are rostered per discharge, the draft shows medications and results side by side, and rejecting a draft costs one click.
3. Approval logged with name and timestamp
Every approval records the clinician's identity, the timestamp and the version approved, for later audit.
4. Override requires a written justification
Clinicians may reject the recommendation, but must complete a free-text form that is reviewed by the specialty lead.
5. Pharmacist checks the medication section
A ward pharmacist verifies the medication section of each summary against the chart before it is released to primary care.
6. Annual training on system limitations
All users complete a training module covering the tool's known limitations before being given access, refreshed each year.
Calibrated reliance
The goal is calibrated reliance: trust that tracks the system's actual reliability for this task, in this population, at this moment. Maximal trust produces automation bias. Maximal distrust produces an expensive system that nobody uses and a workflow that has quietly added a step.
Calibration is supported by design, not by exhortation. Show the basis for the output. Show uncertainty where it is meaningful. Make the system visibly abstain when it should. Route only alerts that carry an available action. Give reviewers feedback on what happened after their decisions.
And write the oversight step honestly in the safety case. 'A clinician reviews all outputs before signing' is a control only if the six conditions above hold. Where they do not, either fix the conditions or stop claiming the control.
Sources & evidence · 13 sources
This module cites public or consensus guidance, scholarly literature.
Content reviewed: September 2026. Publication dates of the individual sources are shown in each citation.
World Health Organization. Ethics and governance of artificial intelligence for health. Geneva: WHO; 2021. ISBN 9789240029200.
WHO guidance setting out ethical principles and governance considerations for health AI, including autonomy, safety, transparency, accountability, equity and responsiveness. A normative guidance document, not a regulation or a technical safety standard.
Open sourceWorld Health Organization. Ethics and governance of artificial intelligence for health: Guidance on large multi-modal models. Geneva: WHO; 2024. ISBN 9789240084759.
Extends the 2021 guidance to large multi-modal and generative models, covering uses, risks and governance across the value chain. Guidance rather than evidence about the performance or safety of any specific system.
Open sourceNational Institute of Standards and Technology. Artificial Intelligence Risk Management Framework (AI RMF 1.0). NIST AI 100-1. Gaithersburg, MD: NIST; 2023. doi:10.6028/NIST.AI.100-1
A voluntary framework organising AI risk work into Govern, Map, Measure and Manage. Useful as a structure for a safety case and for assigning ownership; it is not sector-specific, not mandatory, and does not certify any system.
Open sourceAutio C, Schwartz R, Dunietz J, et al. Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile. NIST AI 600-1. Gaithersburg, MD: NIST; 2024. doi:10.6028/NIST.AI.600-1
A companion profile enumerating risks characteristic of generative AI — including confabulation, information security and data privacy — with suggested actions. A cross-sector risk catalogue, not clinical guidance.
Open sourceGoddard K, Roudsari A, Wyatt JC. Automation bias: a systematic review of frequency, effect mediators, and mitigators. J Am Med Inform Assoc. 2012;19(1):121-127. doi:10.1136/amiajnl-2011-000089
The standard systematic review of automation bias in clinical decision support, covering how often incorrect advice is accepted and which factors moderate it. The included studies vary in design and setting, so treat it as establishing the phenomenon rather than a rate that transfers to your deployment.
Open sourceObermeyer Z, Powers B, Vogeli C, Mullainathan S. Dissecting racial bias in an algorithm used to manage the health of populations. Science. 2019;366(6464):447-453. doi:10.1126/science.aax2342
An empirical dissection of one widely used population-health algorithm whose cost-based target understated need for Black patients at equal illness. Evidence about that algorithm and that proxy choice; it does not establish a general prevalence of bias across healthcare AI.
Open sourceLekadir K, Frangi AF, Porras AR, et al. FUTURE-AI: international consensus guideline for trustworthy and deployable artificial intelligence in healthcare. BMJ. 2025;388:e081554. doi:10.1136/bmj-2024-081554
An international consensus guideline built on six principles — fairness, universality, traceability, usability, robustness and explainability — with 30 lifecycle best practices. Supports the lifecycle framing of safety work; following it is not proof that any specific system is safe, and it does not prescribe the six-column chain used in this module.
Open sourceRay CE, Wilson GM, Hughes AM, et al. Alert fatigue measurement in clinical decision support: a systematic review. J Am Med Inform Assoc. 2026;33(8):1523-1531. doi:10.1093/jamia/ocag064
A review of 22 systematic reviews finding that alert-fatigue measurement is poorly standardised; alert quantity, override and acceptance rates dominate. The authors recommend a significant sustained decrease in appropriate alert response from an established baseline as a more defensible operational measure — which is why override rate alone is treated here as a signal, not a diagnosis.
Open sourceGhassemi M, Oakden-Rayner L, Beam AL. The false hope of current approaches to explainable artificial intelligence in health care. Lancet Digit Health. 2021;3(11):e745-e750. doi:10.1016/S2589-7500(21)00208-9
Argues that current post-hoc explanation methods do not reliably convey why an individual prediction was made and should not substitute for rigorous validation. Supports the boundary drawn in step 5; it is an argued position paper rather than an empirical result about a specific tool.
Open sourceKleinberg J, Mullainathan S, Raghavan M. Inherent Trade-Offs in the Fair Determination of Risk Scores. ITCS 2017. doi:10.4230/LIPIcs.ITCS.2017.43
The formal result behind the fairness trade-off taught in step 4: calibration within groups and balance in error rates cannot generally be satisfied simultaneously when base rates differ, except in special cases such as perfect prediction or equal base rates. A mathematical result, not a statement about any healthcare system.
Open sourceChouldechova A. Fair prediction with disparate impact: a study of bias in recidivism prediction instruments. Big Data. 2017;5(2):153-163. doi:10.1089/big.2016.0047
Complementary evidence for the same incompatibility, showing how a calibrated instrument can still produce unequal error rates across groups when prevalence differs. Developed in criminal justice, not healthcare; the formal argument transfers, the setting does not.
Open sourceGreshake K, Abdelnabi S, Mishra S, et al. Not what you've signed up for: compromising real-world LLM-integrated applications with indirect prompt injection. AISec '23. doi:10.1145/3605764.3623985
Reused from Module 5. Demonstrates that content an application retrieves can act as instruction, which is the mechanism behind the security hazards in step 6. It shows feasibility in general LLM applications; it does not establish incidence in clinical deployments.
Open sourceDebenedetti E, Zhang JZ, Balunović M, et al. AgentDojo: a dynamic environment to evaluate prompt injection attacks and defenses for LLM agents. NeurIPS 2024 Datasets and Benchmarks. arXiv:2406.13352
Reused from Module 5. A benchmark showing that current defences reduce but do not eliminate successful injection against tool-using agents, which is why containment and reversibility, rather than filtering alone, carry the safety argument here.
Open source
Applied practice: AI Safety Incident Lab
Work a near miss in an AI-drafted discharge summary through the safety-case chain: failure mode, hazard, harm, control, monitoring signal, stop rule and owner.
Capstone Board Pack
Add what you just learned to your own strategy document while it is fresh.