Module 4 · 100 min · 3 chapters
Evaluation, Security & Implementation
Three chapters. Evaluating a trajectory rather than a paragraph: task success, tool selection, action accuracy, recovery, latency, cost and drift. The threat surface that appears once a system can act: prompt injection, excessive agency, exfiltration, memory poisoning, the safety case, audit trail and regulatory routing by intended use. And getting to a running service: sandboxed starts, shadow mode, pilot gates, operating model, sourcing, TCO, incidents and decommissioning.
Case-led opening · fictional healthcare service
Judge the path, not the paragraph.
Northlake's referral agent produced a polished update: 'The referral is complete and the practice has been contacted.' It sounded exactly right. The trace told a different story: the agent had read an old referral copy, called the messaging tool twice after a timeout and never verified delivery.
From Module 2 you already know what a trace captures. Here, a trace becomes evaluation evidence. The question is not simply whether the final words look useful, but whether the path to them was accurate, bounded and recoverable.
Fictional educational case. The Northlake Referral Coordination Agent and its traces are invented for teaching. No benchmark, safety rate or operational outcome is claimed.
Where value can appear
Trust the path, not the polish
A fluent final answer can conceal a wrong read, a wrong target or a missing side effect.
Find failure before exposure
Scenario-based evaluation makes weak boundaries visible while the service is still changeable.
Keep working after launch
Production behaviour can move when models, tools, policies, data or workflows change.
Make improvement attributable
Versioned evidence tells a team whether a change helped, harmed or simply changed the path.
Opportunity first, boundaries explicit
This module narrows the system to the work it can support. The value is operational; clinical judgement, diagnosis, treatment and interpretation of new symptoms remain with qualified humans wherever the workflow touches them.
Sources & evidence · 11 sources
This module cites primary legal, public or consensus guidance, technical documentation, vendor documentation.
Content reviewed: September 2026. Publication dates of the individual sources are shown in each citation.
NIST AI 600-1, AI Risk Management Framework: Generative AI Profile.
Risk-management framing for measurement, monitoring and managing generative AI risks. It does not provide a universal agent benchmark.
Open sourceNIST AI 100-2e2025, Adversarial Machine Learning: A Taxonomy and Terminology of Attacks and Mitigations, March 2025.
Current edition of the recognised terminology for adversarial threats and mitigations; useful when designing scenarios beyond the happy path. Supersedes the earlier e2023 edition.
Open sourceOWASP. GenAI Security Project — LLM application security guidance (moving community resource).
Security risks and mitigations relevant to tool-using LLM applications. Community guidance, not a validation protocol. Continuously updated rather than versioned; checked 10 September 2026.
Open sourceAnthropic, Building effective agents (19 December 2024).
A vendor perspective on evaluating agentic workflows and keeping architectures understandable; not evidence that a particular system is safe.
Open sourceOWASP GenAI Security Project. OWASP Top 10 for LLM and Generative AI Applications, 2026 edition (published 3 August 2026).
Community guidance on prompt injection, excessive agency, sensitive information disclosure and risks in tool-using applications. Edition-pinned; checked 10 September 2026. Not peer-reviewed evidence and not a certification scheme.
Open sourceEuropean Commission. Regulatory framework for AI — overview and application timeline.
Moving official overview page of the AI Act's risk-based structure and application timeline; checked 10 September 2026. Secondary to EUR-Lex — rely on the consolidated legal text for obligations and dates.
Open sourceRegulation (EU) 2016/679 (General Data Protection Regulation).
Authoritative source for personal and health data processing, minimisation, roles and impact-assessment questions.
Open sourceEuropean Commission, MDCG 2025-6 / AIB 2025-1: FAQ on MDR/IVDR and the AI Act.
Official guidance on interplay between medical-device routes and the AI Act. It does not classify an individual product.
Open sourceNIST AI Risk Management Framework 1.0 and Generative AI Profile (AI 600-1).
Recognised framework for governing, mapping, measuring and managing AI risk across the lifecycle.
Open sourceNHS England. Medical devices and digital tools (guidance hub).
NHS England guidance on assessing and assuring digital health tools; the earlier DTAC pages have been reorganised, so check the current hub. Local governance remains necessary.
Open sourceWHO. Ethics and governance of artificial intelligence for health (2021).
International principles for responsible AI in health, including accountability, inclusiveness and human control.
Open source