# Healthcare AI Learning Evaluation Guide v0.4

## Purpose

Assess healthcare AI claims, evidence and adoption decisions.

**Core rule:** evidence maturity belongs to a specific claim and intended use, not automatically to a company or product.

## A. Classify the claim(s)

Identify which type(s) of claim are being made:

- Model performance
- Clinical performance
- Workflow / productivity
- Adoption / usability
- Safety
- Clinical outcome
- Economic / ROI
- Regulatory / compliance
- Agentic capability / autonomy

One source may contain multiple claims that need separate evidence profiles.

## B. Source provenance

Record where the evidence comes from:

- Vendor marketing
- Vendor-hosted case study
- Customer / institution report
- Conference / preprint
- Peer-reviewed study
- Independent prospective evaluation
- Controlled / randomized evaluation
- Replicated / systematic evidence

Source type is not an automatic quality score; assess methodology separately.

## C. Intended use

Define: patient/population, user, setting, workflow, input, output, decision/action, and intended benefit. If unclear, say so before assigning strong maturity.

## D. Separate four things

- **CLAIMED** — asserted by the source
- **DEMONSTRATED** — supported by evidence provided
- **PLAUSIBLE BUT NOT DEMONSTRATED** — reasonable, but not shown here
- **UNKNOWN** — information missing

## E. Evidence quality

Assess: patient-level N, samples/encounters, prevalence, representativeness, inclusion/exclusion, missingness, ground truth, labeling, subgroup coverage, retrospective vs prospective design, internal vs external validation, single- vs multi-site, comparator/control, patient-level and temporal separation, predefined endpoints, uncertainty, independence/conflicts.

## F. Metrics

Ask whether the metric is right for the decision.

- **Classification:** sensitivity, specificity, PPV/NPV, calibration, AUPRC.
- **GenAI:** factuality, clinically significant omission, fabrication, completeness, edit rate, task success.
- **Workflow:** time, workload, adoption, override/escalation/incident rates.
- **Economic:** TCO and realized value, not just theoretical savings.

## G. Value chain

Model performance → Decision quality → Workflow change → Action → Outcome → Economic value.

Do not skip links: faster documentation ≠ shorter working day; more capacity ≠ positive ROI; better prediction ≠ better patient outcomes.

## H. Generalizability

Consider sites, populations, countries, languages, devices, workflows, prevalence, specialties and subgroups. External validation ≠ universal generalizability.

## I. Safety & human factors

False positives/negatives, hallucinations, omissions, automation bias, overreliance, alert fatigue, escalation failure, unclear ownership, subgroup harm, drift/change control. Always ask: *What happens when the system is wrong? Who notices?*

## J. Agentic AI (only when relevant)

Assess per use case/action, not per platform:

- Observe
- Reason
- Act
- Access
- Permissions
- Approval
- Escalation
- Reversibility
- Auditability

## K. Regulation & governance

AI Act, MDR/IVDR, GDPR, CE/FDA status, intended-use boundaries, security, oversight, logging, monitoring, updates/change control, accountability.

Regulatory clearance ≠ clinical utility. Security certification ≠ automatic GDPR compliance.

## L. Operational / economic value

Baseline, KPI, effect size, implementation/integration/training/support/monitoring cost, TCO, who pays, who benefits, and whether released capacity is actually realized.

## M. Evidence profile

Score only relevant dimensions (1–5):

1. Little usable evidence
2. Early evidence
3. Credible evidence
4. Strong real-world evidence
5. Strong replicated outcome evidence

Every score MUST include:

- **Why this score?**
- **What would move it up?**

Five should be difficult to earn. This is a structured decision aid, not a validated scientific scoring system.

## N. Named maturity stage

- CLAIM
- EARLY EVIDENCE
- VALIDATED
- REAL-WORLD DEMONSTRATED
- OUTCOME / VALUE DEMONSTRATED

## O. Recommended action

- INSUFFICIENT EVIDENCE
- INVESTIGATE FURTHER
- PILOT CANDIDATE
- CONSIDER ADOPTION

Always explain why, and what must be true before moving to the next stage.

## P. Five questions that matter most

Generate the five highest-information-value questions for this specific case, not a generic checklist.

## Q. Bottom line

In plain language: What is genuinely impressive? What should I be cautious about? What should I do next?

---

**Note:** Marketing is not evidence, but vendor evidence is not automatically invalid. "Not demonstrated" is different from "demonstrated not to work".
