Evidence pack questions
Ten questions, in the order that usually works, for a vendor or data-science team presenting performance figures.
The questions to ask the vendor or data-science team
Ten questions, in the order that usually works. The first three settle what is being claimed; the middle four test whether the evidence supports it here; the last three decide what happens next.
About this aid
An educational review aid for structuring a conversation with a vendor, a data-science team or a project sponsor. It is not a validated instrument and has no regulatory or clinical standing; it does not replace clinical safety assessment, information-governance review or regulatory conformity assessment.
1. What exactly is the intended use, and which decision changes?
- Why it matters
- Everything downstream depends on this. A system without a named decision has no evaluable claim.
- Weak answer
- A description of the technology instead of a decision, or a list of possible uses.
2. In which population was this evaluated, and how does it compare with ours?
- Why it matters
- Case mix, event rate and data capture determine whether the numbers travel.
- Weak answer
- Only a total sample size, with no description of the patients or the setting.
3. What is the operating threshold, and what alert or workload volume does it produce here?
- Why it matters
- Without an operating point there is no workload plan, no PPV and no staffing conversation.
- Weak answer
- “The threshold is configurable”, offered as an answer rather than as a next step.
4. What are the discrimination and calibration results — not just the headline figure?
- Why it matters
- If a displayed risk drives the decision, calibration is as important as ranking.
- Weak answer
- AUC alone, or calibration described as good without a plot or a table.
5. What is the PPV, and what false-positive burden falls on the responding service?
- Why it matters
- This is the number the service experiences daily, and it is a common operational failure point.
- Weak answer
- PPV from a higher-risk population than the one you plan to run on.
6. Has it been evaluated externally or locally, without tuning on the evaluation data?
- Why it matters
- Tuning on the data used to report the result reproduces the developer's optimism.
- Weak answer
- “We recalibrated to your site and report performance on the same records.”
7. What do subgroup results look like, and how many events sit behind each one?
- Why it matters
- Averages hide the groups where performance is unknown, and small groups need honesty rather than reassurance.
- Weak answer
- A significance test offered in place of event counts and intervals.
8. Has the system run prospectively on live data, and what changed when it did?
- Why it matters
- Live data completeness, timing or alert volume can differ materially from retrospective estimates.
- Weak answer
- No silent-mode run, or one whose alert volume is not reported.
9. What comparative evidence supports the claimed workflow or patient benefit?
- Why it matters
- This separates a demonstrated impact from a modelled expectation.
- Weak answer
- Projected benefits derived from retrospective accuracy, presented as results.
10. Is there a named plan for monitoring this after go-live, and who is accountable for reviewing it?
- Why it matters
- A deployment with no monitoring plan cannot detect its own failure. What exactly gets watched, who owns it and what triggers a stop is Module 7's territory; this decision only needs to know a plan exists.
- Weak answer
- Monitoring described as available in the platform, with no plan or owner confirmed.
Used well, this is not an interrogation. Good suppliers answer most of it quickly and tell you plainly which rungs they have not reached — and that answer, given openly, is a better signal of quality than any single number in the pack.
Before you use this
This tool is a prompt for judgement, not a substitute for it. These short sections give you what you need to interpret the answers you get.
The full course behind this
This tool is one step of a longer module in the Practitioner course, where it sits alongside the cases, exercises and evidence that make it usable.
Module 6 — Evaluating Healthcare AIProvenance: this tool restates criteria taught in that module, which cites official EU legal texts, reporting guidelines and peer-reviewed literature. Educational decision aid only, not legal or clinical advice; regulatory points were checked in September 2026.