A pathology AI evaluation should answer a specific question about a defined workflow, user group, input data, and intended output. A result from one laboratory is not automatically evidence that the same software will perform consistently in another laboratory. Site, scanner, staining, case mix, and workflow differences all need to be considered.
Start with intended use
Write the intended use before selecting metrics. State whether the software is for research, quality review, triage, measurement, or clinician decision support. Define the intended users, specimen and image types, exclusions, output, and the decision the output may inform. An educational evaluation does not establish clinical safety, regulatory status, or permission to use software for diagnosis.
Characterize every site
Record scanner models, magnification, image formats, staining protocols, tissue preparation, storage conditions, case mix, prevalence, and missing or rejected images. Document how cases were selected and whether any site contributed to development or tuning. These details help distinguish a model limitation from a data or workflow difference.
Measure the human-AI workflow
Model performance is only one part of a clinical workflow evaluation. Measure how intended users interpret the output, whether it changes review time or decisions, how often users override it, and what happens when the output is unavailable or outside its validated scope. A concurrent decision-support use still requires professional judgment and review of the underlying case.
Plan monitoring before launch
Set a review cadence for data drift, input quality, subgroup performance, error reports, overrides, and changes in scanners or staining. Assign owners for investigation, escalation, communication, and rollback. Monitoring should continue through the software lifecycle rather than ending when an initial evaluation is complete.
The practical conclusion is simple: a multi-laboratory evaluation is a documented human-AI and workflow study, not a single accuracy number. Treat the output as decision support unless a specific intended use and applicable authorization say otherwise, and keep a pathologist or other qualified professional responsible for the clinical decision.
Sources
- Good Machine Learning Practice for Medical Device Development: Guiding Principles (U.S. Food and Drug Administration)
- Artificial Intelligence Risk Management Framework (National Institute of Standards and Technology)
- Ethics and governance of artificial intelligence for health (World Health Organization)
Written by
Digital Pathology Solutions Editorial Team
Medical AI and digital pathology




