Evaluation Frameworks for Clinical AI Incorporating Validation Strategies, Real-World Applicability, and Ethical Principles: Scoping Review
Clinical Snapshot
PICO Framework
| P — Population | Existing evaluation and reporting frameworks for clinical artificial intelligence systems, sourced from peer-reviewed literature, grey literature, and organisational guidelines |
| I — Intervention | Systematic mapping and critical analysis of framework characteristics across three domains: methodological rigour, validation strategies (internal and external), and ethical integration using a UNESCO-based scoring matrix |
| C — Comparator | No formal comparator; frameworks compared descriptively against each other and against UNESCO AI ethical principles across 10 domains |
| O — Outcomes | Characterisation of framework structure, methodological rigour, validation alignment with intended clinical use, and degree of ethical integration; identification of gaps via a dot plot-based gap map |
Bottom Line
This scoping review maps 46 clinical AI evaluation frameworks identified from 3,363 records, revealing a rapidly expanding but deeply fragmented landscape. The central finding is sobering: only 11.4% of frameworks achieve full methodological rigour with validation aligned to intended clinical use, and fewer than 11% demonstrate high ethical compliance against UNESCO principles. Most frameworks remain oriented toward investigational rather than clinical implementation contexts, with critical gaps in external validation, real-world applicability assessment, and human oversight provisions. For senior clinicians and health system leaders, this review confirms what many have suspected: the proliferation of AI tools in healthcare is outpacing the development of robust, standardised evaluation infrastructure. The absence of validated, universally endorsed evaluation frameworks creates genuine patient safety risk when AI tools are deployed without adequate clinical validation. While the scoping methodology precludes certainty grading or practice-changing recommendations, the findings provide a structured evidence base for institutional AI governance committees, regulators including the TGA, and clinical leaders to demand more rigorous validation evidence from AI developers. Institutions should not accept investigational-grade validation as sufficient for clinical deployment. This review is a useful but preliminary contribution; a systematic review with quality appraisal and meta-analysis of framework performance outcomes is now warranted.
Key Findings
P Value: Not reported; descriptive synthesis only
Effect Size: Not applicable (scoping review; descriptive proportional findings): 88% of frameworks targeted investigational use; 11.4% achieved full methodological rigour; 31.8% reported technical performance metrics; 15.9% reported clinical performance indicators; 5 of 46 frameworks (10.9%) achieved ≥80% UNESCO ethical compliance; 4 frameworks scored <10% on ethical compliance
Primary Outcome: Mapping and characterisation of 46 clinical AI evaluation frameworks across methodological rigour, validation strategies, and UNESCO ethical alignment, revealing a fragmented and predominantly investigational-use landscape with critical gaps in clinical validation and ethical integration
Nnt Or Sensitivity: Not applicable to scoping review design. Within included frameworks: area under the curve, sensitivity, and specificity were the most commonly reported technical metrics (31.8% of frameworks); predictive values and calibration metrics were reported in only 15.9% of frameworks. Most frequently addressed UNESCO principles: awareness and education (71.1%) and transparency and explainability (70%). Least represented: human oversight (24.4%) and adaptive governance (33.3%)
Confidence Interval: Not reported; no inferential statistics presented
Clinical Application
The review's findings are directly actionable for institutions developing internal AI governance frameworks. The UNESCO-based scoring matrix, if validated, could serve as a practical audit tool for ethics committees and digital health governance boards. However, the absence of a validated, standardised framework emerging from this review means clinicians and institutions must continue to navigate a fragmented landscape without a single endorsed standard Highly relevant to the Australian healthcare context. The Therapeutic Goods Administration (TGA) regulates software-as-a-medical-device (SaMD) under the Medical Devices framework, and the Australian Digital Health Agency (ADHA) has published AI in Health guidance. The finding that only 11.4% of reviewed frameworks achieve full methodological rigour, and that human oversight is addressed in only 24.4% of frameworks, directly informs TGA SaMD conformity assessment requirements and RACGP guidance on AI-assisted clinical decision support. The Australian Government's AI Ethics Framework (2019) and the National Health and Medical Research Council (NHMRC) ethical guidelines for AI in health align with the UNESCO principles assessed in this review. PBS and MBS implications are indirect but significant: AI tools influencing prescribing or diagnostic decisions require robust validation frameworks before integration into reimbursed care pathways. Australian health services considering AI procurement should use this review's gap map to interrogate vendor-supplied validation evidence against the dimensions identified as most frequently absent. Health informaticians, clinical AI developers, hospital digital health governance committees, health regulators, and clinical leaders responsible for evaluating, procuring, or implementing AI-based clinical decision support tools across any clinical specialty
Abstract
BACKGROUND: AI shows substantial potential in health care; however, the absence of standardized evaluation frameworks limits its safe and effective clinical implementation because of inconsistent validation requirements and fragmented ethical principles. Existing guidelines vary in structure, methodological rigor, and ethical integration, creating uncertainty. OBJECTIVE: This study aimed to systematically map, characterize, and critically analyze existing evaluation frameworks for clinical AI, focusing on three core dimensions: methodological rigor, validation strategies (internal validation, including reporting of technical and clinical performance; external validation, including real-world applicability), and alignment with the United Nations Educational, Scientific and Cultural Organization (UNESCO) AI ethical considerations. METHODS: A scoping review was conducted following PRISMA-ScR (Preferred Reporting Items for Systematic Reviews and Meta-Analyses extension for Scoping Reviews) guidelines. Six databases (PubMed, Embase, BVS, EBSCOhost, ProQuest, and Sage) and the Enhancing the Quality and Transparency of Health Research Network were searched without language or date restrictions up to February 2026. Eligible documents included peer-reviewed papers, gray literature, and organizational guidelines describing evaluation or reporting frameworks for clinical AI. Editorials, commentaries, and conference abstracts lacking a clearly defined evaluative framework or clinical applicability were excluded. Two reviewers independently screened records and extracted data. Data were extracted across three domains: (1) general characteristics, (2) methodological rigor and validation parameters, and (3) ethical integration and were synthesized using a dot plot-based gap map. Ethical adherence was assessed using a 10-domain UNESCO-based scoring matrix. No formal risk-of-bias assessment was conducted, consistent with scoping review methodology. RESULTS: From 3363 records, 46 frameworks met the inclusion criteria. Mapping revealed a rapidly expanding but fragmented landscape. Most frameworks targeted investigational use (88%), with limited focus on clinical applicability. Frameworks varied in structure, methodology, and scope, with a predominance of reporting guidelines and few validated tools. Most (63%) were developed through multi-institutional collaborations, and 32.6% incorporated transdisciplinary participation. Only 31.8% reported technical metrics (commonly area under the curve, sensitivity, and specificity), and 15.9% provided clinical indicators (eg, predictive values or calibration). Only 11.4% achieved methodological rigor, incorporating validation aligned with intended use, while most relied on partial validation strategies, highlighting a gap between model development and clinical evaluation. Ethical integration was heterogeneous: only 5 frameworks achieved high compliance (≥80%), whereas 4 scored <10%. The most frequently addressed UNESCO principles were awareness and education (71.1%) and transparency and explainability (70%), while human oversight (24.4%) and adaptive governance (33.3%) were least represented. Findings indicate a misalignment between framework design, validation requirements, and clinical implementation. CONCLUSIONS: Evaluation frameworks for clinical AI remain heterogeneous and oriented toward investigational contexts. Critical gaps persist in methodological rigor, validation aligned with intended use, and fragmented ethical coverage. These findings highlight the need for standardized, robust, and ethically grounded frameworks to enable safe, reliable, and scalable integration of AI into clinical practice.
References
- 1.López Medina, D. C., Oliveros-Navarro, A., Moreno Angel, N., Herrera-Arellano, A. C., & Henao-Pérez, M. (2026). Evaluation frameworks for clinical AI incorporating validation strategies, real-world applicability, and ethical principles: Scoping review. Journal of Medical Internet Research. https://doi.org/10.2196/78168
Related Research
Journal of critical care
Artificial intelligence and computerized decision support in adult intensive care: A systematic review of randomized controlled trials
3 Aug 2026
Simulation in healthcare : journal of the Society for Simulation in Healthcare
Future of Biometric Technology in Healthcare Simulation
2 Aug 2026
Rheumatic diseases clinics of North America
Demystifying Artificial Intelligence: Key Concepts with Examples in Rheumatology
2 Aug 2026
This content is for educational purposes for healthcare professionals only and does not constitute clinical advice. Clinical decisions should be based on individual patient assessment, current guidelines, and appropriate specialist consultation. Editorial Standards · Privacy Policy · Terms of Service