Research Appraisals
Evidence-based critical appraisals of the latest medical research, systematically evaluated using Oxford CEBM methodology.
Showing 2 appraisals
Journal of medical Internet research
Evaluation Frameworks for Clinical AI Incorporating Validation Strategies, Real-World Applicability, and Ethical Principles: Scoping Review
BACKGROUND: AI shows substantial potential in health care; however, the absence of standardized evaluation frameworks limits its safe and effective clinical implementation because of inconsistent validation requirements and fragmented ethical principles. Existing guidelines vary in structure, methodological rigor, and ethical integration, creating uncertainty. OBJECTIVE: This study aimed to systematically map, characterize, and critically analyze existing evaluation frameworks for clinical AI, focusing on three core dimensions: methodological rigor, validation strategies (internal validation, including reporting of technical and clinical performance; external validation, including real-world applicability), and alignment with the United Nations Educational, Scientific and Cultural Organization (UNESCO) AI ethical considerations. METHODS: A scoping review was conducted following PRISMA-ScR (Preferred Reporting Items for Systematic Reviews and Meta-Analyses extension for Scoping Reviews) guidelines. Six databases (PubMed, Embase, BVS, EBSCOhost, ProQuest, and Sage) and the Enhancing the Quality and Transparency of Health Research Network were searched without language or date restrictions up to February 2026. Eligible documents included peer-reviewed papers, gray literature, and organizational guidelines describing evaluation or reporting frameworks for clinical AI. Editorials, commentaries, and conference abstracts lacking a clearly defined evaluative framework or clinical applicability were excluded. Two reviewers independently screened records and extracted data. Data were extracted across three domains: (1) general characteristics, (2) methodological rigor and validation parameters, and (3) ethical integration and were synthesized using a dot plot-based gap map. Ethical adherence was assessed using a 10-domain UNESCO-based scoring matrix. No formal risk-of-bias assessment was conducted, consistent with scoping review methodology. RESULTS: From 3363 records, 46 frameworks met the inclusion criteria. Mapping revealed a rapidly expanding but fragmented landscape. Most frameworks targeted investigational use (88%), with limited focus on clinical applicability. Frameworks varied in structure, methodology, and scope, with a predominance of reporting guidelines and few validated tools. Most (63%) were developed through multi-institutional collaborations, and 32.6% incorporated transdisciplinary participation. Only 31.8% reported technical metrics (commonly area under the curve, sensitivity, and specificity), and 15.9% provided clinical indicators (eg, predictive values or calibration). Only 11.4% achieved methodological rigor, incorporating validation aligned with intended use, while most relied on partial validation strategies, highlighting a gap between model development and clinical evaluation. Ethical integration was heterogeneous: only 5 frameworks achieved high compliance (≥80%), whereas 4 scored <10%. The most frequently addressed UNESCO principles were awareness and education (71.1%) and transparency and explainability (70%), while human oversight (24.4%) and adaptive governance (33.3%) were least represented. Findings indicate a misalignment between framework design, validation requirements, and clinical implementation. CONCLUSIONS: Evaluation frameworks for clinical AI remain heterogeneous and oriented toward investigational contexts. Critical gaps persist in methodological rigor, validation aligned with intended use, and fragmented ethical coverage. These findings highlight the need for standardized, robust, and ethically grounded frameworks to enable safe, reliable, and scalable integration of AI into clinical practice.
24 July 2026
Read appraisal →Philosophy, ethics, and humanities in medicine : PEHM
Moral diversity and the challenge of responsibility in AI-CDSS
The increasing integration of artificial intelligence in clinical decision support systems (AI-CDSS) has fueled expectations of more personalized and effective diagnostics and therapies. By incorporating machine learning methods, AI-CDSS promise enhanced predictive accuracy, improved stratification, and innovative individualized care. However, this technological optimism is accompanied by complex ethical challenges, including issues of explainability, trust, autonomy, and data security. At the core of these debates lies the question of responsibility, which involves both its attribution and diffusion, as well as the underlying normative standards guiding moral action. In the context of healthcare practice, responsibility is further complicated by moral diversity-the coexistence of varying moral values, cultural beliefs, and ethical frameworks among healthcare professionals, patients, and institutional stakeholders. This plurality challenges the establishment of a unified normative standard necessary for ethically sound responsibility attribution. This paper offers an analysis of moral diversity and AI-CDSS as a challenge for responsibility in healthcare environments. Using a relational concept of responsibility the study examines key areas in which moral diversity affects responsibility in AI-mediated decision-making. This includes algorithmic bias, healthcare professional and patient interaction and the role of patients. Through these examples, the paper explains how different normative standards intensify ethical complexity in AI-supported clinical contexts. It argues that greater ethical sensitivity to moral diversity is essential-both in the development of AI-CDSS and in their application within morally value-laden healthcare situations.
4 July 2026
Read appraisal →