Research Appraisals
Evidence-based critical appraisals of the latest medical research, systematically evaluated using Oxford CEBM methodology.
Showing 2 appraisals
Journal of medical Internet research
Shadow AI in Swedish Health Care: Qualitative Analysis of Physicians' Free-Text Answers
BACKGROUND: The rapid emergence of artificial intelligence (AI) has outpaced its formal adoption in health care organizations, contributing to the emergence of Shadow AI, defined here as the use of unauthorized AI tools by medical professionals. Under the European Union Medical Device Regulation, AI tools used for clinical purposes must undergo conformity assessment before use; general-purpose tools such as ChatGPT have not done so, rendering their clinical application unauthorized at the regulatory level. While Shadow AI offers potential efficiency gains and higher performance, it poses significant risks to data privacy, clinical safety, and regulatory compliance. Despite its growing prevalence, empirical research on the purposes for which physicians use Shadow AI remains scarce. OBJECTIVE: This study explores the purposes for which physicians describe using Shadow AI in their work. METHODS: We conducted a cross-sectional survey of physicians employed in Swedish health care organizations (N=357; response rate~64%). Data were collected between December 2023 and January 2024 via a verified online panel. We conducted a qualitative content analysis of free-text responses on the use of unauthorized AI tools. We applied theoretical lenses from the sociology of professions and paradox theory to interpret the empirical findings. RESULTS: Physicians use Shadow AI for several purposes, which we grouped into 4 categories: clinical work and decision-making, administrative work, research and professional development, and technological interest and curiosity. More specifically, Shadow AI is used as a colleague and second opinion for clinical decision support (eg, differential diagnoses and rare cases), administrative tasks such as patient communication and documentation, and research aimed at staying up to date and exploring developments in generative AI. Physicians described using these tools compensated for perceived gaps in institutional systems, reducing workload, and accessing knowledge considered difficult to obtain through conventional channels. The findings reveal a tension between physicians' drive to improve their practice and the regulatory and organizational constraints that render such use unauthorized. CONCLUSIONS: Shadow AI used by physicians presents both opportunities and risks for health care professionals and organizations. Shadow AI indicates gaps where formal hospital systems may fail to meet health care professionals' needs and signals a way for physicians to strengthen their experience-based knowledge. It represents a renegotiation of professional boundaries, as physicians bypass institutional constraints to maintain professional efficacy. The findings highlight a paradox in which the same tools that pose regulatory and safety risks also address real gaps in clinical and administrative support, suggesting that governance approaches must account for this tension rather than relying on prohibition alone.
30 July 2026
Read appraisal →Journal of medical Internet research
Evaluation Frameworks for Clinical AI Incorporating Validation Strategies, Real-World Applicability, and Ethical Principles: Scoping Review
BACKGROUND: AI shows substantial potential in health care; however, the absence of standardized evaluation frameworks limits its safe and effective clinical implementation because of inconsistent validation requirements and fragmented ethical principles. Existing guidelines vary in structure, methodological rigor, and ethical integration, creating uncertainty. OBJECTIVE: This study aimed to systematically map, characterize, and critically analyze existing evaluation frameworks for clinical AI, focusing on three core dimensions: methodological rigor, validation strategies (internal validation, including reporting of technical and clinical performance; external validation, including real-world applicability), and alignment with the United Nations Educational, Scientific and Cultural Organization (UNESCO) AI ethical considerations. METHODS: A scoping review was conducted following PRISMA-ScR (Preferred Reporting Items for Systematic Reviews and Meta-Analyses extension for Scoping Reviews) guidelines. Six databases (PubMed, Embase, BVS, EBSCOhost, ProQuest, and Sage) and the Enhancing the Quality and Transparency of Health Research Network were searched without language or date restrictions up to February 2026. Eligible documents included peer-reviewed papers, gray literature, and organizational guidelines describing evaluation or reporting frameworks for clinical AI. Editorials, commentaries, and conference abstracts lacking a clearly defined evaluative framework or clinical applicability were excluded. Two reviewers independently screened records and extracted data. Data were extracted across three domains: (1) general characteristics, (2) methodological rigor and validation parameters, and (3) ethical integration and were synthesized using a dot plot-based gap map. Ethical adherence was assessed using a 10-domain UNESCO-based scoring matrix. No formal risk-of-bias assessment was conducted, consistent with scoping review methodology. RESULTS: From 3363 records, 46 frameworks met the inclusion criteria. Mapping revealed a rapidly expanding but fragmented landscape. Most frameworks targeted investigational use (88%), with limited focus on clinical applicability. Frameworks varied in structure, methodology, and scope, with a predominance of reporting guidelines and few validated tools. Most (63%) were developed through multi-institutional collaborations, and 32.6% incorporated transdisciplinary participation. Only 31.8% reported technical metrics (commonly area under the curve, sensitivity, and specificity), and 15.9% provided clinical indicators (eg, predictive values or calibration). Only 11.4% achieved methodological rigor, incorporating validation aligned with intended use, while most relied on partial validation strategies, highlighting a gap between model development and clinical evaluation. Ethical integration was heterogeneous: only 5 frameworks achieved high compliance (≥80%), whereas 4 scored <10%. The most frequently addressed UNESCO principles were awareness and education (71.1%) and transparency and explainability (70%), while human oversight (24.4%) and adaptive governance (33.3%) were least represented. Findings indicate a misalignment between framework design, validation requirements, and clinical implementation. CONCLUSIONS: Evaluation frameworks for clinical AI remain heterogeneous and oriented toward investigational contexts. Critical gaps persist in methodological rigor, validation aligned with intended use, and fragmented ethical coverage. These findings highlight the need for standardized, robust, and ethically grounded frameworks to enable safe, reliable, and scalable integration of AI into clinical practice.
24 July 2026
Read appraisal →