Evidence-Based Medicine

Research Appraisals

Evidence-based critical appraisals of the latest medical research, systematically evaluated using Oxford CEBM methodology.

Showing 41 appraisals

Systematic ReviewEvidence: Weak
25CEBM

Cognitive, affective & behavioral neuroscience

Emotion detection unveiled: A cognitive-computational synthesis of physiological models, machine learning, and datasets.

This comprehensive survey synthesizes state-of-the-art advancements in emotion recognition based on physiological signals, specifically focusing on the paradigm shift occurring between 2021 and 2025. Crucially, we move beyond a technical review by establishing a novel Cognitive-Computational Synthesis Framework (CCSF). This framework explicitly maps multimodal physiological manifestations (e.g., electroencephalogram (EEG), electrocardiogram (ECG), and galvanic skin response (GSR)) to underlying cognitive processes, such as attentional allocation, arousal regulation, and perceptual bias, providing a theoretical foundation for explainable AI (XAI) in affective computing. We meticulously examine the transition from traditional machine learning to advanced deep learning architectures, highlighting how recent innovations in Transformers, self-supervised learning, and diffusion models have shattered previous performance plateaus. While earlier dimensional models were often limited to 70-75% accuracy, this survey details how modern architectures now achieve benchmarks exceeding 95% on seminal datasets like SEED and DREAMER. Furthermore, the survey provides a rigorous analysis of 40 key studies (identified via PRISMA protocols), evaluating them based on their validation strategies, cross-subject generalizability, and adversarial robustness. By bridging the gap between raw physiological data and cognitive theory, this work offers a strategic roadmap for the next generation of robust, interpretable, and real-time emotion recognition systems.

3 Aug 2026

Read appraisal →
Randomised Controlled TrialEvidence: Weak
35CEBM

Science advances

Human-inspired time-series health evaluation with an adaptive multimodal electronic skin

Electronic skin powered with artificial intelligence could enable next-generation robotic and medical devices, yet integrating multimodal sensors and analyzing heterogeneous, multifrequency time series remain challenging. Most wearable machine learning architectures are time-invariant and trained for a specific task, limiting transfer across modalities and users. We present a multimodal electronic skin that captures diverse physiological signs with an adaptive learning framework that rapidly generalizes to unseen tasks with minimal labeled data. Our streamlined end-to-end framework uses a spectral variational autoencoder to denoise and compress multifrequency biosignals into a shared, unified second-wise latent space that preserves the spectral-temporal structure, followed by a transformer to capture temporal dependencies to support diverse downstream tasks with data-efficient learning. We demonstrate robust adaptation with 94.7% accuracy in activity recognition and 90.2% precision in fatigue assessment across various users and daily activities regardless of device and user variations, highlighting a scalable route to generalized physiological time-series analytics and human performance assessments.

3 Aug 2026

Read appraisal →
observationalEvidence: Moderate
75CEBM

Journal of the American Medical Informatics Association : JAMIA

Sociodemographic bias in large language model clinical trial screening

OBJECTIVE: To assess whether large language model (LLM)-based clinical trial screening judgments vary by patient sociodemographic characteristics. MATERIALS AND METHODS: We conducted a cross-sectional evaluation of Phase II-III US adult randomized controlled trial (RCT) protocols (2023-2024). Physician-validated clinical vignettes were evaluated in a control version and 33 sociodemographic identity variants differing only by labels. Nine LLMs assessed eligibility and related domains. Mixed-effects models estimated adjusted differences vs control. RESULTS: Across 58 protocols and 5.3 million evaluations, eligibility judgments were largely stable across identities. Race and ethnicity showed minimal effects after accounting for socioeconomic status. Homelessness produced the largest negative eligibility shift and pronounced effects in adherence, resources, and trust. DISCUSSION AND CONCLUSION: LLMs applied explicit eligibility criteria consistently, but disparities emerged in domains requiring inference about behavior or resources, underscoring the need for careful deployment to promote fair trial access.

2 Aug 2026

Read appraisal →
otherEvidence: Weak
30CEBM

Rheumatic diseases clinics of North America

Artificial Intelligence Regulation in the United States: Current Landscape and Implications for Rheumatology

Artificial intelligence (AI) is increasingly embedded in clinical tools used in rheumatology, including imaging interpretation, longitudinal disease monitoring, and electronic health record-based decision support. AI has moved from the periphery of biomedical research to an operational component of clinical care, increasingly embedded in electronic health records, imaging platforms, and decision support systems. In rheumatology, where care is longitudinal, AI systems offer substantial promise-but also poses distinct risks. Together, these commitments-advancing humanity, ensuring equity, engaging impacted individuals, improving workforce well-being, monitoring performance, innovating and learning, and promoting sustainability-operationalize trustworthy AI systems across a patient's care trajectory.

2 Aug 2026

Read appraisal →
Systematic ReviewEvidence: Weak
20CEBM

Journal of food science

Artificial Intelligence in Food-Nutrition-Health Research: From Multimodal Data Integration to Precision Intervention

Artificial intelligence (AI) is transforming food-nutrition-health research by enabling pattern recognition in complex, high-dimensional datasets that traditional hypothesis-driven approaches cannot address. This review systematically synthesizes research progress of AI across the food-nutrition-health continuum from 2020 to 2025. By examining 181 systematic reviews through PRISMA-guided selection, we provide a comprehensive overview and prospects across four dimensions: technical foundation, application scenarios, existing challenges, and future prospects. We propose a tripartite framework comprising (1) a data layer enabling multisource fusion of food composition, health monitoring, and individual characteristic data; (2) a technological layer of nondestructive testing (spectroscopy, nuclear magnetic resonance [NMR], imaging); and (3) an algorithmic layer progressing from machine learning to deep learning architecture. Key applications include food component analysis and safety detection; nutrition-disease association modeling; pathogen identification; and personalized dietary intervention systems. Despite rapid progress, critical challenges persist, insufficient model generalization across populations, algorithmic opacity limiting clinical trust, data privacy vulnerabilities, and lack of standardized multi-omics integration protocols. Future directions emphasize multimodal fusion models, explainable artificial intelligence (XAI), federated learning for privacy-preserving collaboration, gene-guided precision nutrition, and development of intelligent wearable devices and functional food. This review provides a roadmap for transitioning from population-averaged guidelines to dynamic, individualized health optimization through AI-enabled food system.

1 Aug 2026

Read appraisal →
observationalEvidence: Weak
35CEBM

PloS one

Development of a scale for measuring the perception of artificial intelligence among mental health consumers

BACKGROUND: Artificial Intelligence (AI) has emerged as a transformative force revolutionizing various sectors, including healthcare, particularly the mental health field. However, the acceptance and integration of AI technologies in different healthcare systems can be influenced by various factors, including cultural, social, and individual aspects. Nevertheless, there is a need for a valid and reliable tool for assessing AI's perception among healthcare consumers. AIM: To develop and validate a tool for the perception of AI among healthcare consumers and apply the tool to assess AI's perception among mental health consumers in the Jordanian healthcare system. METHOD: A cross-sectional descriptive correlational design was utilized in the study. Data was collected from a convenience sample of 431 mental health consumers visiting mental health clinics of the International Medical Corps and university hospitals in Jordan. Structured interviews were conducted using an AI Perception (AIP) questionnaire developed by the authors. The questionnaire's content validity was assessed by an expert panel. Using Principal Component Analysis (PCA), the construct validity of the tool was evaluated, and its internal consistency was examined using Cronbach's alpha. Descriptive statistics were used to assess the levels of AI perception among participants. RESULTS: The final AIP tool consisted of 20 items across 4 domains and has demonstrated strong internal consistency across its four domains: AI acceptance and readiness (α = 0.92), AI perceived importance (α = 0.92), AI perceived risk (α = 0.9), and AI perceived challenges (α = 0.85). The construct validity of the four-domain structure of the tool was supported by PCA. Additionally, the mean scores for each domain indicated the average level of agreement with AI perception items among participants. Specifically, the mean score for AI acceptance and readiness was [Formula: see text]). AI perceived importance was (2.18 [Formula: see text]0.83), AI perceived risk was (2.58[Formula: see text]0.92), and AI perceived challenge was (2.78 [Formula: see text] 0.87). CONCLUSION: The findings of this study resulted in developing a valid and reliable 20-item tool to assess AI's perception among mental health consumers. The tool can be used to assess the predictors of AI's readiness among mental health consumers. Therefore, aiding policymakers and other stakeholders in understanding the AI adoption barriers from the perspective of end-users. In addition, this study developed the AIP tool that can be validated and used among other populations in future research.

31 July 2026

Read appraisal →
Systematic ReviewEvidence: Moderate
50CEBM

ACS applied materials & interfaces

Advances and Challenges in Wearable Sensors for Health Monitoring

Analytical tools may revolutionize healthcare by enabling accessible, rapid, and decentralized testing. Wearable (bio)sensors, in particular, provide frequent or continuous patient monitoring through non- to minimally invasive measurements. This approach yields unprecedented amounts of health-related information, leading to more informed clinical decision-making and closer patient follow-up. In this mega-review article, we bring together leading researchers in the field to discuss the state of the art in wearable devices for health monitoring. We begin by providing a broad overview of the field through citation network analysis. We then review the application of chemical (bio)sensors in biofluids (e.g., sweat, saliva, tears, interstitial fluid, and cerebrospinal fluid), highlighting the challenges and advantages associated with each. Subsequently, we discuss the construction of wearable devices and their main formats (e.g., smart contact lenses, textiles, mouthguards, watches/wristbands, and implantable systems). Physical sensors are addressed in a dedicated section focusing on the assessment of heart rate, blood pressure, and body temperature. The role of soft electronics in wearable devices is also examined, as these technologies are essential for enhancing user comfort and sensor reliability, which demands advances in materials science. Furthermore, we present strategies for signal acquisition and transmission, as well as approaches for on-body energy harvesting and device self-powering. The use of artificial intelligence and machine learning is then discussed as a means of enhancing analytical performance and managing the large volumes of data generated by wearable devices. Finally, business, regulatory, and ethical considerations are examined. We expect that this review will provide an overview of sensing and biosensing technologies for health-related applications, identify promising research directions, and inspire future developments.

30 July 2026

Read appraisal →
diagnosticEvidence: Weak
35CEBM

Scientific reports

Recognition of everyday activities using experiment data from wearable sensors: a deep learning-based framework

Tracking everyday activities is vital for detecting changes in older adults' health, allowing timely support to promote well-being. Wearable sensors and deep learning provide continuous monitoring, making them a supportive tool in detecting such changes. However, a more refined method is needed to recognise precise activities with a minimal set of sensors. This study aimed to develop a method to recognise everyday activities among older adults by utilising wearable sensors and a deep learning model. This is a small-scale home lab experiment to develop a method to recognise 14 everyday activities. We compared five models that recognised everyday activities with different sensor signal counts and accuracy. Our results showed that sensor placement is important. Based on the results, we proposed a two-sensor method (pelvis and right hand) to collect and correctly recognise everyday activities among older adults. This model, which utilises two sensors, classified 12 activities with an accuracy of 89.3%. Another model recognised all 14 activities with a lower accuracy of 88.2% using five sensors. We also explored a one-sensor approach, which showed low recognition performance and struggled to distinguish activity variability. The two-sensor-based system will allow for large-scale data collection on everyday activities of older adults.

26 July 2026

Read appraisal →
otherEvidence: Insufficient
20CEBM

Journal of medical Internet research

The Next Generation of Wearables Won't Need the Cloud

Wearable health and fitness devices have typically relied on cloud computing to deliver insights. In this News and Perspectives article, JMIR Correspondent Michelle Falci reports on the advances facilitating on-device data processing and the potential of these next generation wearables.

25 July 2026

Read appraisal →
Randomised Controlled TrialEvidence: Weak
30CEBM

Scientific reports

Integrating skeleton based representations for robust yoga pose classification using deep learning models

Yoga is a popular form of exercise worldwide due to its spiritual and physical health benefits, but incorrect postures can lead to injuries. Automated yoga pose classification has therefore gained importance to reduce reliance on expert practitioners. While human pose keypoint extraction models have shown high potential in action recognition, systematic benchmarking for yoga pose recognition remains limited, as prior works often focus solely on raw images or a single pose extraction model. In this study, we introduce a curated dataset, "Yoga-16", which addresses limitations of existing datasets, and systematically evaluate three deep learning architectures-VGG16, ResNet50, and Xception-using three input modalities: direct images, MediaPipe Pose skeleton images, and YOLOv8 Pose skeleton images. Our experiments demonstrate that skeleton-based representations outperform raw image inputs, with the highest accuracy of 96.09% achieved by VGG16 with MediaPipe Pose skeleton input. Additionally, we provide interpretability analysis using Grad-CAM, offering insights into model decision-making for yoga pose classification with cross validation analysis.

17 July 2026

Read appraisal →
otherEvidence: Weak
35CEBM

NEJM catalyst innovations in care delivery

Thinking about the Impact of Artificial Intelligence on U.S. Health Care Costs and Spending Growth

This article critically examines the potential impacts of artificial intelligence (AI) on prices, total spending, and the rate of spending growth in the areas of prescription drug innovation and expanded patient access to care - including developments in remote patient monitoring, chronic care management, direct-to-consumer health care, clinical decision support, and nonclinical administrative labor. Based on observed industry practices, economics, and public policy, the article presents a thought exercise on the most likely effects for patients, providers, payers, and the broader health care system. The authors argue that under the still-dominant fee-for-service payment model, as well as the highly consolidated hospital and insurance markets, AI is more likely to increase total costs and spending growth in the short to medium term rather than slow them, even as it delivers substantial access and clinical quality improvements for patients. The cost-bending potential of AI varies distinctly by fee-for-service versus value-based payment. Regulators need to introduce policy and reimbursement levers for AI to slow cost growth. The authors conclude that, without significant changes in the health care system's financial incentives and market structures, AI will not slow cost growth.

16 July 2026

Read appraisal →
Systematic ReviewEvidence: Weak
45CEBM

Journal of medical Internet research

User Acceptability and Adoption of AI-Generated Lifestyle Intervention Recommendations: Scoping Review and Theoretical Integration

BACKGROUND: Artificial intelligence (AI)-generated lifestyle recommendations are increasingly used to support health behavior change. However, AI advice does not necessarily mean that users will accept or adopt those recommendations. Although prior reviews have examined AI-enabled lifestyle interventions and health behavior technologies, fewer have focused on whether users accept and adopt AI-generated recommendations. OBJECTIVE: This scoping review aimed to map user acceptability and adoption of AI-generated lifestyle recommendations in user-facing systems used by end users or caregivers. Objectives were to characterize systems and evaluation contexts, clarify how recommendation-level outcomes were conceptualized and measured, synthesize shaping factors, and develop an evidence-informed framework to guide future research, evaluation, and design. METHODS: Following JBI (Joanna Briggs Institute) guidance and the PRISMA-ScR (Preferred Reporting Items for Systematic Reviews and Meta-Analyses extension for Scoping Reviews), we searched Ovid MEDLINE, Ovid Embase, APA PsycInfo via ProQuest, Web of Science Core Collection, Scopus, ACM Digital Library, and IEEE Xplore from database inception to May 5, 2026. The initial search was conducted on November 21, 2025, and an updated search was conducted on May 5, 2026. Eligible studies reported empirical end-user or caregiver data, evaluated AI-generated lifestyle recommendation content delivered without manual review or editing, and reported an acceptability or adoption outcome linked to recommendations. English empirical papers and conference papers were included. Data were charted on study, system, outcome, measurement, factor, and theoretical characteristics. Quality was assessed with the Mixed Methods Appraisal Tool. Findings were synthesized descriptively and through evidence mapping. RESULTS: Searches yielded 12,997 records; 8570 unique records were screened, and 21 studies were included. Most were published in 2025 or 2026 (17/21, 81%). Large language model-centered systems were the most common format (12/21, 57.1%). Outcomes were concentrated in acceptability-related perceptions, such as satisfaction or enjoyment, perceived quality or fit, and persuasiveness, whereas adoption-related outcomes were assessed less often and mainly reflected intention, in-study uptake, or short-term enactment. Factors clustered across system capabilities, content properties, individual states and capacities, and contextual constraints. Findings informed an integrative perception-intention-enactment framework positioning acceptability and adoption as a system-content-user-context process. CONCLUSIONS: This review extends prior AI and digital health reviews by shifting attention toward how users perceive, intend to follow, and enact AI-generated lifestyle recommendations. Acceptability and adoption appear to depend on systems eliciting and adapting to context, content being actionable and credible, users having the capacity to interpret, trust, and engage with recommendations while retaining control, and resources, routines, and social contexts allowing enactment. The framework can guide theory-driven evaluation, outcome selection, and system design by identifying where recommendation processes may succeed or fail, but should be interpreted as preliminary and evidence-informed rather than causal. By integrating implementation, behavioral, and human-AI perspectives, this review provides a foundation for moving AI-generated lifestyle recommendations from technically plausible outputs toward user-centered, context-sensitive, and behaviorally actionable support.

16 July 2026

Read appraisal →
observationalEvidence: Weak
25CEBM

International journal of medical informatics

Examining user-AI interaction patterns in health-information queries

In this study, we examine how individuals utilize generative artificial intelligence (GAI) when seeking health-related information. Using a dataset of user-GAI chat logs available on Hugging Face, we analyzed real-world interactions in which users posed health-related questions to a generative model. We applied a combination of data and text-analytic methods to categorize these interactions, including supervised machine learning techniques such as Support Vector Machines (SVMs). SVMs were selected for their efficiency and strong performance in high-dimensional text classification tasks, and used to identify recurrent themes in user queries and interactions. We found that users frequently consult AI chatbots for symptom exploration, medical education, mental health support, and general health advice. The findings suggest that GAI tools may not only function as informational resources, but also as preliminary support tools that can shape users' health knowledge and encourage them to seek consultation with medical professionals.

16 July 2026

Read appraisal →
otherEvidence: Weak
50CEBM

JMIR mental health

AI Agents Are Coming: 5-Stage Taxonomy of Language-Based AI Systems for Psychiatry, Psychotherapy, and Counseling

The rapid evolution of large language models has accelerated the development of agentic artificial intelligence (AI) systems capable of pursuing autonomous goals, creating an urgent need for structural frameworks in psychiatry and psychotherapy. While existing classifications often draw parallels to autonomous driving, this paper argues that the mental health domain requires a distinct, domain-specific theoretical foundation, as the 2 domains differ fundamentally in their semantic, ideographic, and epistemological demands. Furthermore, they differ in their end goals, for which we introduce terms such as agentic guidance capability. To guide clinicians and researchers through these developments, we propose a 5-stage taxonomy for language-based AI systems that differentiates technical functionality from clinical effectiveness. The taxonomy progresses from level 1 (knowledge level), in which systems perform static benchmark tasks, to level 2 (elementary level), characterized by dynamic engagement in specific therapeutic microskills. At level 3 (integration level), systems achieve consistency across and within modules, as well as basic case-level conceptualization suitable for blended therapy under human oversight. Level 4 (saturation level) describes therapist-in-the-loop systems capable of autonomous functioning with minimal supervision, whereas level 5 (mastery level) represents AI systems that are technically capable of performing autonomous therapy. By distinguishing technical functionality from clinical effectiveness, we conclude that level 4 or level 5 performance does not automatically translate into full treatment effectiveness, even if high treatment fidelity can be achieved. We conclude by emphasizing the need to shift benchmarking from static knowledge tests to dynamic evaluations of therapeutic capabilities in order to safely navigate the transition toward autonomous care.

15 July 2026

Read appraisal →
otherEvidence: Insufficient
30CEBM

JMIR cancer

Internet of Things-Enhanced Mathematical Oncology: Conceptual Framework for Adaptive Cancer Care Modeling

The Internet of Things (IoT) is transforming various industries, including health care. IoT-based systems are increasingly prevalent in consumer health applications, while intelligent or smart devices equipped with sophisticated sensors are gaining recognition for their potential to improve clinical care practice and decision-making. Cancer care is a particularly promising area for IoT applications, enabling real-time and personalized interventions. However, empirical research on the effects of IoT in this field is limited due to the complexities inherent in cancer as a dynamic disease and the paucity of IoT-generated data available for research. This presents an opportunity to apply mathematical modeling to understand the effects of IoT under various scenarios. These analytical and "in silico" mathematical approaches are instrumental with limited data. Such models support the analysis of treatment uncertainty and patient response while balancing patient preferences, clinical outcomes, and system-level constraints. Grounded in mathematical oncology and health informatics, this paper proposes a conceptual framework that integrates real-time IoT data as dynamic inputs into adaptive mathematical models to simulate cancer dynamics. By exploring applications across multiple levels of analysis, the study demonstrates how IoT-enhanced mathematical models could inform implementation and optimize oncology services, addressing a critical gap in current research.

15 July 2026

Read appraisal →
Systematic ReviewEvidence: Weak
30CEBM

Vaccine

Mobile health strategies to improve HPV vaccination uptake among parents of adolescent girls: a scoping review

BACKGROUND: Human papillomavirus (HPV) vaccination is a highly effective primary prevention strategy against cervical cancer. However, despite strong evidence of vaccine efficacy and safety, uptake among adolescent girls remains suboptimal globally, particularly in low- and middle-income countries. Mobile health (mHealth) technologies have emerged as promising tools to address parental concerns, reduce vaccine hesitancy, and strengthen parental engagement in vaccination decision-making. OBJECTIVE: This scoping review aimed to synthesize evidence on the types of mHealth strategies implemented to improve HPV vaccine uptake among parents of adolescent girls and to examine the outcomes of these interventions. METHODS: A systematic search of major databases identified studies evaluating mHealth interventions targeting parental decision-making regarding HPV vaccination. Eligible interventions included text messaging, mobile applications, chatbots, web-based educational platforms, and other digital health tools. Data were charted and synthesized narratively. KEY FINDINGS: mHealth interventions were associated with significant improvements in HPV vaccine uptake across multiple settings. SMS-based interventions consistently increased vaccination rates, while chatbot, web-based, and mobile app interventions improved parental knowledge, vaccine literacy, intention to vaccinate, and, in several studies, actual vaccine uptake. However, intervention effectiveness varied by socioeconomic context, digital access, and cultural factors. CONCLUSION: mHealth technologies represent cost-effective and scalable strategies to improve HPV vaccine uptake by enhancing parental knowledge, reducing hesitancy, and supporting informed decision-making. Future implementation should prioritize culturally tailored, equity-focused approaches and integrate digital interventions with community-based engagement to maximize impact.

12 July 2026

Read appraisal →
qualitativeEvidence: Moderate
80CEBM

JMIR aging

Smart Speaker-Based Applications to Support Social Connectedness in Older Adult Residents in Affordable Housing: User-Centered Design Study

BACKGROUND: Older adults in affordable housing face heightened risks of social isolation and loneliness due to limited social networks, transportation barriers, chronic conditions, and inadequate technology access. Smart speakers offer potential for enhancing social connectedness in this underserved population, yet technology interventions are rarely designed with meaningful input from older adults themselves. User-centered design (UCD) approaches can address this gap by engaging end users throughout the development process to ensure technology solutions align with their needs and living contexts. OBJECTIVE: This study aimed to engage older adults in affordable housing in an iterative UCD process to develop prototype scenarios for smart speaker-based applications that promote social connectedness while addressing safety, community-building, and wellness needs. METHODS: We conducted a 3-stage UCD study with 29 older adults (mean age 70, SD 6.8 years; 23/29, 79% African American; 20/29, 69% high school education or less) living alone in affordable housing between April 2021 and April 2022. Stage 1 included 5 focus groups (n=25) combining needs assessment discussions with rapid brainstorming activities. Stage 2 involved research team synthesis of focus group transcripts and brainstorming data to create a Design Strategies Map, development of initial prototype scenarios, 5 evaluation focus groups (n=18) to gather feedback, and iterative scenario refinement. Stage 3 comprised 4 validation focus groups (n=17) to assess refined scenarios and identify implementation recommendations. Participants included both smart speaker users (n=13) and nonusers (n=16). Data were analyzed using thematic analysis for needs assessment, content analysis for brainstorming ideas and feedback, and matrix analysis for systematic comparison across scenarios. RESULTS: Participants generated 153 ideas for smart speaker use, with Health and Safety and Daily Assistance being the most frequent categories. Analysis revealed that social connection needs were inseparable from safety concerns related to living alone. Through iterative co-design, we developed 7 prototype scenarios across 4 functional categories: Checking-In (peer and management safety verification with privacy controls), Social Companion (conversational artificial intelligence-based companionship and emotional support), Community Involvement (virtual bulletin boards and activity coordination), and Wellness Check (system-initiated monitoring of activity and behavioral patterns as health indicators with user-controlled interventions). Participants emphasized requirements for personalization, opt-in/opt-out controls, "Do-Not-Disturb" functionality, and safeguards preventing replacement of human connection. CONCLUSIONS: Older adults in affordable housing engaged in technology design and provided valuable insights that challenge assumptions about their needs and preferences. The prototype scenarios addressed the dual imperatives of social connection and safety while living alone, offering a foundation for developing technology-based applications tailored to underserved populations. Implementation should prioritize user control, privacy protection, and human-in-the-loop design ensuring that technology facilitates rather than replaces human connection and community programming, alongside consideration of user characteristics to build trust and ensure effective, sustained use of the intended technology platform.

9 July 2026

Read appraisal →
otherEvidence: Weak
65CEBM

BMJ open

Implementation determinants of a planned machine learning-enabled surgical scheduling system in a high-volume orthopaedic centre in Canada: qualitative findings.

OBJECTIVES: Elective non-emergent surgical wait times have increased across countries such as Canada, straining operating room (OR) resources and affecting patient outcomes and healthcare spending. Manual scheduling systems in Ontario orthopaedic centres create wide variations in wait times, with recent declines in meeting benchmark targets despite increased procedure volumes. Challenges stem from fragmented referral processes, outdated scheduling methods and resource constraints. Artificial intelligence and machine learning (ML) offer potential solutions for optimising scheduling; however, their implementation remains inconsistent. This study aims to identify determinants affecting the rollout of a new ML-driven automated scheduling system at a high-volume elective orthopaedic surgery centre. DESIGN: A qualitative description approach supported by implementation science frameworks. SETTING: A high-volume elective orthopaedic surgery unit at a Canadian tertiary care centre. PARTICIPANTS: 17 individuals from clinical, administrative and leadership roles who were directly involved in surgical scheduling. INTERVENTIONS: A new ML-driven automated surgical scheduling system. OUTCOMES: Perceptions of the proposed new surgical scheduling system (barriers and enablers of implementation, recommendations for improvement). RESULTS: Three main themes were identified, capturing challenges and enablers in the existing scheduling system: system functionality, process-related factors and resource constraints.Participants described substantial inefficiencies in the existing manual scheduling system, including outdated software, fragmented information systems, inconsistent communication and resource constraints. Across interest-holder groups, there was broad but variable perceived support for a planned ML-enabled scheduling system, particularly for improving duration prediction, access to scheduling data and reporting, alongside concerns about system complexity, workflow fit, training and resource implications. Interest-holders emphasised the importance of user-friendly design, interoperability, responsive training, phased implementation and ongoing feedback. CONCLUSIONS: This pre-implementation qualitative study identified significant process and resource limitations in manual orthopaedic surgical scheduling, but interest-holder support for a well-designed ML-driven system is strong. While participants anticipated potential benefits for scheduling accuracy, throughput and resource allocation, these perceived advantages will require meaningful user engagement, robust training, phased rollout and evaluation in subsequent implementation and outcome studies.

8 July 2026

Read appraisal →
otherEvidence: Weak
60CEBM

BMJ global health

Enhancing malaria-in-pregnancy monitoring: stakeholder experiences and data integration into the BornFyne-PNMS digital platform in Cameroon

Malaria in pregnancy (MiP) is a significant public health concern in high-burden malaria countries and a key predictor of high-risk pregnancies in Cameroon. One of the challenges of the National Malaria Control Programme (NMCP) in Cameroon is to generate reliable, timely, accurate and actionable data to inform decision-making, effective monitoring and planning. Discussions with the NMCP at the Ministry of Public Health (MoPH) identified poor data quality, lack of clear guidelines, manual data processing at the district level and delays in feedback on data as key bottlenecks in effective implementation of malaria control in Cameroon. This paper documents the data gaps for MiP identified by the digital health platform, BornFyne, project team in collaboration with the MoPH. We outlined lessons learnt in identifying and defining MiP-related indicators through stakeholder engagement across various levels, and identifying priority MiP elements for health facility registers for integration into the BornFyne digital platform. These lessons and insights will provide valuable guidance for other countries looking to incorporate malaria-related content into their antenatal care packages and/or to integrate similar content into their digital health platforms, and suggestions for strengthening the WHO digital adaptation kit MiP content.

8 July 2026

Read appraisal →
observationalEvidence: Weak
55CEBM

BMJ open

The Cyber Paranoia and Fear Scale-Updated (CPFS-U): development and implications for digital health engagement

OBJECTIVES: To update and revalidate the Cyber Paranoia and Fear Scale to reflect current technological contexts and examine its relevance to digital health readiness and engagement. METHODS: Using an online community sample (n=433), exploratory factor analysis was conducted to examine the factor structure of the revised item pool. Items were refined through consultation with Patient and Public Involvement and Engagement groups to ensure contemporary relevance and clarity. RESULTS: Analysis supported a four-factor structure representing artificial intelligence (AI) and digital dependence, technological risk awareness, perceived data vulnerability and surveillance-related mistrust. The updated Cyber Paranoia and Fear Scale-Updated (CPFS-U) demonstrated good internal consistency and supported construct validity. Cyber-paranoia and fear were conceptually and empirically distinct from general paranoia and anxiety, highlighting the specific cognitive and emotional responses elicited by digital technologies. CONCLUSIONS: The CPFS-U offers a psychometrically robust, modernised measure for understanding individuals' responses to digital and AI-based technologies. Its application in digital health research and practice can inform risk communication, user engagement strategies and the design of trustworthy digital interventions. By identifying individuals who may disengage due to online mistrust, the CPFS-U has the potential to inform more inclusive and psychologically informed digital health systems.

8 July 2026

Read appraisal →
otherEvidence: Weak
25CEBM

BMJ health & care informatics

Predicting health and disease: a conceptual framework for AI in preventive and precision medicine

The conventional medical approach of treating symptoms as they appear with restricted screening often limits intervention to slowing disease progression rather than fully reversing it. A new approach leveraging artificial intelligence (AI) and computational technologies across expanding multimodal biomedical datasets holds the promise to enable predicting actionable future changes in health before symptom onset. This article presents a conceptual framework for a preventive paradigm in precision medicine and healthcare, integrating recent advancements in AI and biomedical datasets. Key remaining challenges facing computational systems, real-world clinical validation and implementation, and preventive interventions together with recommendations and prioritised future directions are highlighted. Such an approach could pave the way for more proactive and preventive medical interventions to effectively address the growing burden of chronic disease.

4 July 2026

Read appraisal →
Systematic ReviewEvidence: Weak
10CEBM

The Nigerian postgraduate medical journal

The Role of Digital Therapeutics and Artificial Intelligence in Chronic Disease Management: A Narrative Review

We conducted research showing that chronic disease management continued to challenge healthcare systems, payers, and patients. At the same time, digital therapeutics emerged as a promising and potentially transformative approach. They supported better management of long-term conditions, improved patient outcomes, and helped streamline healthcare delivery. We carried out a comprehensive literature search to identify both original reports and reviewed publications. The search covered multiple databases, including Google Scholar and PubMed, and we also gathered relevant information from credible online sources such as the World Health Organization and India's National Crime Records Bureau. Using these findings, the narrative explained how digital therapeutics reshaped chronic disease management by highlighting key benefits, the technologies that enabled these solutions, and the expected impacts on both clinical and economic outcomes. We also discussed the obstacles the sector encountered as it developed, and we considered the future prospects of digital therapeutics in chronic care. In conclusion, digital therapeutics offered personalised care, improved patient engagement and adherence, provided real-time monitoring and feedback, enhanced accessibility and convenience, and were cost-effective in managing chronic disease; however, challenges such as cybersecurity concerns, reliability of data, the digital divide, and a lack of extensive clinical validation needed to be addressed for widespread adoption. While the evidence to date suggested clear clinical and economic promise, realizing that promise required coordinated action stronger clinical trials to build robust evidence, clear regulatory pathways to ensure safety and efficacy, investment in secure interoperable infrastructure, and targeted efforts to close the digital divide so vulnerable populations were not left behind; policymakers, clinicians, payers, and technology developers had to collaborate to translate innovation into equitable, scalable care improvements.

3 July 2026

Read appraisal →
otherEvidence: Weak
25CEBM

Journal of medical Internet research

How Does That Large Language Model Make You Feel?

People are increasingly turning to commercially available large language models (LLMs) for emotional support. In this News and Perspectives article, JMIR Correspondent Simon Spichak reports on the role of LLMs in mental health, speaking with experts about safety concerns, research gaps, and next steps.

2 July 2026

Read appraisal →
Systematic ReviewEvidence: Weak
45CEBM

International journal of medical informatics

The role of artificial intelligence in virtual emergency care: a systematic review

BACKGROUND: The integration of artificial intelligence (AI) into virtual emergency care represents a potentially transformative approach to healthcare delivery, yet the evidence base remains poorly characterized. This systematic review comprehensively evaluates the current state of AI applications in virtual emergency care settings. METHODS: We systematically searched eight databases (Embase, PsycINFO, MEDLINE, PubMed, Scopus, Web of Science, CINAHL, Cochrane Library) from inception through March 2025. Of 7,098 records identified and 4,935 screened after deduplication using Covidence, 8 studies met inclusion criteria following exclusion of one study lacking AI components. Studies were assessed using PROBAST + AI for risk of bias and quality assessment, TRIPOD + AI for reporting quality, and GRADE for certainty of evidence. RESULTS: The eight included studies (total participants: approximately 0.5 million) evaluated diverse AI applications including decision trees, machine learning ensembles, and graph neural networks across multiple virtual emergency contexts. Performance varied widely (accuracy 77.5-100%, sensitivity 63-100%, specificity 60% in single study reporting). All clinical studies demonstrated serious risk of bias. TRIPOD + AI compliance averaged only 36.9% (range 30.9-48.1%). GRADE assessment revealed very low to low certainty evidence across all outcomes, with no studies measuring actual clinical outcomes. CONCLUSIONS: Current evidence is insufficient to support widespread clinical implementation of AI in virtual emergency care. While preliminary results suggest potential benefits in triage accuracy and resource efficiency, critical gaps exist in validation, clinical outcome assessment, and reporting standards. Future research must prioritize prospective controlled trials with real patient data, clinical outcome measurements, and adherence to reporting guidelines.

2 July 2026

Read appraisal →
otherEvidence: Weak
45CEBM

Physiology (Bethesda, Md.)

Wearable Sensing for Clinical Physiology Monitoring: Emerging Paradigms

Advances in wearable sensing technology are driving a new era of personalized health monitoring. In contrast to the hard, rigid form factors of conventional wearable sensors, these emerging skin-interfaced systems support high-quality physiological measurements across biophysical, biochemical, and kinematic signals of interest. These platforms enable continuous monitoring of complex physiological processes with unprecedented detail as a result of a seamless, conformal skin interface and advanced wireless communications capabilities. These platforms integrate flexible materials, miniaturized electronics, and wireless communication to provide detailed physiological data during daily activities. This review examines how skin-interfaced wearables are advancing patient care, remote monitoring, and large-scale health studies. We highlight critical barriers to clinical adoption including interpreting data, validating devices, and integration into health care systems. Key opportunities include sustainable manufacturing, point-of-care fabrication, and development of disease-specific digital biomarkers. By addressing these challenges through collaboration among engineers, clinicians, and data scientists, wearable sensors can expand patient access to advanced physiological monitoring and transform personalized medicine.

2 July 2026

Read appraisal →
otherEvidence: Weak
70CEBM

Advances in skin & wound care

Mobile Health App Needs Among Patients With Diabetic Foot Ulcers in China: A Qualitative Study From the Perceptions of Patients, Caregivers, and Health Care Professionals

OBJECTIVE: To explore the specific needs of patients with diabetic foot ulcers (DFUs), caregivers, and health care professionals (HCPs) for a mobile health (mHealth) app, aiming to inform the design and development of effective mHealth service solutions. METHODS: This descriptive qualitative study was conducted from June to September 2024 in the wound care clinics of 2 local hospitals. Participants included patients with DFUs, caregivers, and HCPs directly involved in their care. Interview data were analyzed, synthesized, and refined using content analysis. RESULTS: Five key themes emerged: the pressing need to implement mHealth app services, convenient and personalized access to information, continuous and specialized health guidance, a multidisciplinary approach to disease management, and free access alongside privacy and legal protections. CONCLUSIONS: This study provides valuable insights for the design and development of an mHealth app for DFU. During the development process, it is essential to consider user needs and ensure the app meets the expectations of patients and related groups for personalized, continuous, and specialized health guidance; free access; privacy protection; and multidisciplinary collaboration. A balance should be struck between convenience and security of the app's features, to encourage user engagement and enhance the app's value and effectiveness.

2 July 2026

Read appraisal →
Systematic ReviewEvidence: Weak
15CEBM

Diagnostic and interventional imaging

Artificial intelligence in emergency musculoskeletal imaging: A critical review of current applications.

Artificial intelligence (AI) is increasingly shaping emergency musculoskeletal imaging, where rapid and accurate diagnosis is often challenged by high imaging volumes and time pressure. These constraints increase the risk of missed injuries and underscore the need for tools that support faster and reliable assessments. AI systems show promise in improving workflow efficiency by prioritizing urgent studies, guiding modality selection, and reducing reporting delays. Deep learning models can enhance abnormality detection by identifying fractures, soft tissue injuries, and infectious processes, and can provide structured classifications that support clinical decision making. Pediatric-focused AI systems address age-specific developmental considerations and offer valuable support for clinicians with varying levels of pediatric musculoskeletal expertise. Large language models further expand the role of AI by improving report clarity, generating structured impressions, and facilitating communication with clinicians, patients, and families. Despite these advances, challenges remain, including limited external validation, dataset bias, and medicolegal considerations. This review summarizes current AI applications across these domains and highlights key strengths, limitations, and future directions for safe and effective integration into emergency musculoskeletal imaging.

2 July 2026

Read appraisal →
otherEvidence: Weak
65CEBM

Journal of medical Internet research

Using a Large Language Model to Support Thematic Analysis of Patient Experiences in Chronic Illness Management: Comparative Qualitative Study

BACKGROUND: Qualitative health research often focuses on how patients experience and manage chronic illnesses, a topic that has been extensively studied in the literature. With the emergence of large language models (LLMs), such as Claude (Anthropic PBC) and ChatGPT (OpenAI), new opportunities are arising to support and scale the thematic analysis of narrative health data. However, their role and added value compared to traditional human-led approaches remain underexplored, particularly in complex clinical contexts such as multimorbidity. OBJECTIVE: We aim to evaluate the methodological contribution of LLM-assisted analysis by examining its ability to replicate and extend established qualitative insights, in comparison with traditional thematic analysis. METHODS: Semistructured interviews were conducted with 30 individuals living with two or more chronic illnesses. Transcripts were analyzed using both manual thematic coding and Claude 3.5 Sonnet. A structured comparison was conducted to identify shared and unique themes across the two approaches. The analysis examined thematic overlap, differences in subtheme identification, and variation in the level of detail between the methods. RESULTS: Both approaches identified similar core themes related to the patient experience, including health care navigation and challenges, support systems and family dynamics, and emotional challenges and coping. Manual analysis produced more contextually detailed interpretations, while the LLM approach identified a larger number of subthemes. Each method also revealed distinct themes: the manual analysis included themes such as faith, caregiving roles, and a proactive mindset, whereas the LLM identified themes such as future planning and multiple health conditions. The findings show both similarities and differences between the two approaches. The LLM analysis also demonstrated efficiency in processing large volumes of qualitative data. CONCLUSIONS: A hybrid approach that integrates artificial intelligence-assisted and human-led thematic analysis can enhance both analytical depth and scalability. These findings support the use of LLMs as a complementary tool in qualitative research, while highlighting the importance of combining automated pattern detection with human interpretation.

30 June 2026

Read appraisal →
observationalEvidence: Weak
55CEBM

JMIR cancer

Large Language Models for Breast and Cervical Cancers Communication: Mixed Methods Evaluation Study Assessing Linguistic Quality, Safety, and Accessibility

BACKGROUND: Effective communication about breast and cervical cancers remains a public health challenge, with widespread misinformation and barriers to cancer-related language understanding. Large language models (LLMs) offer potential for scalable health communication, yet trade-offs between quality, safety, and accessibility of general-purpose and medical-domain LLMs remain underexplored. OBJECTIVE: This study aimed to propose a comprehensive evaluation framework and systematically assess the performance of LLMs in generating breast and cervical cancer information, with a focus on linguistic quality, safety and trustworthiness, and communication accessibility and affectiveness. METHODS: This mixed methods evaluation study assessed outputs from 5 general-purpose and 3 medical LLMs using real-world breast and cervical cancer-related questions curated from publicly available medical datasets. LLM-generated responses were evaluated in a controlled offline setting. Primary outcomes included linguistic quality (fluency, coherence, and accuracy), safety and trustworthiness (toxicity, bias, and harm potential), and communication accessibility and affectiveness (readability, empathy, and clarity). Qualitative ratings were performed by domain experts, while quantitative metrics were compared across models. Statistical analyses included Welch ANOVA to detect differences in metric scores, Games-Howell tests for pairwise comparisons, and Hedges g to assess effect sizes. RESULTS: General-purpose LLMs, particularly Llama 3 and Gemma, demonstrated superior linguistic quality and affectiveness but often produced complex outputs that may limit accessibility. In contrast, medical LLMs (eg, MedAlpaca and BioMistral) generated simpler content suitable for broader audiences but scored lower in safety and empathy due to higher levels of hallucination, bias, and toxicity. CONCLUSIONS: While LLMs show promise for improving digital cancer communication, our findings reveal a trade-off between domain specialization and overall communication quality and safety. Future development of health-focused LLMs should prioritize hybrid modeling strategies to enhance trust, clarity, and clinical relevance in patient-facing tools.

29 June 2026

Read appraisal →
Systematic ReviewEvidence: Weak
40CEBM

Journal of medical Internet research

Applications, Challenges, and Future Directions of Large Language Models in Health Care Communication: Scoping Review

BACKGROUND: Effective health care communication is crucial in the medical field. However, effective communication in clinical practice still faces numerous obstacles, and large language models (LLMs) offer various possibilities for improving the quality of medical communication. To date, there are no published reviews on the use of LLMs in health care communication. OBJECTIVE: This review sought to summarize the applications and challenges of LLMs in health care communication and to identify directions for future research. METHODS: A comprehensive literature search was conducted in PubMed, Embase, Web of Science, and the Cochrane Library from January 2018 to November 2025. The search and selection process followed the PRISMA-ScR (Preferred Reporting Items for Systematic Reviews and Meta-Analyses extension for Scoping Reviews) guideline and the PRISMA-S (Preferred Reporting Items for Systematic Reviews and Meta-Analyses literature search extension) checklist. Eligible studies used LLMs to facilitate health care communication among the public, patients, and clinicians. Following rigorous data extraction and cross-checking, we conducted a quantitative analysis of characteristics of the included literature. Furthermore, using communication accommodation theory as a framework, we identified application patterns of LLMs in health care communication and summarized current challenges and future directions. RESULTS: Ninety-six studies were included in this review, all published between 2023 and 2025, summarizing 4 patterns of LLM application in health care communication: transforming medical information (n=30), facilitating dynamic interaction (n=38), empowering communication capabilities (n=10), and optimizing clinical workflows (n=18). The role of LLMs in health care communication is undergoing a paradigm shift from "static information processing" to "dynamic intelligent interaction." Although they show great promise for practical applications, current evaluation methods and dimensions exhibit significant heterogeneity. Furthermore, LLMs still face multiple challenges in their practical application in health care communication, including technical reliability issues, social trust and adoption, interaction and access barriers, and clinical integration challenges. CONCLUSIONS: Unlike previous studies that merely touched upon the challenges and future directions, this scoping review uses communication accommodation theory to systematically map the application patterns and developmental landscape of LLM-mediated health care communication. Health care communication powered by LLMs holds significant innovation potential and is currently still in the early stages of rapid development. Future research should focus on optimizing model performance, strengthening ethical governance frameworks, enhancing human-machine collaboration models, and ensuring responsible application of LLMs in health care through rigorous empirical validation.

28 June 2026

Read appraisal →
observationalEvidence: Weak
25CEBM

Journal of health organization and management

Unlocking potential: impact of mHealth affordance realization on patient well-being

PURPOSE: This study investigates how realized affordances of mobile health technologies influence the relationship between technology use, patient empowerment and well-being in diabetes self-management. DESIGN/METHODOLOGY/APPROACH: The research proposes a technology use-empowerment-well-being (TEW) model based on affordance theory. Data were collected through a survey of 257 diabetes patients using mobile technologies for self-management and analyzed using structural equation modeling with SmartPLS. FINDINGS: Results indicate that patient empowerment fully mediates the relationship between mobile technology use and patient well-being. Importantly, realized affordances moderately strengthen the relationship between technology use and patient empowerment, demonstrating that patients derive greater benefits when they actively utilize technology's capabilities for diabetes management. RESEARCH LIMITATIONS/IMPLICATIONS: Reliance on self-reported data and the cross-sectional design limit causal inferences about the evolution of technology-empowerment-well-being relationships over time. PRACTICAL IMPLICATIONS: Healthcare technology designers should focus on creating clear, tangible benefits for patients that support specific self-management behaviors. Healthcare providers should educate patients about the benefits of technology and support them in realizing these affordances through training and ongoing assistance. ORIGINALITY/VALUE: This study integrates affordance theory with patient empowerment concepts, providing empirical evidence that realized affordances - not merely technology availability - are critical for improving self-management outcomes. The TEW model introduces a novel framework for understanding how technology empowers patients and enhances well-being in chronic disease management.

27 June 2026

Read appraisal →
qualitativeEvidence: Weak
40CEBM

JMIR formative research

Supporting Student Mental Health With the Safespace Generative AI Chatbot: Mixed Methods Feasibility Study

BACKGROUND: Generative artificial intelligence (GenAI) chatbots have the potential to provide personalized mental health support to individuals at scale. OBJECTIVE: This study evaluates the feasibility and usage patterns of the Safespace GenAI chatbot, an artificial intelligence (AI)-driven smartphone app that offers a large language model-powered interactive chatbot to support mental health. METHODS: Using a mixed methods approach, we explored baseline attitudes toward GenAI chatbots and chatbot usage patterns, conducted a qualitative content analysis of participants' experiences, and descriptively assessed patterns related to preintervention depressive symptoms. The study included an initial sample of 42 university students, 20 of whom actively used the chatbot over 2 to 4 weeks, generating 286 user-chatbot interactions. RESULTS: Preintervention surveys indicated that the majority of participants anticipated that the chatbot would be helpful (27/42, 64%) and that they trusted its privacy safeguards (39/42, 93%). Usage patterns suggested that the highest levels of interaction occurred early in the morning and late at night, when peer and professional support may be inaccessible. The qualitative analysis indicated that participants appreciated using the chatbot for reflection as a blended-care tool between their counseling sessions, while also naming technical barriers and specific design needs required to sustain engagement. In addition, our exploratory analyses descriptively showed that participants with elevated depression scores engaged in emotional disclosure during 99% (38 sessions with 8 participants) of their sessions, compared to 84% (26 sessions of 12 participants) of those with low symptoms. Due to the small sample size, future adequately powered studies are needed to inferentially examine these observed patterns. CONCLUSIONS: These findings provide initial insights into the usage and engagement dynamics of the Safespace GenAI chatbot and highlight directions for future research to optimize GenAI-driven mental health interventions.

27 June 2026

Read appraisal →
otherEvidence: Weak
75CEBM

JMIR research protocols

Cross-Cultural Perspectives on Metabolic Syndrome Management and Attitudes Toward Using Digital Health Tools in Two Distinct Populations: Protocol for a Qualitative Descriptive Study

BACKGROUND: Metabolic syndrome (MetS) is defined by the presence of at least 3 out of 5 clinical risk factors, including abdominal obesity, elevated blood pressure, high fasting glucose, elevated triglycerides, and low high-density lipoprotein cholesterol. Individuals with MetS face significantly increased risks of cardiovascular disease, type 2 diabetes, and all-cause mortality. OBJECTIVE: This study aims to explore the multilevel factors influencing individuals' engagement in self-management of MetS. Specifically, it seeks to identify key barriers and facilitators to effective self-management and to assess participants' attitudes toward the use of digital health tools in supporting disease management. METHODS: This qualitative descriptive study will be conducted at 2 international sites: the University of Memphis (United States) and King Saud bin Abdulaziz University for Health Sciences (Saudi Arabia), in collaboration with local primary care clinics. Participants will be recruited using purposive sampling with a maximum variation strategy, aiming for 20 to 30 individuals per site who meet eligibility criteria. Data will be collected through semi-structured, 60-minute one-on-one interviews. An abductive thematic analysis approach (integrating both inductive and deductive reasoning) will be used to analyze the data with NVivo 15 software. RESULTS: This study will begin participant enrollment in August 2026. Thematic analysis will be used to examine barriers and facilitators to MetS self-management and participants' views on digital health tools across 2 international sites. CONCLUSIONS: Findings will offer culturally grounded insights into how social, dietary, and familial factors influence chronic disease management. This knowledge can inform the design of equitable, context-sensitive digital health interventions. INTERNATIONAL REGISTERED REPORT IDENTIFIER (IRRID): PRR1-10.2196/91903.

27 June 2026

Read appraisal →
otherEvidence: Weak
55CEBM

Journal of medical Internet research

Personal Health Large Language Models and the Negotiation of Medical Authority in Clinical Care: Opportunities, Risks, and Governance

Personal health large language models (PH-LLMs) are patient-facing conversational systems that synthesize user-entered information, patient-generated health data, wearable data, and selected personal health records-where users choose to connect them-into personalized, longitudinal, action-oriented health narratives. Unlike generic health chatbots that mainly provide one-off responses to isolated questions, PH-LLMs may generate continuing interpretations, priorities, and candidate next steps that patients bring into clinical encounters. In contrast to electronic health record-tethered clinical artificial intelligence (AI), they often originate outside institutional oversight and may be selected, used, or trusted by patients before professional review. This viewpoint examines how PH-LLMs may reshape the negotiation of medical authority by contributing to a shift from the traditional dyadic clinician-patient relationship toward a triadic model of negotiated authority, in which clinicians may increasingly need to mediate among clinical evidence, patient values, and algorithmic narratives. PH-LLMs may support patient participation by organizing symptoms, contextualizing home-monitoring and wearable data, improving health literacy, assisting chronic disease self-management, and preparing patients for more collaborative visits. Patient-facing AI narratives may also introduce distinct risks. At the individual level, these include inaccurate or incomplete responses arising from imprecise queries or missing context, misinterpretation of otherwise accurate information in the absence of clinical context, and contextually biased or poorly matched advice across demographic, cultural, linguistic, disability-related, or socioeconomic contexts. At the system level, they include authority conflict when AI recommendations diverge from clinical judgment, fragmentation of clinical truth, privacy and data-governance concerns, diffusion of accountability when harm results from advice produced outside clinical governance, and inequitable access to premium tools and continuous monitoring devices. To address these challenges, we propose a 3-layer clinical governance framework for patient-brought PH-LLM narratives. The first layer, evidence and provenance, makes AI-generated narratives epistemically legible by clarifying platform identity, data sources, temporal anchoring, uncertainty, and privacy-relevant data-use and retention conditions. The second layer, clinical arbitration and workflow integration, uses risk-stratified intake, proportionate documentation, escalation triggers, and equity-preserving workflows to embed PH-LLM outputs into routine care. The third layer, competence and accountability, defines the communication competencies, AI literacy supports, institutional responsibilities, vendor accountability, and risk-proportionate verification duties needed for triadic care. This framework is a conceptual and governance-oriented proposal rather than a validated clinical protocol. Future empirical work should evaluate its feasibility, documentation burden, equity effects, clinical safety impact, and acceptability among patients, clinicians, and health systems. Governed through these interdependent layers, PH-LLMs may serve as supporting infrastructure for safer, person-centered longitudinal care.

26 June 2026

Read appraisal →
Systematic ReviewEvidence: Weak
40CEBM

Journal of medical Internet research

The Emerging Roles of AI in Self-Directed Stress Management: Systematic Review

BACKGROUND: Stress is widespread and carries substantial mental health, social, and economic burdens. Yet, access to clinician-led stress management remains constrained by service capacity, cost, and stigma. In response, artificial intelligence (AI)-enabled tools have rapidly proliferated as scalable, self-directed options. However, evidence on how these systems support stress management outside formal clinical settings remains fragmented. OBJECTIVE: This systematic review aimed to synthesize empirical evidence on how AI-enabled technologies are used for self-directed stress management. We mapped the emerging functions of these tools, the psychological frameworks informing their design, the populations and settings studied, and the outcomes reported. METHODS: We conducted a PRISMA (Preferred Reporting Items for Systematic Reviews and Meta-Analyses)-compliant systematic review of English-language studies published between 2000 and 2025. Six databases were searched (APA PsycINFO, PubMed, MEDLINE, Scopus, Web of Science Core Collection, ProQuest, and Google Scholar). RESULTS: Of 3008 records identified, 35 studies met the inclusion criteria. The methodological quality of included studies was critically appraised using the Mixed Methods Appraisal Tool (version 2018). Findings illustrated that AI-supported stress management can operate through 5 core functions, including psychological intervention, behavioral support, psychoeducation, companionship, and emotional support, and stress monitoring, detection, and triage. Across the reviewed studies, these functions supported self-directed stress management by helping users identify stress, regulate responses, and engage in coping outside formal clinical care. CONCLUSIONS: AI-enabled systems show preliminary promise for supporting self-directed stress management through multiple user-facing functions grounded in established psychological frameworks.

26 June 2026

Read appraisal →
Systematic ReviewEvidence: Weak
55CEBM

Journal of medical Internet research

Accessibility of Digital Financial Applications for People With Visual Impairment: Scoping Review

BACKGROUND: Routine financial activities are now conducted primarily through digital channels. Many such systems remain inaccessible to more than 2.2 billion people globally living with vision impairment, limiting independent financial management. Constrained access can create financial strain and social disadvantage, reducing access to health-enabling resources, and contributing to avoidable health inequities. OBJECTIVE: This scoping review maps evidence on the accessibility of digital financial services for individuals with visual impairment (VI) as a digital determinant of health. We synthesized barriers and facilitators, characterized study designs, settings, and populations, and identified evidence gaps to inform inclusive design, digital health research priorities, and policy. METHODS: A scoping review was conducted using the Joanna Briggs Institute framework and reported in line with PRISMA-ScR (Preferred Reporting Items for Systematic Reviews and Meta-Analyses extension for Scoping Reviews) guidelines. Eight databases (PubMed, MEDLINE, CINAHL, Scopus, Web of Science, Business Source Complete, ProQuest, and IEEE Xplore) were searched for peer-reviewed papers in English published between 1995 and 2026. Searches featured controlled vocabulary and free-text terms structured in 3 conceptual blocks (VI, digital financial services, and accessibility or usability). A random sample of 20% of titles, abstracts, full texts, and included studies was independently screened or charted by 2 reviewers to calibrate decisions; the remainder were screened and charted by a single reviewer. Data were charted using a standardized extraction form, and results were synthesized descriptively and thematically. RESULTS: Twenty-three studies met the inclusion criteria. Studies were conducted across 12 countries, with the largest number from India (n=7), Indonesia (n=2), Thailand (n=2), and the United States (n=2). Study designs included qualitative studies (n=6), mixed methods studies (n=1), cross-sectional studies (n=4), nonrandomized experimental studies (n=2), and technical or design-focused evaluations (n=6). One study was a large population survey (n=19,136), and the remaining studies with human participants had sample sizes ranging from 4 to 36 participants. Accessibility barriers were reported across all platform types, with authentication-related barriers described in 18 studies and screen reader incompatibility in 17 studies. Reported barriers included reliance on sighted assistance for tasks such as login, verification, and payments, compromising privacy and independence. Facilitators included assistive technology support, logical navigation order, nonvisual feedback mechanisms, and accessible authentication alternatives. Evidence mapping revealed recurrent barrier patterns across Android, iOS, and web platforms. No longitudinal or intervention-based evaluations were identified. CONCLUSIONS: This review provides a focused synthesis of accessibility evidence at the intersection of digital financial services and VI, a domain addressed by neither prior digital accessibility reviews nor financial inclusion for people with disabilities. Authentication methods, interface labeling, and navigation were identified as persistent cross-platform accessibility barriers. The findings carry implications for financial technology developers, accessibility auditors, and policymakers implementing accessibility legislation and extend the digital determinants of health framework by demonstrating how inaccessible financial technology may compound health inequities.

25 June 2026

Read appraisal →
otherEvidence: Weak
30CEBM

Journal of medical Internet research

Solidarity or Segregation? ChatGPT Health and US Health Care Disparities.

ChatGPT Health, an artificial intelligence feature launched by OpenAI in January 2026, integrates personal medical records and consumer health data into a large language model-based chatbot. It positions itself as a digital health tool to help users navigate a fragmented US health care system marked by inaccessible insurance, medical deserts, and workforce shortages. We argue that, rather than closing these gaps, ChatGPT Health risks widening what we call the solidarity gap in health care. By appearing to address unmet needs, it may obscure and entrench the structural conditions that produce health care disparities. We identify specific risks to patient safety, such as the legitimization of dangerous self-rationing and self-medication behaviors, inaccurate triage of clinical emergencies, erosion of effective health communication, and the reinforcement of confirmation and anchoring bias. These risks are particularly acute for individuals facing structural barriers to care. To prevent ChatGPT Health and similar large language model-based health tools from deepening inequities, we propose solidarity-based policy interventions such as upstream reforms that rebuild the institutional foundations of access and health equity and downstream safeguards that embed artificial intelligence tools within accountable, equitable health care infrastructures as complements to rather than substitutes for human care.

24 June 2026

Read appraisal →
otherEvidence: Moderate
50CEBM

Current psychiatry reports

Artificial Intelligence in Child and Adolescent Psychiatry: A Narrative Review of Recent Clinical Applications and Ethical Considerations

PURPOSE OF REVIEW: This narrative review examines recent literature of artificial intelligence (AI) in child and adolescent psychiatry. With increasing mental health disorders in young people along with persistent workforce shortages, AI has emerged as a potential tool to improve efficiency and support clinical decision making. However, using AI raises important ethical concerns which are summarized in our review. RECENT FINDINGS: Most AI applications in child and adolescent psychiatry remain early in development and are not ready for routine clinical use. AI scribes are likely impractical in many child psychiatry settings because of multiparty visits, consent concerns, and sensitive clinical discussions. Many multimodal diagnostic tools using machine learning still require further testing and can be impractical when relying on costly diagnostics like neuroimaging. Similarly, AI-assisted therapeutics requiring physical hardware like robotics and virtual or augmented reality devices can also be prohibitively expensive. One of the few exceptions includes AI-enabled video and eye-tracking approaches for autism diagnosis. Chatbots and robot companions may provide modest improvements for depression, but evidence remains limited in the pediatric population with risk of serious harm. Concerns of AI include misinformation, algorithmic biases, privacy risks, crisis mismanagement, and excessive emotional attachments to chatbots. AI may eventually support child and adolescent psychiatry, but current evidence supports cautious, supervised use rather than broad clinical adoption. Clinicians should help families understand AI's limits, encourage digital literacy, and ensure that AI remains an adjunct to human care rather than a substitute.

22 June 2026

Read appraisal →
qualitativeEvidence: Insufficient
55CEBM

JMIR research protocols

Design, Development, and Validation of a Chatbot to Support Health Care Professionals Experiencing Workplace Aggression: Protocol for a Mixed Methods Study

BACKGROUND: Workplace violence against health care professionals has increased worldwide, leading to negative psychological, professional, and organizational outcomes. Despite existing prevention and reporting programs, underreporting and lack of accessible, confidential support persist. Digital health tools, including chatbots, may offer scalable support, guidance, and follow-up for affected professionals. OBJECTIVE: This study aims to design, develop, and validate a chatbot (Sanidad Segura) to assist health care professionals who experience workplace aggression and evaluate its usability, readability, and exploratory indicators of perceived usefulness and support in a pilot study. METHODS: This study will follow a mixed methods design conducted in two main phases: (1) design, development, and content validation of the chatbot based on literature review, institutional protocols, and expert consensus; and (2) pilot-testing, including usability and readability assessment using standardized instruments, as well as feasibility and acceptability evaluation among health care professionals working in emergency and critical care settings in Almería (Spain). The study is aligned with the Medical Research Council framework for complex interventions, incorporating development and feasibility stages. Quantitative data will be collected using the System Usability Scale and Inflesz readability scale. Qualitative data will be collected through semistructured interviews and analyzed using thematic analysis to explore user experience and identify barriers to and facilitators of use. RESULTS: The study has been funded for a 2-year period starting on December 18, 2024. Quantitative outcomes will include usability scores (System Usability Scale), readability scores (Inflesz), and participants' sociodemographic characteristics. Qualitative findings will identify themes related to usability, user experience, and suggestions for improvement. Integration of quantitative and qualitative findings will be conducted through triangulation to provide a comprehensive understanding of the usability, acceptability, and readability of the chatbot. CONCLUSIONS: This study addresses the increasing incidence of workplace violence against health care professionals through the development of a new chatbot (Sanidad Segura). This intervention seeks to facilitate the identification, support, and follow-up of affected individuals while minimizing the adverse effects of such events on their physical and psychological well-being, social interaction, and professional performance. Sanidad Segura will enable confidential case reporting and provide access to tailored medical, psychological, and legal resources, as well as information about institutional support services. This project represents a crucial step toward implementing an integrated digital framework for the detection, management, and prevention of workplace violence in health care settings. INTERNATIONAL REGISTERED REPORT IDENTIFIER (IRRID): PRR1-10.2196/92511.

19 June 2026

Read appraisal →
observationalEvidence: Weak
55CEBM

Archives of gynecology and obstetrics

Comparing clinical decision-making between colposcopists and large language models in cervical dysplasia management: a pilot prospective multicenter study

PURPOSE: This prospective multicenter study aimed to compare the decision-making abilities of board-certified colposcopists and two commercially available large language models (LLM), ChatGPT-4o and ChatGPT-5, in cervical dysplasia management. METHODS: Twenty-three anonymized real-life patient cases with multiple-choice (MC) questions regarding treatment decisions were used to assess answer quality. Ten board-certified colposcopists and the two LLMs addressed the MC questions. The gold standard was defined by two guideline authors. LLMs were prompted to justify their responses. Concordance rates were calculated and compared across all questions and histopathological subgroups, including cervical intraepithelial neoplasia (CIN), unspecific histopathological results, and cervical cancer cases. RESULTS: Clinicians and LLMs achieved similar overall concordance rates compared to the gold standard (69.6% for clinicians, 69.6% for ChatGPT-4o, and 65.2% for ChatGPT-5). ChatGPT-5 outperformed clinicians in precancerous lesions (81.8% vs. 66.4%), while clinicians excelled in complex cases with unspecific histopathology (86% vs. 60%). Clinicians showed a tendency to overtreat low-grade lesions (CIN I), opting for more intensive surveillance. ChatGPT-4o performed better than ChatGPT-5 in cervical cancer cases, though both models struggled with these scenarios. CONCLUSION: This study highlights the potential of LLMs as decision support tools in cervical dysplasia management, particularly for straightforward cases like precancerous lesions. However, clinicians remain superior in handling complex or ambiguous cases. The tendency of clinicians to overtreat low-grade lesions may offer the potential to test the implementation of a decision support tool for those cases. While LLMs show promise, exploring open-ended clinical scenarios and integrating retrieval-augmented generation could enhance their practical application.

19 June 2026

Read appraisal →
Systematic ReviewEvidence: Weak
40CEBM

Journal of medical Internet research

Ethical Considerations in Personal Health Large Language Models

Personal health large language models (PH-LLMs) have rapidly evolved from research prototypes into consumer-facing, data-linked systems that support symptom triage, medication questions, mental health check-ins, and longitudinal self-management. Their direct-to-consumer use without clinical oversight creates a distinct ethical risk profile that general artificial intelligence governance frameworks do not fully address. This viewpoint focuses on text-based, platform-mediated PH-LLMs and synthesizes PH-LLM-specific challenges across 6 domains: privacy, accuracy, equity, transparency, human-artificial intelligence interaction, and regulatory governance. These risks may be amplified by health literacy gaps, longitudinal data aggregation, persuasive conversational design, and fragmented oversight across the consumer-clinical boundary. Grounded in the 4 principles of biomedical ethics, we propose a governance framework that operationalizes beneficence, nonmaleficence, autonomy, and justice through design and deployment controls, including health literacy-aligned communication, crisis and pharmacological safeguards, hallucination mitigation, role disclosure, granular consent, fairness auditing, and accessible design. We further outline implementation mechanisms, including risk-tiered certification, tiered accountability, and postdeployment oversight through adverse-event reporting, transparency reporting, and independent safety evaluation. This framework is intended as an evidence-informed but partly anticipatory approach to governing PH-LLMs in personal health management.

19 June 2026

Read appraisal →