Evidence-Based Medicine

Research Appraisals

Evidence-based critical appraisals of the latest medical research, systematically evaluated using Oxford CEBM methodology.

Showing 113 appraisals

Systematic ReviewEvidence: Moderate
45CEBM

Telemedicine journal and e-health : the official journal of the American Telemedicine Association

Telehealth Competency Evaluation Tools: A Scoping Review

INTRODUCTION: Telehealth is an integral part of healthcare. Exposure to telehealth education is essential for both students and health care professionals to support its effective adoption and appropriate utilization. This scoping review aims to explore whether and which telehealth competencies are being assessed and the methods used to evaluate them among health care professionals and students. METHODS: An electronic literature search was performed using six electronic databases between January 2024 and February 2025. We included studies where telehealth competencies or their components were evaluated by some type of tool (validated or researcher created). Strict inclusion and exclusion criteria were used, and each article was evaluated by at least two reviewers. RESULTS: Out of 1,217 articles screened by title and abstract, 75 met inclusion criteria, with 36 of those selected for inclusion in this review. Participants were evaluated mostly by researcher-created tools that lacked psychometric properties or adapted nontelehealth evaluation tools. Few validated/reliable tools directly measured telehealth competencies. DISCUSSION: The results of this review illustrate there exists a lack of available validated/reliable evaluation tools to appraise telehealth competencies. Few studies included an evaluation of all telehealth competencies but rather focused only on communication and technology proficiency; none evaluated the appropriate use of telehealth or digital disparities. CONCLUSIONS: More research is needed to develop validated telehealth evaluation tools that can be used across disciplines and encompass all telehealth competencies. Additionally, true research methods are needed to adequately assess the impact of telehealth education on the performance and confidence of learners.

3 Aug 2026

Read appraisal →
Randomised Controlled TrialEvidence: Weak
35CEBM

Science advances

Human-inspired time-series health evaluation with an adaptive multimodal electronic skin

Electronic skin powered with artificial intelligence could enable next-generation robotic and medical devices, yet integrating multimodal sensors and analyzing heterogeneous, multifrequency time series remain challenging. Most wearable machine learning architectures are time-invariant and trained for a specific task, limiting transfer across modalities and users. We present a multimodal electronic skin that captures diverse physiological signs with an adaptive learning framework that rapidly generalizes to unseen tasks with minimal labeled data. Our streamlined end-to-end framework uses a spectral variational autoencoder to denoise and compress multifrequency biosignals into a shared, unified second-wise latent space that preserves the spectral-temporal structure, followed by a transformer to capture temporal dependencies to support diverse downstream tasks with data-efficient learning. We demonstrate robust adaptation with 94.7% accuracy in activity recognition and 90.2% precision in fatigue assessment across various users and daily activities regardless of device and user variations, highlighting a scalable route to generalized physiological time-series analytics and human performance assessments.

3 Aug 2026

Read appraisal →
Systematic ReviewEvidence: Weak
55CEBM

Journal of medical Internet research

Barriers and Facilitators to Implementing Digital Health Technologies for Remote Management of NCDs in Rural Areas: Mixed Methods Systematic Review

BACKGROUND: Digital health technologies (DHTs) have the potential to improve care delivery and outcomes for patients with noncommunicable diseases. Yet their implementation in rural settings remains uneven, and the factors influencing uptake are not well understood. OBJECTIVE: This mixed methods systematic review aimed to identify barriers and facilitators influencing the implementation and use of DHTs for remote management of noncommunicable diseases in rural areas. METHODS: We searched Medline, Embase, and CINAHL from inception to February 12, 2026, using terms related to digital health, noncommunicable diseases, and rural settings. Following the Joanna Briggs Institute methodology for mixed-method systematic review, we synthesized quantitative and qualitative studies. Barriers and facilitators were categorized using the Consolidated Framework for Implementation Research, and study quality was appraised using the Mixed Methods Appraisal Tool. RESULTS: From the initial 1491 records, 14 studies met the inclusion criteria, with most conducted in high-income countries (n=11). Key barriers included technical challenges (software instability and hardware issues), poor internet connectivity, financial constraints, and workforce constraints, such as staff shortages and heavy workloads. Key facilitators included user-friendly technology design, strong leadership, effective teamwork, and ongoing communication. Evidence was predominantly qualitative, with only limited quantitative data available. CONCLUSIONS: DHTs show promise for improving access and continuity of care for cardiovascular disease, hypertension, and diabetes in rural settings; however, their impact is constrained by structural inequities, including limited broadband access, workforce shortages, and financial fragility. These findings highlight important implications for research, policy, and practice, including the need for rigorous mixed methods evaluations sensitive to rural contexts, long-term equity-oriented financing mechanisms, and strengthened organizational readiness to support effective DHT uptake.

2 Aug 2026

Read appraisal →
Systematic ReviewEvidence: Weak
50CEBM

Journal of tissue viability

Clinical decision support tools in wound management: A scoping review of existing and emerging tools.

AIM: This scoping review aimed to identify and map clinical decision support tools (both digital and non-digital) used by clinicians for wound management in acute, primary, and community health settings globally, with a focus on the use of artificial intelligence, and the barriers and enablers to implementation. MATERIALS AND METHODS: Studies published between January 2015, and August 2024 were identified through systematic searches of Scopus, Embase, MEDLINE, and CINAHL, conducted in accordance with JBI methodology and reported using the PRISMA-ScR guidelines. Eligible studies included quantitative, qualitative, and mixed-methods research examining non-digital and AI-supported clinical decision support tools for wound management across healthcare settings. RESULTS: Seventeen studies were included, two evaluating AI-supported clinical decision support tools. Most clinical decision support tools had structured wound assessment and treatment planning, though evidence of clinical effectiveness was limited, with only one tool fully validated. Adoption by nurses was influenced by experience, trust, training, and workflow integration, with senior nurses less likely to rely on clinical decision support tools. AI-enabled tools, and a convolutional neural network-based model, improved assessment consistency, documentation, and workflow efficiency. Key barriers included concerns about trust, clinical autonomy, and usability. CONCLUSIONS: Both traditional and AI-supported clinical decision support tools are used for chronic wound management across acute, primary, and community care, but evidence of effectiveness and validation remains limited. The absence of experimental studies highlights the need for rigorous evaluation, clinician education, and strategies to support integration of AI-enabled clinical decision support tool into routine practice.

2 Aug 2026

Read appraisal →
Systematic ReviewEvidence: Insufficient
5CEBM

Cognitive, affective & behavioral neuroscience

A review of current capabilities and future directions in machine-based emotion recognition

The ability of machines to recognize emotions automatically is becoming increasingly significant across many domains where emotional understanding is essential. Such technology is applied in customer interaction, marketing, healthcare, education, the automotive industry, entertainment, and security. Providing real-time insights into human affective states improves user engagement and enables systems to respond more intelligently. Nevertheless, progress in this field is hindered by the inherent complexity of emotions, cultural differences in expression, and technical limitations that make accurate detection challenging. This paper delivers a broad review of contemporary approaches to emotion recognition. It highlights techniques based on facial expression analysis (FER), oculometrics (OM), microexpressions identification (MER), and speech analysis (SER). Further attention is given to methods involving body posture, gesture, and gait, as well as tactile interaction, text-based emotion recognition, and methods based on self-reporting. In addition, physiological signal-driven methods are discussed in depth, including respiration signals (RS), galvanic skin response (GSR), electroencephalography (EEG), electromyography (EMG), skin temperature (SKT), cardiac signals (ECG, PPG, HRV), and touch dynamics (TD) analysis. This comprehensive overview lays the foundation for advancing research on machine-based emotion recognition.

2 Aug 2026

Read appraisal →
observationalEvidence: Moderate
75CEBM

Journal of the American Medical Informatics Association : JAMIA

Sociodemographic bias in large language model clinical trial screening

OBJECTIVE: To assess whether large language model (LLM)-based clinical trial screening judgments vary by patient sociodemographic characteristics. MATERIALS AND METHODS: We conducted a cross-sectional evaluation of Phase II-III US adult randomized controlled trial (RCT) protocols (2023-2024). Physician-validated clinical vignettes were evaluated in a control version and 33 sociodemographic identity variants differing only by labels. Nine LLMs assessed eligibility and related domains. Mixed-effects models estimated adjusted differences vs control. RESULTS: Across 58 protocols and 5.3 million evaluations, eligibility judgments were largely stable across identities. Race and ethnicity showed minimal effects after accounting for socioeconomic status. Homelessness produced the largest negative eligibility shift and pronounced effects in adherence, resources, and trust. DISCUSSION AND CONCLUSION: LLMs applied explicit eligibility criteria consistently, but disparities emerged in domains requiring inference about behavior or resources, underscoring the need for careful deployment to promote fair trial access.

2 Aug 2026

Read appraisal →
Systematic ReviewEvidence: Weak
15CEBM

Clinical & translational oncology : official publication of the Federation of Spanish Oncology Societies and of the National Cancer Institute of Mexico

Research progress of machine learning applications in gastric cancer diagnosis and therapy

Gastric cancer (GC), a malignant neoplasm originating from the gastric mucosal epithelium, represents one of the most prevalent cancers worldwide. Early detection is critical for improving treatment outcomes and patient prognosis. Recent advances in artificial intelligence (AI), particularly in machine learning, have introduced powerful computational and analytical capabilities that are increasingly being applied in GC research. Machine learning algorithms have shown considerable promise in enhancing the accuracy of GC diagnosis and optimizing therapeutic strategies. This review provides a concise overview of progress in machine learning applications within oncology, examines their current role and clinical utility in GC diagnosis and treatment, and highlights the transformative potential of machine learning in advancing GC management and patient care.

2 Aug 2026

Read appraisal →
Systematic ReviewEvidence: Moderate
60CEBM

Journal of the American Medical Informatics Association : JAMIA

Alert fatigue measurement in clinical decision support: a systematic review

BACKGROUND: Alert fatigue is defined as alert dismissals due to excessive or irrelevant alerts and is frequently cited as a barrier to clinical decision support system use and impact. However, the criteria for determining the presence or absence of alert fatigue are poorly defined. The objective of this systematic review of systematic reviews was to identify operationalized definitions and measures of alert fatigue or alert-related metrics. METHODS: Systematic reviews reporting at least one alert-related metric or measure/operationalization of alert fatigue for physician-directed electronic alerts were included. The Cochrane Library, Embase, and PubMed were searched from database start to 2024. The Revised Assessment of Multiple Systematic Reviews was used to assess study quality and risk of bias. Data were synthesized narratively and with descriptive statistics. RESULTS: A total of 22 studies were included in the review. Studies reported between 1 and 11 alert metrics. Studies were most often of medium quality. Reporting of primary study characteristics was frequently judged to be insufficient. Only one article reported an operational definition of alert fatigue. The most common alert metrics were quantity, override rate, and acceptance rate. DISCUSSION: Alert fatigue measurement methods are not clearly or consistently defined in systematic reviews related to alert fatigue in clinical decision support. Reporting of other primary study characteristics is often limited. We recommend that future efforts use a significant, sustained decrease in appropriate alert response rates from an established baseline as a measure of alert fatigue.

2 Aug 2026

Read appraisal →
otherEvidence: Moderate
65CEBM

Balkan medical journal

Bias and Fairness Across the Healthcare AI Lifecycle: A Clinician-Oriented Review

Artificial intelligence (AI) is increasingly being investigated and, in selected clinical settings, implemented to support diagnosis, triage, and workflow optimization. Although these systems have the potential to improve access, consistency, and efficiency, they may also reproduce or amplify health inequities when bias is introduced during development, evaluation, implementation, or postdeployment use. This clinician-oriented narrative review adopts a practical lifecycle approach to explain how algorithmic unfairness becomes clinically relevant, how clinicians can recognize it, and how institutions can mitigate its impact. We first outline the ethical, clinical, and mathematical dimensions of fairness. We then examine fairness risks and sources of bias across six stages of the healthcare AI lifecycle: problem formulation, data generation, model development, evaluation, implementation, and postdeployment monitoring and governance. Key mechanisms include biased proxy outcomes, unrepresentative or error-prone data and labels, model shortcut learning, hidden stratification, distribution shift, and human-AI interaction effects (e.g., automation bias and alert fatigue), all of which can create feedback loops and contribute to fairness drift over time. For each stage, we identify clinician-facing red flags and practical mitigation strategies, including defining clinically meaningful outcomes, using representative and well-documented datasets, conducting subgroup-stratified evaluations, performing external and prospective validation, justifying decision thresholds, implementing safeguards for human-AI interactions, and maintaining continuous postdeployment monitoring, including postmarket surveillance for regulated medical devices. Fairness cannot be ensured through a single metric, publication, regulatory clearance, or one-time validation. Instead, equitable healthcare AI requires transparent design, rigorous evaluation, local governance, and ongoing monitoring across diverse populations, clinical sites, devices, workflows, and time. Fairness should therefore be regarded as a continuous clinical and institutional responsibility rather than a downstream technical consideration.

1 Aug 2026

Read appraisal →
Randomised Controlled TrialEvidence: Weak
35CEBM

Scientific reports

Manual federated simulation for multiple sclerosis integrating XGBoost algorithm with SHAP explanation

Multiple sclerosis (MS) is a chronic autoimmune disorder of the central nervous system, underscoring the importance of early and accurate diagnosis. In this study investigates the predictive modelling of MS progression in patients with Clinically Isolated Syndrome (CIS), privacy-preserving for a federated and explainable Machine Learning (ML) framework. To address missing data while preserving inter-feature dependencies, Multivariate Imputation by Chained Equations (MICE) with iterative imputers was employed. Classification was performed using the Extreme Gradient Boosting (XGBoost) algorithm. Model interpretability was developed through Explainable Artificial Intelligence (XAI) techniques, specifically Shapley Additive Explanations (SHAP). To ensure data confidentiality and simulate decentralized clinical environments, an in silico federated learning framework was applied. Experimental results demonstrated strong predictive performance, achieving 96.7% accuracy and 99% ROC-AUC during training, 92.5% accuracy in validation, and 81.8% accuracy with an AUC of 88% on the test set. For the Federated Learning (FL) simulation, the model maintained competitive performance, yielding an accuracy of 76.3% and an AUC of 83.9%. The proposed approach supports early diagnosis, enhances clinical trust through interpretability, and promotes secure data collaboration, thereby contributing to more informed and transparent clinical decision-making and improved patient care.

1 Aug 2026

Read appraisal →
Systematic ReviewEvidence: Moderate
50CEBM

ACS applied materials & interfaces

Advances and Challenges in Wearable Sensors for Health Monitoring

Analytical tools may revolutionize healthcare by enabling accessible, rapid, and decentralized testing. Wearable (bio)sensors, in particular, provide frequent or continuous patient monitoring through non- to minimally invasive measurements. This approach yields unprecedented amounts of health-related information, leading to more informed clinical decision-making and closer patient follow-up. In this mega-review article, we bring together leading researchers in the field to discuss the state of the art in wearable devices for health monitoring. We begin by providing a broad overview of the field through citation network analysis. We then review the application of chemical (bio)sensors in biofluids (e.g., sweat, saliva, tears, interstitial fluid, and cerebrospinal fluid), highlighting the challenges and advantages associated with each. Subsequently, we discuss the construction of wearable devices and their main formats (e.g., smart contact lenses, textiles, mouthguards, watches/wristbands, and implantable systems). Physical sensors are addressed in a dedicated section focusing on the assessment of heart rate, blood pressure, and body temperature. The role of soft electronics in wearable devices is also examined, as these technologies are essential for enhancing user comfort and sensor reliability, which demands advances in materials science. Furthermore, we present strategies for signal acquisition and transmission, as well as approaches for on-body energy harvesting and device self-powering. The use of artificial intelligence and machine learning is then discussed as a means of enhancing analytical performance and managing the large volumes of data generated by wearable devices. Finally, business, regulatory, and ethical considerations are examined. We expect that this review will provide an overview of sensing and biosensing technologies for health-related applications, identify promising research directions, and inspire future developments.

30 July 2026

Read appraisal →
observationalEvidence: Moderate
70CEBM

Journal of medical Internet research

Shadow AI in Swedish Health Care: Qualitative Analysis of Physicians' Free-Text Answers

BACKGROUND: The rapid emergence of artificial intelligence (AI) has outpaced its formal adoption in health care organizations, contributing to the emergence of Shadow AI, defined here as the use of unauthorized AI tools by medical professionals. Under the European Union Medical Device Regulation, AI tools used for clinical purposes must undergo conformity assessment before use; general-purpose tools such as ChatGPT have not done so, rendering their clinical application unauthorized at the regulatory level. While Shadow AI offers potential efficiency gains and higher performance, it poses significant risks to data privacy, clinical safety, and regulatory compliance. Despite its growing prevalence, empirical research on the purposes for which physicians use Shadow AI remains scarce. OBJECTIVE: This study explores the purposes for which physicians describe using Shadow AI in their work. METHODS: We conducted a cross-sectional survey of physicians employed in Swedish health care organizations (N=357; response rate~64%). Data were collected between December 2023 and January 2024 via a verified online panel. We conducted a qualitative content analysis of free-text responses on the use of unauthorized AI tools. We applied theoretical lenses from the sociology of professions and paradox theory to interpret the empirical findings. RESULTS: Physicians use Shadow AI for several purposes, which we grouped into 4 categories: clinical work and decision-making, administrative work, research and professional development, and technological interest and curiosity. More specifically, Shadow AI is used as a colleague and second opinion for clinical decision support (eg, differential diagnoses and rare cases), administrative tasks such as patient communication and documentation, and research aimed at staying up to date and exploring developments in generative AI. Physicians described using these tools compensated for perceived gaps in institutional systems, reducing workload, and accessing knowledge considered difficult to obtain through conventional channels. The findings reveal a tension between physicians' drive to improve their practice and the regulatory and organizational constraints that render such use unauthorized. CONCLUSIONS: Shadow AI used by physicians presents both opportunities and risks for health care professionals and organizations. Shadow AI indicates gaps where formal hospital systems may fail to meet health care professionals' needs and signals a way for physicians to strengthen their experience-based knowledge. It represents a renegotiation of professional boundaries, as physicians bypass institutional constraints to maintain professional efficacy. The findings highlight a paradox in which the same tools that pose regulatory and safety risks also address real gaps in clinical and administrative support, suggesting that governance approaches must account for this tension rather than relying on prohibition alone.

30 July 2026

Read appraisal →
otherEvidence: Weak
60CEBM

JMIR research protocols

Large Language Models in German Continuing Medical Education Assessments: Protocol for a Fully Crossed Experimental Study

BACKGROUND: Continuing medical education (CME) is a legal and ethical obligation for physicians in Germany. The rapid rise of large language models (LLMs) such as ChatGPT, Gemini, Claude, and Grok raises concerns about the integrity of CME assessments, as LLMs can already pass German CME tests. OBJECTIVE: This study aims to determine whether the choice of document format (searchable PDF, protected PDF, raster PDF, or vector PDF) and LLM influences the ability of LLMs to solve CME test questions at rates exceeding the passing threshold specified for each CME module (typically 70%). METHODS: In a fully crossed within-subjects repeated-measures design, 18 expired CME articles from 3 major German publishers across 6 specialties will be converted into 3 cheating-impeding PDF formats and processed alongside the original PDF files by 4 current LLMs (GPT-5, Claude Sonnet 4, Grok-4, and Gemini 3). This results in 16 model-format combinations. Each model will answer every article 3 times per file-format condition, with outcomes derived from aggregated run-level results. The primary outcome is the proportion of correctly answered questions; the secondary outcome is the pass/fail rate. RESULTS: The study has been approved by the Witten/Herdecke University Ethics Committee (S-260/2025; dated August 10, 2025) and is preregistered at the Open Science Framework. The study is supported by internal departmental resources only, and no external funding was received. Because this protocol evaluates LLMs using expired CME materials, no human participants are being recruited. Data collection is planned to begin in June 2026 and is expected to last approximately 4 weeks. At the time of manuscript submission, no data have been collected or analyzed. Results are expected to be available after the completion of data collection and statistical analysis in 2026. The analyses will quantify performance differences across document formats; these findings may inform the feasibility of nonsearchable document formats as a temporary measure to reduce LLM-enabled cheating risks in CME contexts. CONCLUSIONS: By quantifying how document format constrains LLM performance, this study aims to evaluate simple technical safeguards that may reduce artificial intelligence-assisted manipulation of CME tests and inform regulators and CME providers about how to balance assessment validity, accessibility, and responsible LLM integration into postgraduate medical education.

30 July 2026

Read appraisal →
Systematic ReviewEvidence: Moderate
70CEBM

Journal of breast imaging

Contrast-Enhanced Mammography for Breast Cancer Screening: A Systematic Review and Meta-Analysis

OBJECTIVE: To evaluate pathway-specific cancer detection and summarise screening outcomes (recall/biopsy/PPV3) for contrast-enhanced mammography when reported, and to compare them with those for low-energy mammography and breast MRI. METHODS: This systematic review and meta-analysis followed PRISMA guidelines. PubMed (MEDLINE), Embase, Scopus, and Web of Science were searched from January 1, 2011, through August 31, 2025, for studies evaluating contrast-enhanced mammography in screening or surveillance populations (including supplemental/add-on screening and surveillance of patients with a personal history of breast cancer). Diagnostic, recall, and problem-solving studies were excluded. Screening pathways were categorized as: (1) direct comparison with low-energy mammography, (2) direct comparison with MRI, and (3) add-on CEM to mammography-based screening. Random-effects meta-analyses were performed when designs were comparable. Comparisons between CEM and MRI were summarized qualitatively because of heterogeneity and limited cancer events across studies. Single-arm screening studies were summarized descriptively. Risk of bias was assessed using Quality Assessment of Diagnostic Accuracy Studies-2 (QUADAS-2). RESULTS: Fourteen studies met eligibility criteria for qualitative synthesis; seven screening and surveillance studies contributed to the meta-analysis. Compared with low-energy mammography, CEM detected 9.3 additional cancers per 1,000 women screened (risk difference [RD], +9.3 per 1,000; 95% CI, 4.0-14.6; P < 0.01; I2 = 0%). Three studies directly compared CEM with MRI in screening settings and were summarized qualitatively because of heterogeneity and limited cancer events. CONCLUSION: In screening and surveillance populations, contrast-enhanced mammography increases cancer detection compared with low-energy mammography. Evidence from available comparative studies suggests that cancer detection with CEM may be broadly comparable to MRI in selected screening contexts; however, the limited number of studies and heterogeneous designs preclude definitive comparative conclusions. These pathway-specific estimates may inform clinical implementation and the design of prospective screening trials.

30 July 2026

Read appraisal →
diagnosticEvidence: Insufficient
10CEBM

Analytical chemistry

Wearable and Multimodal Electrochemical Hydrogel Sensor for Real-Time Non-Invasive Sweat Glucose Monitoring

Noninvasive sweat glucose monitoring is a promising strategy for real-time health management. In this study, we developed a flexible electrochemical sensor platform based on gold nanorods@polylysine-glucose oxidase (AuNRs@PLL-GOx) composite hydrogel, which enables noninvasive, highly sensitive, and multimode detection of sweat glucose. The polylysine (PLL) interfacial layer provides abundant amino groups for glucose oxidase (GOx) immobilization, improves the dispersion of gold nanorods (AuNRs) within the hydrogel matrix, and facilitates interfacial charge transport by maintaining close contact between GOx and the conductive AuNRs network. These effects improve the electron-transfer efficiency and analytical performance of the hydrogel sensor. Furthermore, the platform integrates differential pulse voltammetry (DPV), cyclic voltammetry (CV), and chronoamperometry (i-t) within a single hydrogel system for sweat glucose monitoring. Benefiting from synergistic and multitechnique detection, the sensor demonstrated excellent analytical performance, including wide detection range (up to 160 μM), low detection limit (3.71 μM), and strong anti-interference capability. These results indicate that this hydrogel-based platform is well suited for future wearable biosensors and smart healthcare systems.

30 July 2026

Read appraisal →
otherEvidence: Weak
50CEBM

Proceedings of the National Academy of Sciences of the United States of America

The backfiring effect of weak AI safety regulation

Recent policy proposals aim to improve the safety of general-purpose AI, but there is little understanding of the efficacy of different regulatory approaches. We present a strategic model that explores interactions between safety regulation, general-purpose AI technology creators, and domain specialists-those who adapt the technology for specific applications. Our analysis examines how regulatory measures targeting different parts of the AI development chain affect the outcome of this game. Our model assumes AI technology is characterized by two key attributes: safety and performance. The regulator first sets a minimum safety requirement that applies to one or both players. The general-purpose creator then invests in the technology, establishing its initial safety and performance levels. Next, domain specialists refine the AI for their use cases, updating the safety and performance levels and taking the product to market. Resulting revenue is shared between the specialist and generalist. Our analysis reveals two insights: first, weak safety regulation imposed predominantly on domain specialists can backfire. While it might seem logical to regulate AI use cases, our analysis shows that weak regulations targeting domain specialists alone can reduce safety in a large class of parameterizations. Second, in contrast to the previous finding, we observe that stronger, well-placed regulation can mutually benefit all players. When regulators impose appropriate safety standards on both general-purpose AI creators and domain specialists, the regulation can function as a commitment device, leading to safety and performance gains, surpassing what is achievable under no regulation or regulating only one player.

28 July 2026

Read appraisal →
Systematic ReviewEvidence: Weak
50CEBM

PloS one

The economic burden of Type 2 Diabetes by social determinants of health: A systematic review.

BACKGROUND: The unequal distribution of resources in society generates social gradients that translate into health inequalities and differential use of health care resources and their costs. Non-medical factors such as employment, income, ethnicity and education impact the prevalence and treatment outcomes of patients with type 2 diabetes mellitus (T2DM); however, there is a scarcity of articles assessing the relationship between health inequalities and the economic costs of treatment. Therefore, we conducted a systematic review of published studies examining the cost differences of treating T2DM across social determinants of health (SDH). METHODS: We systematically searched MEDLINE, Embase, PsycINFO, EconLit, and NHS EED for original peer-reviewed articles that provided cost differences of treating T2DM by SDH: education, income, employment, residency and ethnicity. We grouped the studies by each SDH and calculated the percentage differences where possible between the lowest and highest ends of the gradient (education, income and employment). Residency was categorised as rural vs. urban and ethnicity as white or general population vs other ethnic minorities. RESULTS: We included 19 articles retrieved internationally from varying healthcare systems. Results were contextualised given the healthcare financing model. In countries with high out-of-pocket expenses, Black and Hispanic ethnic backgrounds and rural residence were associated with lower direct health care and costs likely to be determined by ability to pay rather than clinical need. Indirect costs such as lost productivity due to absenteeism were also lower in unemployed, and lower income groups. CONCLUSIONS: There are evident health disparities in the direct and indirect economic consequences of T2DM. The effect of decreased healthcare use and costs on treatment outcomes needs to be further explored to inform policies to ensure healthcare delivery is based on clinical need rather than socio-economic factors.

28 July 2026

Read appraisal →
observationalEvidence: Weak
20CEBM

PloS one

Artificial intelligence for early detection of diabetic retinopathy: A vision transformer-based approach

BACKGROUND: Early identification of diabetic retinopathy (DR), which is a primary cause of vision impairment globally, is a crucial phasis for effective intervention and treatment. Traditional screening workflows rely on manual diagnosis by ophthalmologists, which remains the gold standard but can be time-consuming and subject to variability due to human factors. To support and enhance the screening process, artificial intelligence (AI)-based tools have shown promise in automating DR detection, particularly with recent advances in deep learning. However, medical images with long-range dependencies and spatial linkages can be challenging for CNN-based algorithms to handle. METHODS: This paper proposes a Vision Transformer (ViT)-based model, specifically using a Compact Convolutional Transformer (CCT), for early automated detection of DR. The model uses self-attention techniques to improve feature extraction and classification performance; combining three main stages: the CCT tokenizer, transformer encoder, and sequence pooling. The proposed approach was trained on public datasets (EyePACS and APTOS 2019) and evaluated against state-of-the-art deep learning architectures. RESULTS: Our experimental findings demonstrate that ViT performs among the best in the current state of the art with an overall accuracy of 97% and F1-scores above 0.95 across all DR severity levels. Our system is primarily designed for the pre-screening stage of diabetic retinopathy workflows, enabling rapid and reliable identification of potential DR cases for further clinical evaluation. CONCLUSION: These results highlight the potential of transformer-based designs in medical picture analysis, as well as the implications for telemedicine and e-health solutions in real-time, especially in cases of low-resource settings.

28 July 2026

Read appraisal →
otherEvidence: Insufficient
25CEBM

Journal of medical Internet research

Digital Health Technologies Are Bridging the Maternal Mortality Gap

Disparities in maternal outcomes along socioeconomic, racial, ethnic, and geographic lines are well established. In this News and Perspectives article, JMIR Correspondent Anika Nayak reports on the growth and potential of digital health interventions in bridging the rural-urban maternal mortality gap.

26 July 2026

Read appraisal →
otherEvidence: Weak
45CEBM

Medicine

A bibliometric analysis of global trends in AI-driven digital health technologies for diabetes management

BACKGROUND: Digital health technologies are increasingly applied in diabetes care, enabling continuous monitoring, personalized support and remote interventions. Meanwhile, artificial intelligence (AI) is enhancing the precision and effectiveness of these tools. This study aims to map global research trends and thematic developments in AI-driven digital health technologies for diabetes management and to explore their future directions. METHODS: We collected data from the Web of Science Core Collection, including articles and reviews published up to July 12, 2025, using CiteSpace, VOSviewer, and Microsoft Excel to analyze countries/regions, institutions, journals, references, authors, and keywords. RESULTS: A total of 673 publications were included in the analysis. Global publications on AI-driven digital health technologies for diabetes increased steadily, with the USA leading in output. The University of London ranked as the most productive institution. Sensors and diabetes care were the most frequently published and cited journals in this field. Herrero P was among the most prolific authors. The most cited article was "Development and Validation of a Deep Learning Algorithm for Detection of Diabetic Retinopathy in Retinal Fundus Photographs." "diabetes" was the most frequently occurring keyword. Keyword cluster analysis identified 3 primary research hotspots: AI-enabled monitoring, digital health interventions, and AI-based diabetic retinopathy screening. CONCLUSIONS: This study summarizes the evolution of AI-driven digital health technologies in diabetes care. Although challenges remain in data security, standardization and validation, these technologies hold increasing potential for accurate diagnosis, real-time monitoring and personalized care.

26 July 2026

Read appraisal →
Systematic ReviewEvidence: Moderate
80CEBM

Heart (British Cardiac Society)

Impact of diabetes on outcomes in hypertrophic cardiomyopathy: a GRADE meta-analysis

BACKGROUND: Diabetes mellitus (DM) is a common comorbidity in hypertrophic cardiomyopathy (HCM) and may exacerbate arrhythmic risk, promote structural remodelling and worsen heart failure outcomes. Its overall prognostic impact and effect on cardiac structure and function in adults with HCM remain uncertain. METHOD: We systematically searched PubMed, Scopus, Web of Science and Cochrane to January 2025 for observational studies comparing adults with HCM-DM versus HCM without DM. Random-effects meta-analyses were performed to pool ORs for clinical outcomes and standardised mean differences (SMDs) for echocardiographic parameters. Certainty of evidence was assessed using the Grading of Recommendations Assessment, Development and Evaluation (GRADE) framework after Risk Of Bias In Non-randomised Studies - of Interventions (ROBINS-I) evaluation. Subgroup, sensitivity and heterogeneity analyses were undertaken. RESULT: Eight studies encompassing approximately 47 592 patients met inclusion criteria. DM was associated with higher odds of all-cause mortality (OR 1.43, 95% CI 1.29 to 1.58; high certainty), heart failure (OR 1.34, 95% CI 1.25 to 1.43; moderate certainty) and atrial fibrillation (OR 1.41, 95% CI 1.18 to 1.68; high certainty). The association with atrial fibrillation was most pronounced in patients younger than 50 years (OR 2.55) and attenuated in those with body mass index ≥30 kg/m². HCM-DM was also linked to smaller left ventricular end-diastolic volumes (SMD -0.26) and impaired global longitudinal strain (SMD 0.58), consistent with subclinical systolic dysfunction, although heterogeneity was high and certainty low to moderate. Evidence for left ventricular ejection fraction, mass and septal thickness was inconclusive. Results were robust across sensitivity analyses. CONCLUSIONS: DM is a clinically important risk marker in HCM, associated with excess mortality, heart failure and atrial fibrillation, as well as adverse structural-functional changes. These findings support closer rhythm and function monitoring in HCM-DM and highlight the need for prospective studies to determine whether targeted metabolic interventions can improve outcomes. PROSPERO REGISTRATION NUMBER: CRD420250650799.

26 July 2026

Read appraisal →
otherEvidence: Insufficient
20CEBM

Journal of medical Internet research

The Next Generation of Wearables Won't Need the Cloud

Wearable health and fitness devices have typically relied on cloud computing to deliver insights. In this News and Perspectives article, JMIR Correspondent Michelle Falci reports on the advances facilitating on-device data processing and the potential of these next generation wearables.

25 July 2026

Read appraisal →
otherEvidence: Moderate
60CEBM

Journal of medical ethics

AI interventions in cancer screening: balancing equity and cost-effectiveness

This paper examines the integration of artificial intelligence (AI) into cancer screening programmes, focusing on the associated equity challenges and resource allocation implications. While AI technologies promise significant benefits-such as improved diagnostic accuracy, shorter waiting times, reduced reliance on radiographers, and overall productivity gains and cost-effectiveness-current interventions disproportionately favour those already engaged in screening. This neglect of non-attenders, who face the worst cancer outcomes, exacerbates existing health disparities and undermines the core objectives of screening programmes.Using breast cancer screening as a case study, we argue that AI interventions must not only improve health outcomes and demonstrate cost-effectiveness but also address inequities by prioritising non-attenders. To this end, we advocate for the design and implementation of cost-saving AI interventions. Such interventions could enable reinvestment into strategies specifically aimed at increasing engagement among non-attenders, thereby reducing disparities in cancer outcomes. Decision modelling is presented as a practical method to identify and evaluate these cost-saving interventions. Furthermore, the paper calls for greater transparency in decision-making, urging policymakers to explicitly account for the equity implications and opportunity costs associated with AI investments. Only then will they be able to balance the promise of technological innovation with the ethical imperative to improve health outcomes for all, particularly underserved populations. Methods such as distributional cost-effectiveness analysis are recommended to quantify and address disparities, ensuring more equitable healthcare delivery.

25 July 2026

Read appraisal →
observationalEvidence: Moderate
60CEBM

JMIR human factors

Japanese Health Information Technology Usability Evaluation Scale for Sexually Transmitted Infection-Related Chatbots: Development and Psychometric Validation Study

BACKGROUND: The rapid expansion of mobile technology has accelerated the integration of health applications and conversational AI into clinical and public health practices. To ensure these tools are effective and sustainable, usability evaluations and early user engagement during development are essential. The Health Information Technology Usability Evaluation Scale (Health-ITUES) is a validated and flexible usability assessment instrument that is available in multiple languages and applicable across diverse contexts. However, a Japanese version of this scale has not yet been developed. OBJECTIVE: This study aimed to translate and validate a Japanese version of the Health-ITUES, customized for a sexually transmitted infection (STI)-related chatbot, and to support the usability assessment of emerging mobile health tools in Japan. METHODS: We developed a Japanese version of the Health-ITUES using a chatbot under development as a consultation tool for young women regarding STIs. First, the original scale was customized to reflect the chatbot's specific purpose and intended usage context. Following established translation guidelines, we conducted forward translation from English to Japanese, back translation, expert review, and reconciliation. We then evaluated the reliability and validity of the Japanese version in a sample of 301 young women. RESULTS: The Japanese version of the Health-ITUES demonstrated high internal consistency (Cronbach α=0.85-0.98). Confirmatory factor analysis supported acceptable construct validity (root mean square error of approximation is 0.10, comparative fit index>0.90). Additionally, the Health-ITUES scores showed strong correlations with satisfaction and usage intention for the tool (r=0.779 and 0.797, respectively). CONCLUSIONS: The Japanese version of the Health-ITUES provides initial evidence of reliability and validity in an STI-related scenario among young women and may facilitate more rigorous usability evaluations of mHealth and conversational AI tools in Japan.

24 July 2026

Read appraisal →
observationalEvidence: Weak
30CEBM

Journal of medical Internet research

Initial-Visit Specialty Triage in Rare Diseases Using Large Language Models: Retrospective Benchmarking Study

BACKGROUND: Specialty triage at first contact is an overlooked step in early diagnostic pathways for rare diseases. Patients often present with overlapping, multisystem, and atypical manifestations, making first-visit specialty selection challenging and potentially prolonging diagnostic pathways. OBJECTIVE: The aim of this study is to evaluate the accuracy, response time, and consistency of large language models (LLMs) for initial-visit specialty triage in rare diseases across multiple datasets, and to compare their performance with registered nurses and nonmedical participants. METHODS: In this retrospective benchmarking study, we used 5 rare disease datasets: a publication-derived case set, 3 RareBench-derived datasets, and a Facial phenotype-Gene-Disease Dataset-derived set. Fourteen LLMs were evaluated over 5 independent runs per case. Performance was assessed using accuracy, response time, and consistency, with subgroup analyses by model accessibility, reasoning mode, parameter scale, and phenotype count. Human comparison was conducted on the publication-derived case set using registered nurses and nonmedical participants. RESULTS: Across datasets, model accuracy ranged from 0.4378 to 0.7141. Claude-opus-4-5 achieved the highest accuracy (0.7141) and consistency (0.9653), averaging 10.79 seconds per case. GPT-5.1 had the shortest response time (3.39 s/case) and high accuracy (0.6948). Proprietary models had numerically higher average accuracy than open-weight models (0.6973 vs 0.6365). Nonthinking models achieved higher average accuracy than thinking models (0.6789 vs 0.5826) and had shorter response times, although this exploratory comparison was based on a small number of thinking models. Accuracy varied by phenotype count, with higher performance in cases with 1 to 2 or more than 14 phenotypes. On the publication-derived case set, LLMs achieved higher average accuracy than registered nurses and nonmedical participants (0.5978 vs 0.4914 and 0.4573). CONCLUSIONS: LLMs showed potential as assistive tools for initial-visit specialty triage in rare diseases. Model choice, reasoning mode, and phenotype information density influenced performance, but subgroup findings should be interpreted cautiously. Future work should evaluate LLM-based specialty triage in prospective clinical settings and develop clinician-supervised workflows with traceable evidence support.

24 July 2026

Read appraisal →
otherEvidence: Weak
50CEBM

Journal of medical Internet research

Model and Task-Aware Test-Time Scaling Strategies for Large Language and Vision-Language Models in Medicine: Evaluation Study

BACKGROUND: Test-time scaling has emerged as a promising method to enhance the reasoning capabilities of large language models (LLMs) and vision-language models (VLMs) during inference without additional training. While foundational studies established scaling paradigms in general domains, their applicability to the unique complexities of medical AI remains underexplored. OBJECTIVE: This study aims to conduct a comprehensive investigation of test-time scaling in the medical domain. We evaluate the impact of scaling across different model sizes and task complexities. Furthermore, we seek to identify domain-specific bottlenecks and assess model robustness against user-driven perturbations, such as misleading clinical authority. METHODS: This study evaluated a diverse set of general and medical-specific LLMs and VLMs. Experiments used five textual medical benchmarks comprising over 5500 questions and two multimodal benchmarks comprising 7000 samples. Performance was measured under three scaling conditions: increasing token budgets, iterative sequential scaling, and parallel scaling. Robustness was tested by embedding misleading hints with varying tones and levels of simulated clinical expertise into prompts. RESULTS: For nonreasoning LLMs, accuracy saturated quickly, with token usage often remaining under 500 tokens regardless of budget increases. Reasoning models demonstrated significant performance gains on complex tasks as token budgets increased. Notably, we identified distinct domain-specific behaviors. First, current VLMs showed a structural bottleneck in integrating visual clues and experienced limited benefit from token expansion. Second, medically fine-tuned LLMs excelled in clinical question answering but exhibited degraded scaling efficiency on calculation tasks compared to general-domain models. This reflects a disparity between qualitative clinical alignment and procedural logic. Third, while optimal scaling improved robustness, models exhibited a cognitive vulnerability by readily abandoning correct reasoning when confronted with misleading expert physician hints. Regarding scaling strategies, parallel scaling outperformed sequential scaling on easier tasks. Conversely, extended sequential scaling or increased budgets proved essential for complex problem-solving. CONCLUSIONS: Test-time scaling rules from general domains do not perfectly translate to medical AI. Longer reasoning is not universally beneficial. Concise reasoning with parallel scaling is optimal for simpler tasks. An extended chain of thought via sequential scaling or increased budgets is required for complex problems. Furthermore, safe clinical deployment requires addressing fundamental vision-language alignment, balancing clinical and procedural reasoning, and mitigating vulnerabilities to perceived clinical authority.

24 July 2026

Read appraisal →
otherEvidence: Weak
45CEBM

Journal of medical Internet research

Digital Peer Support Intervention for Family Caregivers of Individuals With Neuromuscular Disease: Randomized Controlled Trial

BACKGROUND: Neuromuscular diseases (NMD) affect nerves and muscles, resulting in weakness and often profound disability. Family caregivers of individuals with NMD experience significant psychological burden, stress, and reduced well-being. Digital peer support interventions may help to ameliorate these negative impacts. OBJECTIVE: This study evaluated the effect of a 12-week digital peer support intervention compared to usual care on caregiver mastery, competence, stress, burden, anxiety, and depression among family caregivers of individuals with NMD. METHODS: We conducted a parallel-group multicenter randomized controlled superiority trial in Ontario, Canada. Family caregivers of children or adults with NMD were recruited between August 2022 and September 2023 through 7 sites, social media, and national organizations. Participants were randomized 1:1 to a 12-week digital peer support intervention or usual care. The 12-week intervention comprised access to a trained peer mentor, private app-based communication via aTouchAway (Aetonix, Canada), and weekly moderated digital group discussion forums. The primary outcome was caregiver mastery measured using the Pearlin Mastery Scale, adjusted for baseline score. Secondary outcomes included caregiver stress, competence, burden, anxiety, and depression. We calculated adjusted (for baseline score) mean differences using analysis of covariance and generated multivariable linear regression models exploring associations with the intervention and caregiver age, years of caregiving, care recipient medical diagnosis, care recipient ventilation type, adjusting for baseline outcome scores. Intervention fidelity was evaluated through participant engagement metrics. RESULTS: A total of 100 participants were randomized (n=50 intervention and n=50 control). Participants had a mean age of 46.8 (SD 11.6) years, 70% (n=70) were mothers, with a mean length of caregiving of 11.8 (SD 7.6) years. We found no difference in 12-week Pearlin Mastery Scale scores (adjusted mean difference 0.67, 95% CI -1.7 to 3.1). We also found no difference in any of our secondary outcomes. Mentors and participants sent a mean of 21.3 (SD 33.3) and 17.7 (SD 33.0) messages, respectively. Overall, 62% (n=31) of participants and 92% (n=11) of mentors engaged in at least 1 program element for ≥8 of the 12 weeks. CONCLUSIONS: Our 12-week digital peer support program had no effect on caregiver mastery or other caregiving or psychological outcomes among family caregivers of individuals with NMD. This might be partly due to moderate fidelity and variability in participant engagement. Unlike prior caregiver interventions that incorporated structured psychoeducation or self-management training, this intervention evaluated primarily peer support delivered through a digital platform. This study contributes important evidence regarding the feasibility and limitations of scalable digital peer support programs for caregivers of individuals with NMD. These findings highlight the importance of intervention tailoring, participant matching, and sustained engagement. Future research should evaluate longer-duration and more individualized peer support models targeting caregivers earlier in the caregiving trajectory to improve intervention fidelity and ultimately caregiver well-being.

24 July 2026

Read appraisal →
Systematic ReviewEvidence: Moderate
45CEBM

Journal of medical Internet research

Evaluation Frameworks for Clinical AI Incorporating Validation Strategies, Real-World Applicability, and Ethical Principles: Scoping Review

BACKGROUND: AI shows substantial potential in health care; however, the absence of standardized evaluation frameworks limits its safe and effective clinical implementation because of inconsistent validation requirements and fragmented ethical principles. Existing guidelines vary in structure, methodological rigor, and ethical integration, creating uncertainty. OBJECTIVE: This study aimed to systematically map, characterize, and critically analyze existing evaluation frameworks for clinical AI, focusing on three core dimensions: methodological rigor, validation strategies (internal validation, including reporting of technical and clinical performance; external validation, including real-world applicability), and alignment with the United Nations Educational, Scientific and Cultural Organization (UNESCO) AI ethical considerations. METHODS: A scoping review was conducted following PRISMA-ScR (Preferred Reporting Items for Systematic Reviews and Meta-Analyses extension for Scoping Reviews) guidelines. Six databases (PubMed, Embase, BVS, EBSCOhost, ProQuest, and Sage) and the Enhancing the Quality and Transparency of Health Research Network were searched without language or date restrictions up to February 2026. Eligible documents included peer-reviewed papers, gray literature, and organizational guidelines describing evaluation or reporting frameworks for clinical AI. Editorials, commentaries, and conference abstracts lacking a clearly defined evaluative framework or clinical applicability were excluded. Two reviewers independently screened records and extracted data. Data were extracted across three domains: (1) general characteristics, (2) methodological rigor and validation parameters, and (3) ethical integration and were synthesized using a dot plot-based gap map. Ethical adherence was assessed using a 10-domain UNESCO-based scoring matrix. No formal risk-of-bias assessment was conducted, consistent with scoping review methodology. RESULTS: From 3363 records, 46 frameworks met the inclusion criteria. Mapping revealed a rapidly expanding but fragmented landscape. Most frameworks targeted investigational use (88%), with limited focus on clinical applicability. Frameworks varied in structure, methodology, and scope, with a predominance of reporting guidelines and few validated tools. Most (63%) were developed through multi-institutional collaborations, and 32.6% incorporated transdisciplinary participation. Only 31.8% reported technical metrics (commonly area under the curve, sensitivity, and specificity), and 15.9% provided clinical indicators (eg, predictive values or calibration). Only 11.4% achieved methodological rigor, incorporating validation aligned with intended use, while most relied on partial validation strategies, highlighting a gap between model development and clinical evaluation. Ethical integration was heterogeneous: only 5 frameworks achieved high compliance (≥80%), whereas 4 scored <10%. The most frequently addressed UNESCO principles were awareness and education (71.1%) and transparency and explainability (70%), while human oversight (24.4%) and adaptive governance (33.3%) were least represented. Findings indicate a misalignment between framework design, validation requirements, and clinical implementation. CONCLUSIONS: Evaluation frameworks for clinical AI remain heterogeneous and oriented toward investigational contexts. Critical gaps persist in methodological rigor, validation aligned with intended use, and fragmented ethical coverage. These findings highlight the need for standardized, robust, and ethically grounded frameworks to enable safe, reliable, and scalable integration of AI into clinical practice.

24 July 2026

Read appraisal →
otherEvidence: Weak
65CEBM

Supportive care in cancer : official journal of the Multinational Association of Supportive Care in Cancer

Initial development and usability testing of a digital fertility preservation navigator for newly diagnosed young adult cancer survivors

PURPOSE: Young adults with cancer (YAs) are rarely provided the necessary resources to make informed decisions about fertility preservation (FP) prior to beginning treatment. Digital decision aids are a promising strategy to ensure YAs receive timely, guideline-concordant information. We developed an alpha version Fertilit-e, a digital FP navigator, and conducted usability testing to obtain feedback from YAs. METHODS: We designed an alpha version of Fertilit-e that drew from existing evidence-based FP resources. Feedback was obtained from two waves of participants using a think-aloud framework, and modifications were made to the decision aid. RESULTS: Across waves, participants (n = 12) understood the purpose of Fertilit-e and thought modules were comprehensive and easy to understand. The usability of Fertilit-e was rated as excellent across waves and devices (iPad and computer). Participants suggested reducing the use of and better defining medical terminology and adding text to videos to enhance comprehensibility. CONCLUSIONS: YAs reported that Fertilit-e provided important information regarding FP in an understandable manner. Participants made suggestions for how to improve comprehensibility and usability that will guide future iterations of Fertilit-e. For example, tailoring algorithms and reorganization may direct users to the most relevant information. IMPLICATIONS FOR CANCER SURVIVORS: Digital decision aids may be a feasible and useful strategy for helping YAs navigate decisions regarding whether to pursue FP at the time of diagnosis. Additional research is needed to refine and evaluate such tools before integration into routine care.

24 July 2026

Read appraisal →
qualitativeEvidence: Moderate
70CEBM

PloS one

"We'll have to see how it works": An interview study to understand collaborative practices in interdisciplinary artificial intelligence and healthcare research.

Developing artificial intelligence (AI) algorithms for healthcare is a collaborative effort, bringing data scientists, clinicians, patients and other stakeholders together. By understanding AI as 'sociotechnical' where the social and the technical nature of the work and the models are inseparable, we explore the AI development workflow and how stakeholders navigate the challenges and tensions of sharing and generating knowledge across disciplines. We conducted an inductive thematic analysis of 13 semi-structured interviews with participants in early stages of AI-in-healthcare research consortia in the UK. Our findings identify that participants needed to adapt both the tools used for sharing and the information communicated according to their audience, particularly when working with those with a clinical or patient perspective. We identify the novelty of participating in AI research, how AI knowledge is shared, and the inclusion of clinician and patient stakeholder perspectives as key areas within collaborative AI practices in healthcare. These findings highlight that bringing AI into the mix can introduce new obstacles to interdisciplinary work.

24 July 2026

Read appraisal →
observationalEvidence: Weak
45CEBM

Cell reports. Medicine

Toward generalizable prediction of cancer signal using a cell-free DNA language model.

Cell-free DNA can be used for early cancer detection, minimal residual disease monitoring, and post-treatment risk stratification. However, current assays are often designed for a single purpose and rely on deep or broad sequencing panels that capture only a small fraction of tumor-derived signals, limiting transferability, increasing cost, and reducing scalability. Fragmentia-AI is an artificial intelligence language model that learns fragment-level sequence patterns in tumor-derived cell-free DNA. Instead of focusing on mutations, it uses the structure of cell-free DNA to detect cancer signals in a partially panel-agnostic manner from ultra-low sequencing input, approximately 0.1%-1% of conventional depth. The model performs well across cancer types and clinical settings, including monitoring after surgery or immunotherapy, and in samples with low variant allele frequencies or no detected mutations. Fragment-level analyses identify shorter fragments and tumor-derived sequence patterns across panels of different sizes and ultra-low-pass whole-genome sequencing in multiple cohorts.

23 July 2026

Read appraisal →
otherEvidence: Weak
55CEBM

Current heart failure reports

The Future of Imaging in Heart Failure: Toward Precision Phenotyping, Integration, and Intelligence.

PURPOSE OF REVIEW: Heart failure (HF) is increasingly understood not as a single, uniformly treated diagnosis but as a heterogeneous syndrome requiring aetiological clarification, in which cardiac imaging is central. As the opening article of this journal's 'Imaging in Heart Failure' section, this review surveys the technologies currently reshaping HF imaging and sets out the section's scope and priorities, framing the shift from a descriptive, modality-siloed practice toward an integrated, predictive, patient-specific discipline. RECENT FINDINGS: Artificial intelligence (AI) now delivers expert-level echocardiography automation, guides image acquisition by novices in resource-limited settings, detects aetiologies such as transthyretin amyloid cardiomyopathy from a single acquisition and enables deep phenotyping through radiomics and vendor-agnostic strain analysis. Handheld, AI-enabled point-of-care ultrasound extends imaging-guided triage beyond the echocardiography laboratory. Cardiovascular magnetic resonance (CMR) advances - parametric mapping, four-dimensional flow, diffusion tensor imaging, spectroscopy, and accelerated reconstruction - broaden tissue and metabolic characterisation, including patients with implanted devices. Molecular imaging with novel positron emission tomography tracers and hyperpolarised magnetic resonance is moving from depicting the structural consequences of disease to imaging active pathobiology, while photon-counting computed tomography and image-derived digital twins support one-stop structural assessment and in-silico prediction of therapy response. The convergence of AI, molecular imaging and advanced precision is transforming HF imaging from better pictures into smarter, integrated, personalised data that directly inform care. Realising this promise will require rigorous validation, attention to algorithmic bias and generalisability, demonstrated cost-effectiveness, curricular reform, and equitable access. This section aims to critically appraise these innovations and their translation into practice.

23 July 2026

Read appraisal →
otherEvidence: Moderate
60CEBM

BMJ health & care informatics

Benchmarking large language models for de-identification of electronic health record notes

OBJECTIVES: The rapid evolution of large language models (LLMs) and their growing application in clinical text processing have created an urgent need for reliable de-identification mechanisms. While LLMs show promise in identifying sensitive health information (SHI), their capabilities require rigorous evaluation. This study aims to conduct a comprehensive benchmarking analysis of various LLM-based, traditional rule-based and hybrid de-identification methods. METHODS: Our benchmark analysis used five datasets (i2b2-2006, MIMIC-2008, i2b2-2014, i2b2-2016 and OpenDeID v1) from different countries. We developed three baseline and eight LLM-based models. The experimental setup encompassed nine different settings using various combinations of training and testing sets to assess model robustness and cross-dataset performance. RESULTS: In the baseline models, the approach trained on the combined corpus of all five datasets (setting 3) significantly outperformed the other settings, achieving a strict F1 micro-average score of 0.8172. Regarding LLM-based models, the supervised fine-tuning approach using the same combined configuration (setting 9) achieved the highest performance with a strict F1 score of 0.9447. DISCUSSION: The harmonisation of corpora ensured standardised data formatting and SHI management across five diverse datasets, highlighting the necessity for uniform categorisation to enhance the reliability of de-identification results. CONCLUSIONS: Our findings indicate that while fine-tuned LLMs offer superior accuracy, the observed performance variability across heterogeneous electronic health record sources poses significant technical challenges. Real-world implementation must address these inconsistencies to overcome the ethical and technical hurdles associated with deploying LLMs for handling sensitive health data.

23 July 2026

Read appraisal →
Systematic ReviewEvidence: Weak
35CEBM

Orphanet journal of rare diseases

Clinical outcomes in alpha-mannosidosis: a systematic review of therapeutic approaches

BACKGROUND: Alpha mannosidosis (AM) is a rare lysosomal storage disorder caused by a deficiency in the α-mannosidase enzyme, resulting in impaired glycoprotein metabolism within lysosomes. Enzyme dysfunction is attributed to an autosomal recessive mutation in the MAN2B1 gene. Affected individuals present with a broad spectrum of manifestations, including developmental delays, cognitive decline, musculoskeletal abnormalities, hearing difficulties, and recurrent infections. Current therapeutic options are limited to hematopoietic stem cell transplantation and the more recently developed enzyme replacement therapy. OBJECTIVE: The aim of this review was to evaluate and compare the therapeutic outcomes, benefits and challenges associated with Hematopoietic stem cell transplantation (HSCT) and enzyme replacement therapy (ERT) in the treatment of AM. METHODS: A systematic search across PubMed, MEDLINE, EMBASE, the Cochrane Library, OMIM, and ScienceDirect identified 12 original studies from 307 records. The data are presented narratively due to the scarcity of literature and the heterogeneity of study designs and interventions. RESULTS: A total of 28 patients who received hematopoietic stem cell transplantation showed improvements in preserving neurocognitive function and skeletal stabilization and reduced infection rates, especially when performed at relatively young ages. However, this treatment carries significant risks, including infections, graft-versus-host disease, and increased morbidity and mortality, particularly in older patients. Conversely, enzyme replacement therapy was administered to 75 patients, who demonstrated a favorable safety profile, enhanced respiratory function, reduced skeletal abnormalities, and improved overall quality of life. However, enzyme replacement therapy has limited efficacy in preventing neurocognitive decline and requires lifelong administration. CONCLUSION: Both interventions yield better outcomes when initiated early, particularly before cognitive deterioration becomes significant. This review emphasizes the importance of a timely diagnosis to optimize treatment outcomes and prevent severe complications.

23 July 2026

Read appraisal →
observationalEvidence: Weak
50CEBM

JMIR cardio

Mobile Apps for Heart Rate Variability: App Store Search and Content Analysis

BACKGROUND: Heart rate variability (HRV) is a noninvasive indicator of autonomic nervous system activity that is increasingly used for health and performance monitoring. Digital and mobile technologies are increasingly providing opportunities for remote HRV monitoring outside of laboratory-based settings. OBJECTIVE: This study aimed to describe the landscape of mobile apps that measure, analyze, and provide feedback on HRV, with a focus on how HRV is measured, analyzed, interpreted, and communicated to users. A secondary aim was to assess the transparency of these apps, including the extent to which they disclose the evidence underpinning their HRV metrics and feedback. METHODS: This study was an app store search and content analysis. Searches were conducted in the Google Play Store and Apple iTunes Store. Apps were eligible for inclusion if they had functionality to record, analyze, or provide feedback on HRV and were available in English. Data were extracted from app descriptions, screenshots, websites, and, where necessary, contact with developers. Data were extracted on app metadata (developer, release and update dates, and pricing), alongside information about HRV measurement, analysis, and feedback. This included the type of sensor used; HRV measurement characteristics (sensor placement, recording duration, and body position or standardization procedures); methods to calculate and interpret HRV (ie, metrics derived and how they were interpreted for users); and additional app functionality such as reminders, the ability to log self-reported stressors, and the type of feedback or guidance provided based on HRV. We used previously published criteria for assessing the quality of information on the internet, which included authorship, scientific attribution, currency of updates, and data privacy. RESULTS: Of 746 apps identified, 206 met eligibility criteria. Of these, 132 were primary measurement apps, 59 were aggregators, and 15 were hybrid. Photoplethysmogram was the most common sensing modality (n=117, 56.8%), followed by multiple sensors (n=60, 29.1%). Full data extraction across app metadata and HRV measurement and analysis data was only achievable for 93 (45.1%) apps, representing a transparent subset with sufficient available information for content analysis. The most commonly reported HRV metrics were root mean square of successive differences (n=51) and SD of normal-to-normal intervals (n=48), while frequency-domain power (n=22) and low frequency to high frequency ratios (n=15) were less common. Most apps presented data as personalized trends or individualized ranges (76/93, 81.7%), emphasizing user-specific context rather than isolated values. Although 86% (80/93) offered contextual guidance (eg, readiness or recovery scores), many relied on proprietary algorithms that were not transparently described, limiting independent assessment of how these scores were derived and validated. CONCLUSIONS: Consumer HRV apps are widely available but vary considerably in how data are collected, processed, and contextualized. While many offer personalized trends and guidance, methodological transparency is often limited, particularly regarding the proprietary algorithms underlying the feedback scores.

20 July 2026

Read appraisal →
Randomised Controlled TrialEvidence: Moderate
65CEBM

Journal of global health

Animal-assisted therapy on psychological and physical outcomes: a meta-analysis of randomised controlled trials

BACKGROUND: Adults are at heightened risk of anxiety, stress, and depression; animal-assisted therapy (AAT) may serve as an effective approach to promote psychological well-being. In this study, we compared the effectiveness of AAT in improving depression, anxiety, stress, pain, and gait among adults with or without illness. METHODS: We systematically searched six electronic databases (PubMed, CINAHL, Embase, Web of Science, the Cochrane Library, and Scopus) and included all studies published up to August 2024. We used comprehensive meta-analysis software to complete the quantitative synthesis. We reported pooled effect sizes as Hedges' g with corresponding 95% confidence intervals (CIs), after applying a random-effects model. Furthermore, we assessed heterogeneity using Cochran's Q test and the I2 statistic. We applied the Cochrane Risk of Bias 2.0 tool to appraise the methodological quality of the included studies. The synthesis process followed PRISMA guidelines. RESULTS: From the 13,345 studies identified, 35 randomised controlled trials involving 2391 adults were included. Across diverse populations, AAT was associated with reductions in depression (Hedges' g = -0.403; 95% CI = -0.536, -0.271), anxiety (Hedges' g = -0.661; 95% CI = -1.069, -0.253), and stress (Hedges' g = -1.062; 95% CI = -1.849, -0.275) at post-intervention, although substantial between-study variability was observed. CONCLUSIONS: We demonstrated that AAT significantly improves anxiety, depression and stress in adults, but has no meaningful effect on pain or gait. Subgroup and meta-regression analyses indicate that psychological benefits depend on population and intervention characteristics, and are not moderated by age or gender. REGISTRATION: PROSPERO: CRD42024570108.

19 July 2026

Read appraisal →
Systematic ReviewEvidence: Moderate
75CEBM

International journal of technology assessment in health care

Consideration of intersectoral costs and benefits in health economic evaluations of pneumococcal conjugate vaccines: a systematic review.

OBJECTIVES: This systematic review explores which intersectoral costs and benefits (ICBs) are considered in the economic evaluation of pneumococcal conjugate vaccines (PCVs) across different age groups and the related impacts on the results of economic evaluations. METHODS: Seven databases were searched from 2009 to 2024 for full economic evaluations (Ees) of PCVs from a societal perspective. ICBs were presented narratively and in tabular form. Studies were appraised on quality, and PRISMA guidelines were followed. RESULTS: In all, 57 studies were included. ICBs were captured in the following sectors: patient and family, paid labor (productivity), nonpaid productivity, other sectors (education, leisure, consumption), unspecified sectors for ecological effects (herd immunity, serotype replacement, cross-protection), and equity. In cases where the comparator was no vaccination, 10 studies reported that PCVs were dominant when ICBs were included. After excluding dominant analyses, the median QALY-based incremental cost-effectiveness ratio (ICER) reduction ranged from -3.25 percent (for patient/family costs and productivity loss) to -46.02 percent (in children with ecological effects). Among DALY-based studies, the inclusion of ecological effects reduced the median ICER by -22.97 percent. In head-to-head comparisons, 16 studies reported that PCVs were dominant when ICBs were included. CONCLUSION: Inclusion of ICBs, except for serotype replacement, showed favorable cost-effectiveness results. This study provides a framework for identifying the intersectoral costs and benefits for the health economic evaluation of PCVs in children and adults.

19 July 2026

Read appraisal →
Systematic ReviewEvidence: Weak
35CEBM

JMIR mHealth and uHealth

mHealth Technologies for the Care of Children With Congenital Heart Disease: Scoping Review

BACKGROUND: Mobile information technology (IT) is increasingly being used in the health care sector, and it can play a critical role in both the care of children with congenital heart disease (CHD) and the quality of life of their families. OBJECTIVE: This study aimed to conduct a scoping review of the application of mobile health (mHealth) technologies in the care of children with CHD. We summarized the forms of mHealth interventions and effects on CHD to provide a reference for future research in this field. METHODS: We searched PubMed; Embase; Web of Science; the Cochrane Library; CINAHL; China National Knowledge Infrastructure; Wanfang Data; the Chinese Biomedical Database; VIP Chinese Science and Technology Journal Database; National Guideline Clearinghouse of the United States; the website of the Registered Nurses' Association of Ontario, Canada; the Guidelines International Network; the American Heart Association; and the American Association of Cardiovascular and Pulmonary Rehabilitation. The search period was from the establishment of the databases to June 12, 2025. The retrieved literature was screened and analyzed. RESULTS: A total of 519 Chinese- and English-language articles were identified, with 44 (8.5%) studies meeting the inclusion criteria. The primary forms of mHealth interventions for patients with CHD included mobile apps, wearable devices, and remote monitoring equipment. The findings indicated that mHealth technologies could improve exercise capacity, nutritional status, psychological well-being, and quality of life in children with CHD. CONCLUSIONS: The application of mHealth in the care of children with CHD is feasible and demonstrates positive effects. Future research should emphasize peer education and patient privacy protection while further exploring remote education and health management based on theoretical frameworks and intelligent ITs to enhance quality of life for both children with CHD and their parents.

19 July 2026

Read appraisal →
Systematic ReviewEvidence: Weak
15CEBM

Antonie van Leeuwenhoek

Advancements in the study of gut microbiome in disease diagnosis

This review summarizes disease-associated changes in gut microbial composition and evaluates the diagnostic performance of models constructed with different machine-learning algorithms. The review seeks to answer questions related to the relationship between the human gut microbiome and disease progression, how different machine learning algorithms affect disease diagnosis using gut microbiome data, and how disease-specific microbial communities impact diagnostic models. Multiple studies report that gut microbiome dysbiosis is commonly observed in many diseases, though patterns vary between conditions and cohorts. Large-scale computational analyses are increasingly applied to identify microbial signatures and to build diagnostic models; however, model performance often depends on data source, preprocessing and choice of algorithm. Overall, evidence indicates disease-associated shifts in gut microbial composition, and that diagnostic model accuracy is sensitive to cohort, sequencing and modeling choices. While certain taxa recur across studies for some diseases, heterogeneity between cohorts limits immediate clinical translation; thus, harmonized study designs and external validation are required. Future work should prioritize reproducible multi-cohort analyses, transparent reporting (e.g., PRISMA for reviews) and prospective validation before clinical deployment.

18 July 2026

Read appraisal →
Randomised Controlled TrialEvidence: Moderate
70CEBM

Journal of behavioral addictions

Efficacy of a mobile-based approach-avoidance task training (PROTECTapp) for problematic usage of the internet in young adults: A randomized controlled trial

BACKGROUND AND AIMS: Problematic usage of the internet (PUI) has been linked to impaired mental health and academic functioning in young adults. This randomized controlled trial evaluated the efficacy of a 3-week mobile-based approach-avoidance task (AAT) training (PROTECTapp) for reducing PUI in university students. METHODS: Ninety-two participants (Mage = 22.00 years, 69.6% women) with elevated levels of PUI were randomized to the PROTECTapp intervention (n = 45) or a waitlist control group (n = 47). Primary outcomes were PUI severity and internet-related craving. Secondary outcomes included motivation to change, psychopathological symptoms and academic functioning. Participants were assessed at baseline and postintervention; the intervention group completed additional 3- and 12-week follow-ups. RESULTS: Intention-to-treat analyses indicated greater reductions in PUI following the PROTECTapp intervention compared to the waitlist (p = .003; d = -0.80, 95% CI [-1.23, -0.38]). No significant effects emerged for craving or broader psychological outcomes (ps > .05), though favorable effects were observed on motivation to change (ambivalence: p = .020; d = -0.24, 95% CI [-0.65, 0.17]; taking steps: p = .002; d = 0.45 [0.04, 0.87]). Satisfaction with the intervention was moderate (M = 18.32 of 32), and participants completed on average 35.52 training sessions. Adverse events were reported infrequent (7.1%). DISCUSSION AND CONCLUSIONS: PROTECTapp is a promising mobile-based intervention to reduce PUI and enhance motivation to change in young adults. Its brevity, scalability, and safety profile highlight its potential as a low-threshold preventive or adjunctive intervention for young individuals at-risk.

18 July 2026

Read appraisal →
Systematic ReviewEvidence: Insufficient
15CEBM

Analytical methods : advancing methods and applications

Emerging trends in amino acid detection: wearable devices and machine learning-assisted signal processing

As critical metabolic biomarkers, amino acids exert essential physiological functions, and their abnormal levels are closely associated with a range of diseases such as cancer, neurodegenerative disorders, and cardiovascular conditions. In recent years, amino acid analysis technologies have achieved remarkable advancements in the field of personalized healthcare. This review navigates the latest progress in amino acid analysis, with a focus on two prominent emerging trends. The first trend encompasses the development of flexible and wearable sensing devices for non-invasive, continuous amino acid monitoring in diverse biofluids, such as blood, sweat, and interstitial fluid. Particular attention is given to their design principles, operational mechanisms, practical applications, and key performance metrics. The second trend involves the application of machine learning (ML) for processing and interpreting complex response signals. Specifically, this review discusses how various ML approaches, including classical chemometric linear regression models, deep learning models, support vector machines (SVMs), ensemble learning, and tree-based models, address common challenges in complex environments, such as signal interference and nonlinear drift. Furthermore, this review outlines the current challenges and proposes future research directions, aiming to advance amino acid analysis toward more intelligent, integrated, and personalized health monitoring systems.

17 July 2026

Read appraisal →
otherEvidence: Weak
50CEBM

Clinica chimica acta; international journal of clinical chemistry

Cortisol as a stress biomarker: analytical perspectives and challenges

Point-of-care testing (POCT) has emerged as a transformative approach in modern healthcare for the rapid detection of physiological abnormalities through minimally or non-invasively collected samples. Biological matrices such as saliva, urine, hair, blood, and interstitial fluid contain clinically significant biomarkers that may serve as indicators of physiological disorders. Among these, cortisol is the key stress biomarker and exerts substantial effects on body metabolism, acting on both peripheral tissues and the CNS. This review comprehensively outlines the evolution of cortisol detection strategies, progressing from conventional laboratory-based immunoassays to advanced analytical platforms, including electrochemical biosensors, wearable devices, and microfluidic systems. Accurate and reliable detection of elevated cortisol levels is crucial for improving diagnostic, therapeutic, and preventive strategies for stress-related disorders. By tracing the analytical window, the work describes the detection of cortisol from traditional immunoassays to innovative biosensing platforms. Moreover, recent advances in nanomaterials, sensor design, and data integration have enabled continuous, on-site monitoring of cortisol levels, thereby enhancing their applicability across diverse clinical and non-clinical settings. This integration of physiological insight with technological advancement provides a comprehensive overview of developments in cortisol assessment, connecting fundamental endocrine science with practical diagnostic applications. The review underscores the potential of next-generation POCT systems to improve early diagnosis, therapeutic monitoring, and personalized healthcare through real-time biomarker analysis.

17 July 2026

Read appraisal →
observationalEvidence: Insufficient
15CEBM

Biomedical physics & engineering express

NurtureNest: an IoT-wearable predictive analytics framework for real-time maternal risk assessment

Pregnancy-related complications are increasing globally, necessitating timely and accurate risk prediction for effective clinical intervention. This paper presents NurtureNest, an internet of things-enabled machine learning framework for automated pregnancy risk assessment. The system integrates wearable sensor data from smartwatches with user-provided clinical parameters via a mobile application, enabling continuous remote monitoring. Historical clinical data are used to train multiple machine learning models for multi-class classification of pregnancy risk into low-, medium-, and high-risk categories. Ensemble learning techniques are independently trained and evaluated alongside conventional machine learning classifiers to assess their effectiveness in pregnancy risk prediction. Experimental results show that LightGBM achieves the highest performance with 92.86% test accuracy and 95.04% cross-validation accuracy. Model performance is validated using ROC analysis and feature importance evaluation. The proposed framework enables early risk detection and supports timely clinical decision-making, improving maternal healthcare outcomes.

17 July 2026

Read appraisal →
otherEvidence: Weak
35CEBM

NEJM catalyst innovations in care delivery

Thinking about the Impact of Artificial Intelligence on U.S. Health Care Costs and Spending Growth

This article critically examines the potential impacts of artificial intelligence (AI) on prices, total spending, and the rate of spending growth in the areas of prescription drug innovation and expanded patient access to care - including developments in remote patient monitoring, chronic care management, direct-to-consumer health care, clinical decision support, and nonclinical administrative labor. Based on observed industry practices, economics, and public policy, the article presents a thought exercise on the most likely effects for patients, providers, payers, and the broader health care system. The authors argue that under the still-dominant fee-for-service payment model, as well as the highly consolidated hospital and insurance markets, AI is more likely to increase total costs and spending growth in the short to medium term rather than slow them, even as it delivers substantial access and clinical quality improvements for patients. The cost-bending potential of AI varies distinctly by fee-for-service versus value-based payment. Regulators need to introduce policy and reimbursement levers for AI to slow cost growth. The authors conclude that, without significant changes in the health care system's financial incentives and market structures, AI will not slow cost growth.

16 July 2026

Read appraisal →
Systematic ReviewEvidence: Weak
45CEBM

Journal of medical Internet research

User Acceptability and Adoption of AI-Generated Lifestyle Intervention Recommendations: Scoping Review and Theoretical Integration

BACKGROUND: Artificial intelligence (AI)-generated lifestyle recommendations are increasingly used to support health behavior change. However, AI advice does not necessarily mean that users will accept or adopt those recommendations. Although prior reviews have examined AI-enabled lifestyle interventions and health behavior technologies, fewer have focused on whether users accept and adopt AI-generated recommendations. OBJECTIVE: This scoping review aimed to map user acceptability and adoption of AI-generated lifestyle recommendations in user-facing systems used by end users or caregivers. Objectives were to characterize systems and evaluation contexts, clarify how recommendation-level outcomes were conceptualized and measured, synthesize shaping factors, and develop an evidence-informed framework to guide future research, evaluation, and design. METHODS: Following JBI (Joanna Briggs Institute) guidance and the PRISMA-ScR (Preferred Reporting Items for Systematic Reviews and Meta-Analyses extension for Scoping Reviews), we searched Ovid MEDLINE, Ovid Embase, APA PsycInfo via ProQuest, Web of Science Core Collection, Scopus, ACM Digital Library, and IEEE Xplore from database inception to May 5, 2026. The initial search was conducted on November 21, 2025, and an updated search was conducted on May 5, 2026. Eligible studies reported empirical end-user or caregiver data, evaluated AI-generated lifestyle recommendation content delivered without manual review or editing, and reported an acceptability or adoption outcome linked to recommendations. English empirical papers and conference papers were included. Data were charted on study, system, outcome, measurement, factor, and theoretical characteristics. Quality was assessed with the Mixed Methods Appraisal Tool. Findings were synthesized descriptively and through evidence mapping. RESULTS: Searches yielded 12,997 records; 8570 unique records were screened, and 21 studies were included. Most were published in 2025 or 2026 (17/21, 81%). Large language model-centered systems were the most common format (12/21, 57.1%). Outcomes were concentrated in acceptability-related perceptions, such as satisfaction or enjoyment, perceived quality or fit, and persuasiveness, whereas adoption-related outcomes were assessed less often and mainly reflected intention, in-study uptake, or short-term enactment. Factors clustered across system capabilities, content properties, individual states and capacities, and contextual constraints. Findings informed an integrative perception-intention-enactment framework positioning acceptability and adoption as a system-content-user-context process. CONCLUSIONS: This review extends prior AI and digital health reviews by shifting attention toward how users perceive, intend to follow, and enact AI-generated lifestyle recommendations. Acceptability and adoption appear to depend on systems eliciting and adapting to context, content being actionable and credible, users having the capacity to interpret, trust, and engage with recommendations while retaining control, and resources, routines, and social contexts allowing enactment. The framework can guide theory-driven evaluation, outcome selection, and system design by identifying where recommendation processes may succeed or fail, but should be interpreted as preliminary and evidence-informed rather than causal. By integrating implementation, behavioral, and human-AI perspectives, this review provides a foundation for moving AI-generated lifestyle recommendations from technically plausible outputs toward user-centered, context-sensitive, and behaviorally actionable support.

16 July 2026

Read appraisal →
observationalEvidence: Weak
25CEBM

International journal of medical informatics

Examining user-AI interaction patterns in health-information queries

In this study, we examine how individuals utilize generative artificial intelligence (GAI) when seeking health-related information. Using a dataset of user-GAI chat logs available on Hugging Face, we analyzed real-world interactions in which users posed health-related questions to a generative model. We applied a combination of data and text-analytic methods to categorize these interactions, including supervised machine learning techniques such as Support Vector Machines (SVMs). SVMs were selected for their efficiency and strong performance in high-dimensional text classification tasks, and used to identify recurrent themes in user queries and interactions. We found that users frequently consult AI chatbots for symptom exploration, medical education, mental health support, and general health advice. The findings suggest that GAI tools may not only function as informational resources, but also as preliminary support tools that can shape users' health knowledge and encourage them to seek consultation with medical professionals.

16 July 2026

Read appraisal →
diagnosticEvidence: Weak
35CEBM

Biosensors & bioelectronics

Functional wearable hydrogel microneedle platform for continuous ketone monitoring: Translating from rodents to humans.

Diabetic ketoacidosis (DKA) is a life-threatening complication of diabetes, driven by excessive ketone production; it is most common in type 1 diabetes. Continuous ketone monitoring (CKM) could enable earlier detection and prevention of DKA, yet no wearable CKM solution is clinically available. We present a wearable CKM patch that couples hydrogel microneedles (HMNs) for painless interstitial fluid (ISF) access with an enzymatic, electrochemical biosensor for on-patch detection of ketone bodies. The HMN patch and ketone biosensor were individually characterized and optimized in vitro, then integrated into a single platform. We validated the performance of the integrated CKM sensing patch across species: in vivo testing in healthy and diabetic rats, translation to a swine model, and a pilot evaluation in human participants. The CKM sensing patch reliably tracked dynamic ketone fluctuations in all models. To our knowledge, this is the first HMN-based biosensor tested in humans. These results establish a minimally invasive, wearable approach for continuous ketone tracking with the potential to transform outpatient DKA monitoring and early intervention.

16 July 2026

Read appraisal →
Randomised Controlled TrialEvidence: Insufficient
30CEBM

Analytical chemistry

Biodegradable Self-Powered Electrotherapy Patch for Integrated Smart Wound Management

Smart patches based on multimodal wearable devices enable real-time physiologic monitoring and proactive interventions to promote wound healing. Herein, we describe a biodegradable wearable electrotherapy patch (E-patch) that integrates noninvasive self-powered electrical stimulation (ES) therapy for tissue regeneration and a multiplexed electrochemical biosensor array for continuous monitoring of wound status. Custom-developed supercapacitor arrays (SCs) supply stable energy for ES, and constructed wearable biosensors enable sensitive monitoring of biomarkers in the wound exudate. As-fabricated wearable E-patches degrade harmlessly after operation, significantly reducing the environmental pollution pressure associated with flexible electronics. In vitro studies demonstrated that an applied electric field (EF) significantly promotes cell-directed alignment, which is crucial for tissue regeneration and remodeling. In vivo investigations in the Sprague-Dawley (SD) rat model illustrate that combination therapy dramatically accelerates wound healing. Overall, this work provides a promising strategy toward integrated smart wound management and future feedback-assisted wearable therapeutic systems.

15 July 2026

Read appraisal →
observationalEvidence: Weak
20CEBM

PloS one

Transcriptomic characterization of key psoriasis-associated genes based on single-cell RNA-seq and machine learning

BACKGROUND: Psoriasis is a multifaceted skin and systemic disorder driven by a complex interplay of genetic, immunological, and environmental factors. Genetic predisposition plays a pivotal role, with the IL-17/IL-23 immune axis recognized as a central pathogenic pathway. Ongoing research, however, continues to uncover additional critical drivers, cytokines, intracellular signaling networks, and potential therapeutic targets. METHODS: Single-cell RNA sequencing (scRNA-seq) datasets comprising both psoriatic and healthy samples were obtained from the Gene Expression Omnibus (GEO). Cell-type proportions were estimated using Cell-type Identification by Estimating Relative Subsets of RNA Transcripts (CIBERSORT), and weighted gene co-expression network analysis (WGCNA) was applied to explore correlations between cell types and gene signatures. Machine learning algorithms were subsequently employed to identify four psoriasis-associated key genes: DEFB4A, GJB2, SERPINB3, and SERPINB13. Their expression was validated in bulk RNA-seq datasets. Using scRNA-seq data, we further investigated the lesional regulatory roles of these genes and their associated pathway alterations, and we proposed targeted therapeutic strategies. RESULTS: A series of algorithms identified 271 hub genes significantly associated with psoriasis lesions and basal cells. Machine learning analysis refined this set to four key genes in psoriasis: DEFB4A, GJB2, SERPINB3, and SERPINB13. CONCLUSIONS: These four psoriasis-associated driver genes were upregulated in lesional skin. We also screened small-molecule compounds targeting these genes, offering potential therapeutic strategies.

14 July 2026

Read appraisal →
Randomised Controlled TrialEvidence: Weak
35CEBM

Daru : journal of Faculty of Pharmacy, Tehran University of Medical Sciences

Deep fused explainable neural architecture for the automatic recognition of therapeutic medicinal plants

BACKGROUND: Medicinal plants have been utilized for centuries for their beneficial properties and significant impact on traditional medical systems worldwide. Indian medicinal plants and their therapeutic properties have recently gained increased attention from the global healthcare community. Accurate identification and classification of medicinal plants ensure proper therapeutic applications. However, botanists manually identify these plants, which is subjective, labor-intensive, and time-consuming. OBJECTIVE: The main objective of this study is to develop an accurate and automated classification of these medicinal plants for diverse applications such as conservation initiatives, quality assurance, and the formulation of novel herbal treatments. METHODS: A Deep Fused Explainable Neural Architecture termed MedPlantNet is proposed for robust recognition of medicinal plant leaves. The deep discriminative features are extracted from the medicinal plant images using a fine-tuned Deep Convolutional Neural Network (DCNN). A customized Support Vector Machine is developed to efficiently classify these extracted deep features. To enhance model interpretability, a meta-visualization framework incorporating Grad-CAM and Occlusion Sensitivity visualization is implemented to identify the regions of the image that contribute to the classification decisions. We utilized the Mendeley benchmark Medicinal plant leaf database with 30 classes to train and evaluate the model. RESULTS: To assess the effectiveness of the proposed model, accuracy, specificity, sensitivity, and AUC are calculated. The proposed system outperforms existing state-of-the-art (SOTA) algorithms for medicinal plant classification, achieving an optimal maximum validation accuracy of 99.87% and AUC of 1. To further validate the generalizability external validation was performed on Bangladeshi medicinal plant dataset and MedLeaves dataset where the proposed model MedPlantNet attained 99.27% and 99.80% validation accuracy respectively. CONCLUSION: MedPlantNet provides an automatic framework for medicinal plant identification based on deep feature extraction, classification and explainable visualization techniques. The results obtained indicate its potential utility in supporting biodiversity conservation, quality assurance and herbal medicine research.

13 July 2026

Read appraisal →
Systematic ReviewEvidence: Weak
30CEBM

Vaccine

Mobile health strategies to improve HPV vaccination uptake among parents of adolescent girls: a scoping review

BACKGROUND: Human papillomavirus (HPV) vaccination is a highly effective primary prevention strategy against cervical cancer. However, despite strong evidence of vaccine efficacy and safety, uptake among adolescent girls remains suboptimal globally, particularly in low- and middle-income countries. Mobile health (mHealth) technologies have emerged as promising tools to address parental concerns, reduce vaccine hesitancy, and strengthen parental engagement in vaccination decision-making. OBJECTIVE: This scoping review aimed to synthesize evidence on the types of mHealth strategies implemented to improve HPV vaccine uptake among parents of adolescent girls and to examine the outcomes of these interventions. METHODS: A systematic search of major databases identified studies evaluating mHealth interventions targeting parental decision-making regarding HPV vaccination. Eligible interventions included text messaging, mobile applications, chatbots, web-based educational platforms, and other digital health tools. Data were charted and synthesized narratively. KEY FINDINGS: mHealth interventions were associated with significant improvements in HPV vaccine uptake across multiple settings. SMS-based interventions consistently increased vaccination rates, while chatbot, web-based, and mobile app interventions improved parental knowledge, vaccine literacy, intention to vaccinate, and, in several studies, actual vaccine uptake. However, intervention effectiveness varied by socioeconomic context, digital access, and cultural factors. CONCLUSION: mHealth technologies represent cost-effective and scalable strategies to improve HPV vaccine uptake by enhancing parental knowledge, reducing hesitancy, and supporting informed decision-making. Future implementation should prioritize culturally tailored, equity-focused approaches and integrate digital interventions with community-based engagement to maximize impact.

12 July 2026

Read appraisal →
qualitativeEvidence: Moderate
75CEBM

Journal of medical Internet research

Addressing the Key Challenges of Decentralized Clinical Trials in Europe: Multistakeholder Perspective Delphi Study

BACKGROUND: Decentralized clinical trials (DCTs) represent an emerging model in clinical research, accelerated by the restrictions imposed during the COVID-19 pandemic. By leveraging digital technologies and local health care resources, DCTs aim to increase accessibility and reduce participant burden compared to traditional site-based models, which often face recruitment failures and high attrition rates. While various regulatory initiatives in Europe, such as the Accelerating Clinical Trials in the European Union program and the European Medicines Regulatory Network recommendation paper (updated in October 2025), have sought to facilitate their implementation, the widespread adoption of DCTs remains limited due to significant operational, regulatory, and technological challenges, including platform fragmentation and gaps in digital literacy. OBJECTIVE: This study aimed to identify and prioritize actionable solutions to the main challenges of DCT implementation in Europe from a multistakeholder perspective, gathering insights to address specific ethical, legal, and operational barriers. METHODS: Building on a preceding strengths, weaknesses, opportunities, and threats analysis, a 2-round Delphi study was conducted, involving 26 experts in clinical trials, ethics, law, regulation, and patient engagement between March and May 2023. In the first round, 309 open-ended responses were collected via REDCap (Research Electronic Data Capture; Vanderbilt University) surveys and underwent systematic inductive content analysis using ATLAS.ti (ATLAS.ti Scientific Software Development GmbH) with independent double coding. This process resulted in 244 unique proposals that were categorized according to 6 key challenges. In the second round, 39 synthesized proposals were evaluated using a 4-point Likert scale. Consensus was defined as ≥80% agreement on the appropriateness of each proposal. RESULTS: High levels of consensus were achieved, with 32 out of the 39 proposals reaching the threshold and 14 achieving 100% unanimity. Overall, 82% of the proposals were rated as "appropriate" or "very appropriate." Key recommendations included providing support and training for health care professionals, enhancing investigational medicinal product and biological sample logistics through validated technologies, improving collaboration with local health care providers, fostering regulatory harmonization while respecting national specificities, strengthening capacity-building initiatives, and promoting accessible, user-friendly digital tools supported by hybrid trial models. Conversely, proposals such as peer-to-peer participant support and the centralization of ethics reviews at the European Union level failed to reach consensus. CONCLUSIONS: The study offers a prioritized compilation of expert-driven recommendations for overcoming current barriers to DCT implementation in Europe. The adoption of these recommendations could support the development of more inclusive, efficient, and sustainable decentralized research frameworks across diverse health care systems.

12 July 2026

Read appraisal →
Systematic ReviewEvidence: Insufficient
60CEBM

JMIR research protocols

Applications of Large Language Models in Ovarian Cancer Management: Protocol for a Systematic Review and Meta-Analysis

BACKGROUND: Ovarian cancer (OC) is a highly fatal gynecologic malignancy with complex management challenges and limited long-term survival for advanced stages. Large language models (LLMs)-including systems such as GPT-4, Claude, Google Gemini, and others-are emerging artificial intelligence (AI) tools capable of performing health care-related tasks such as diagnostic support, treatment planning, report generation, and patient communication. However, their applications in OC care have not yet been comprehensively assessed. OBJECTIVE: This protocol outlines a systematic review and meta-analysis aimed at evaluating the use, performance, and clinical impact of LLMs in OC management. We will examine how LLMs have been applied across various domains (eg, diagnosis, prognosis, treatment planning, and patient engagement), the metrics used to assess their performance (eg, accuracy, sensitivity, and area under the curve), and their strengths and limitations. METHODS: This review will be conducted in accordance with PRISMA-P (Preferred Reporting Items for Systematic Reviews and Meta-Analyses Protocols) guidelines. A comprehensive search strategy will be implemented across biomedical, technical, and Chinese-language databases (eg, PubMed, Embase, Web of Science, IEEE Xplore, and China National Knowledge Infrastructure) from inception to December 31, 2025. Eligible studies include clinical evaluations, validation studies, and real-world implementation reports involving LLMs in OC care. Two independent reviewers will perform screening, data extraction, and quality appraisal using validated tools (eg, version 2 of the Cochrane risk-of-bias tool for randomized trials, Risk of Bias in Nonrandomized Studies of Interventions, Quality Assessment of Diagnostic Accuracy Studies 2, and Prediction Model Study Risk of Bias Assessment Tool+AI). Outcomes of interest include model performance metrics, clinical process impacts, safety concerns, and usability. Meta-analyses will be conducted where feasible using random-effects models in R (meta, metafor, and mada packages), including bivariate models for sensitivity and specificity. RESULTS: The review is currently in progress. The PROSPERO registration has been completed, and the literature search and selection process is underway. Study selection, data extraction, and quality assessment are expected to be completed by mid-2026. Final results will include pooled performance metrics (eg, accuracy, F1-score, and area under the curve), qualitative insights into clinical integration, and identification of limitations such as reporting bias or insufficient external validation. CONCLUSIONS: This systematic review will provide the first comprehensive synthesis of evidence on the application of LLMs in OC care. It will identify promising use cases, highlight safety and reporting challenges, and inform future research directions. The findings are expected to support evidence-based integration of LLMs into gynecologic oncology workflows while promoting transparency and methodological rigor in AI evaluation.

12 July 2026

Read appraisal →
Systematic ReviewEvidence: Weak
25CEBM

Current diabetes reports

Digital Management of Early-Onset Type 2 Diabetes: Empowerment, Challenges, and Future Outlook

PURPOSE OF REVIEW: Early-onset type 2 diabetes (EOT2D), defined as a diabetes diagnosis before 40 years of age, is rising globally and associated with an aggressive disease course and early complications. This review examines the role of digital health technologies (DHT) in addressing the unique clinical and life-course challenges of EOT2D. RECENT FINDINGS: DHT, including continuous glucose monitoring, mobile health applications, digital therapeutics, telemedicine, remote patient monitoring, wearable devices, and artificial intelligence-based analytics, have demonstrated modest improvements in glycemic control, weight management, and patient engagement in people with type 2 diabetes. However, evidence in adults with EOT2D remains limited. Compared with usual-onset T2D, people with EOT2D may derive particular benefits due to higher digital literacy, greater lifestyle variability, and longer anticipated disease duration. Although DHT shows promise for improving empowerment and care integration in EOT2D, important gaps persist, including a lack of EOT2D-specific trials, digital divide-related inequities, interoperability challenges, and reimbursement barriers. Future research should prioritize tailored interventions and hybrid care models to optimize long-term outcomes in this high-risk population.

11 July 2026

Read appraisal →
Systematic ReviewEvidence: Weak
60CEBM

BMC oral health

The association between oral hygiene and metabolic dysfunction-associated steatotic liver disease - a systematic review

BACKGROUND: Given the increasingly evident links between oral health and systemic health, the aim of this review was to investigate whether oral hygiene practices influence the risk or progression of metabolic dysfunction-associated steatotic liver disease (MASLD). The findings support the potential importance of good oral hygiene and regular dental visits, especially for patients with MASLD. METHODS: This systematic review was registered with PROSPERO (CRD420251041530), and PubMed, Scopus, Embase and Web of Science were used for literature research up to June 2026. In accordance with the Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) guidelines, screening, study selection and quality check were performed using Rayyan and the Joanna Briggs Institute (JBI) critical appraisal checklists. Studies assessing the potential link between the frequency of oral hygiene procedures and the incidence, prevalence or progression of MASLD were selected for a systematic synthesis. The old terms non-alcoholic fatty liver disease (NAFLD) and non-alcoholic steatohepatitis (NASH) were considered as well. RESULTS: A total of 570 results were identified of which six studies met the inclusion criteria. Five of them reported a correlation between higher tooth brushing frequency and a lower incidence, prevalence or progression of NAFLD/NASH compared to individuals with lower tooth brushing frequency. Notably, one study revealed significantly worse liver parameters among NASH patients without regular dental visits than among those who visited a dentist more than once a year. CONCLUSIONS: An association between oral hygiene habits and NAFLD/NASH was found in all included studies. A greater frequency of daily tooth brushing as well as regular dental visits might have beneficial effects on both the development and progression of fatty liver disease. However, there is not enough evidence to establish causality yet and the impact of confounding factors must be further investigated. CLINICAL TRIAL REGISTRATION: Not applicable.

11 July 2026

Read appraisal →
Systematic ReviewEvidence: Weak
55CEBM

Journal of epidemiology and community health

Analysis of the potential association between physical activity and skin autofluorescence: a systematic review.

INTRODUCTION: Accumulation of advanced glycation end products, measured by skin autofluorescence (SAF), has been shown to be associated with several chronic non-communicable diseases, particularly cardiovascular diseases (CVDs). The promotion of physical activity (PA) as a strategy for the prevention of CVD by modifying healthy habits has been widely studied. AIM: To assess the evidence for the association between PA and SAF in the general adult population. METHODS: A systematic review was conducted following Preferred Reporting Items for Systematic Reviews and Meta-Analyses guidelines and the Synthesis Without Meta-analysis framework. A search was performed in MEDLINE (via PubMed), Web of Science, Scopus, Cochrane Library and SportDiscuss (via EBSCOhost), from inception to September 2024. Study quality was assessed using the National Heart, Lung and Blood Institute tools and the certainty of evidence was evaluated with Grading of Recommendations, Assessment Development and Evaluation. Vote counting based on the direction of effect was used as the standardised synthesis metric. RESULTS: In the systematic review, 17 studies were included. The qualitative synthesis showed a predominant consistency in favour of a beneficial association. Specifically, 58.8% of the studies reported a statistically significant inverse association, indicating that higher levels of PA or exercise frequency are correlated with lower SAF levels. The remaining studies (41.2%) reported non-significant results, though several showed favourable trends. No studies reported a positive association between PA and SAF. The quality of the studies was generally fair, and the certainty of evidence low. CONCLUSIONS: PA is inversely associated with SAF. Therefore, while causality cannot be proven, it is hypothesised that PA may reduce SAF and thus have a positive impact on health.

11 July 2026

Read appraisal →
qualitativeEvidence: Moderate
70CEBM

JMIR mHealth and uHealth

Perceived Sensitivity of Sensor-Based Digital Health Data: Qualitative Interview Study

BACKGROUND: Digital health tools are increasingly used in mental health care to passively collect patient data and analyze health status outside of clinical settings. While technologies such as digital phenotyping, affective computing, and computational behavioral analysis offer new insights into symptom manifestation in daily life, they generate large volumes of potentially sensitive data that raise significant data privacy concerns, requiring high levels of patient awareness and consent. Empirical research is lacking on stakeholder understandings toward the sensitivity of these data and expectations for data stewardship, perspectives that are critical for developing robust informed consent and data protection policies for digital health data use. OBJECTIVE: This study aimed to explore key stakeholder perspectives on the sensitivity of computer perception (CP) data, trust in existing data protections, willingness to share CP data externally, and desire for transparency of CP data transactions outside of the clinical space. METHODS: As part of a larger, multisite study, we conducted qualitative interviews (n=40) via Zoom (Zoom Communications, Inc) with 20 adolescents (aged 12-17 years) familiar with CP tools and their caregivers (n=20). Interviews consisted of a series of open-ended questions regarding stakeholders' perspectives on privacy, data security, and the use and exchange of CP data. We developed a qualitative codebook to identify and label thematic patterns in responses to questions addressing the topics above, using thematic content analysis to identify themes inductively. Each interview was coded by merging work from at least two separate coders, and several team members contributed to qualitative analysis. RESULTS: Most adolescents and caregivers viewed CP data as highly sensitive and expressed a reluctance to share these data beyond their clinical teams. While many participants expressed trust in existing data protections to protect CP data, they often misunderstood or overestimated the extent of protections to safeguard CP data. CONCLUSIONS: Our findings underscore the critical need for clear and effective patient communication and education about the risks, benefits, and protections associated with CP data through informed consent protocols. To promote greater transparency, understanding, and trust, we recommend 5 strategies: educating patients about data protection; studying secondary data exchange and reidentification risks; strengthening transparency regulations; improving data traceability mechanisms, such as distributed ledger technologies, to enhance data traceability and auditability; and adopting dynamic consent models.

11 July 2026

Read appraisal →
observationalEvidence: Weak
30CEBM

PloS one

Intergenerational patterns of digital use: Evidence from a large cross-sectional study

This study aims to determine the level of digital use across six generations (Greatest, Silent, Baby Boomers, and Generations X, Y and Z) within a district in North East England (population ~200,000). A cross-sectional descriptive study was conducted via an online and paper-based survey from February to May 2022, mailed to households in the district (N = 98,260). Respondents estimated their weekly use of digital tools and the internet at home. One-way ANOVA and Tukey-Kramer HSD tests were used to analyse differences in digital resource uses across six generations. To account for underlying sociodemographic characteristics, a follow-up ANCOVA analysis was also performed. Content Analysis was used for qualitative responses from an open-ended survey question. A total of 9,181 completed surveys were analysed. The sample was skewed towards older, homeowner adults (mean age 63, 60% female). Findings revealed that respondents spend less time online than other British cohorts. Baby Boomers and Generation X self-reported statistically significant differences in the level of digital use compared to all other generations. Younger generations (Y and Z) self-reported, on average, a larger amount of time spent on both digital tools and the internet. Members of Greatest and Silent Generations had the lowest hours spent on digital tools and the internet. The results suggest that public health initiatives should prioritise strategies bridging the digital divide between generations. Targeted training programs pairing younger, tech-savvy individuals with older adults could enhance digital literacy. Additionally, integrating user-friendly digital health platforms that cater to varying levels of technological proficiency will encourage wider adoption. These strategies not only foster intergenerational collaboration but also drive successful digital health and care transformations, ensuring equitable access to technological advancements for all age groups.

11 July 2026

Read appraisal →
Systematic ReviewEvidence: Moderate
65CEBM

International journal of qualitative studies on health and well-being

Between duty and constraint: a qualitative systematic review of healthcare providers' ethical challenges and moral stressors in caring for undocumented migrants

PURPOSE: Healthcare providers play a critical role in delivering care to undocumented migrants, who face systemic barriers to healthcare access. While research has documented undocumented migrants' legal and policy barriers, less is known about the ethical challenges and moral stressors experienced by providers. This systematic review synthesizes qualitative studies on providers' experiences when delivering care to undocumented migrants. METHODS: The review followed PRISMA guidelines. PubMed, Embase, CINAHL, and the Cochrane Library were searched for relevant qualitative studies. Studies were screened by title, abstract, and full text using predefined inclusion and exclusion criteria. Quality assessment was conducted using the CASP checklist. Data were synthesized using the Qualitative Analysis Guide of Leuven (QUAGOL), integrating Graneheim and Lundman's approach to qualitative content analysis. RESULTS: The systematic search identified 37 qualitative studies. Analysis revealed 58 themes and subthemes, organized under five key concepts: experiences, perceptions, attitudes, practices and coping mechanisms, and ethical challenges. Providers reported moral distress, emotional exhaustion, and professional dilemmas arising from legal constraints, resource limitations, and conflicting obligations. Ethical tensions centered on beneficence vs. non-maleficence, professional duty vs. legal compliance, and the moral dilemma of deservingness. Across these findings, the synthesis identifies ethical burden-shifting as a central analytical contribution: restrictive systems transfer the moral and practical consequences of exclusionary arrangements onto providers and undocumented migrants at the point of care. CONCLUSION: This review shows that providers' ethical challenges are not only individual clinical dilemmas but structurally generated moral stressors. Addressing these challenges requires structural interventions, including policy reforms that reconcile professional ethics with legal constraints and institutional support to mitigate moral distress.

10 July 2026

Read appraisal →
Systematic ReviewEvidence: Weak
60CEBM

Journal of medical Internet research

The Association Between eHealth Literacy and Health Behaviors During and Since the COVID-19 Pandemic: Systematic Review and Meta-Analysis

BACKGROUND: Since COVID-19, health information seeking, service navigation, and routine care have become increasingly digitally mediated. It remains unclear whether the association between eHealth literacy and health behaviors is consistent across behavioral domains, populations, and analytic frameworks. OBJECTIVE: This systematic review and meta-analysis synthesized COVID-19 and post-COVID-19 evidence on the association between eHealth literacy and health behaviors and examined variation across study contexts. METHODS: We conducted a PRISMA (Preferred Reporting Items for Systematic Reviews and Meta-Analyses) 2020-compliant, PROSPERO-registered review (CRD420251009048). PubMed, Embase, Web of Science Core Collection, CINAHL Ultimate, and Scopus were searched from January 1, 2020, through March 27, 2026. Eligible observational studies assessed eHealth literacy and reported analyzable associations with health behavior outcomes collected in 2020 or later. Health behaviors were classified as health decision-making, health-promoting, or health management behaviors. Correlation coefficients, grouped odds ratios (ORs), and continuous ORs were synthesized separately using random-effects models with Knapp-Hartung adjustment. Certainty was assessed using the Grading of Recommendations Assessment, Development, and Evaluation framework. RESULTS: In total, 19 studies were included: 10 contributed correlation coefficients, 6 grouped ORs, and 3 continuous ORs. In the correlation-based synthesis, higher eHealth literacy was associated with more favorable health behaviors (pooled r=0.43, 95% CI 0.36-0.51; 95% prediction interval 0.22-0.61), with substantial heterogeneity (I2=80.50%). In the grouped OR synthesis, higher eHealth literacy was also associated with more favorable health behaviors (pooled OR 2.12, 95% CI 1.47-3.05; 95% prediction interval 1.11-4.06), with moderate heterogeneity (I2=46.33%). In the continuous OR synthesis, all 3 studies showed positive associations, but the pooled effect was not statistically significant (pooled OR 1.07, 95% CI 0.89-1.30; 95% prediction interval 0.84-1.37), with very high heterogeneity (I2=97.98%). Subgroup analyses showed a significant difference only by geriatric status in the grouped OR synthesis. Certainty was low for the grouped OR synthesis and very low for the other 2 syntheses. CONCLUSIONS: Higher eHealth literacy was generally associated with more favorable health behaviors in the correlation-based and grouped OR syntheses, whereas evidence from the continuous OR synthesis was inconclusive. Given the predominantly cross-sectional evidence base, heterogeneity, risk-of-bias concerns, and low to very low certainty, the association should be interpreted as contextual and associative rather than causal or uniform. This review is innovative in synthesizing COVID-19 and post-COVID-19 evidence, applying a functional classification of health behaviors, and analyzing distinct effect measures separately. Unlike previous reviews that summarized the association more broadly, it avoids a single mixed pooled effect and provides a cautious, context-specific interpretation. In practice, interventions should pair eHealth literacy improvement with trustworthy digital services, clinician support, and behavior-specific conditions that help translate digital information into sustained health-related action.

10 July 2026

Read appraisal →
Systematic ReviewEvidence: Weak
45CEBM

Cell host & microbe

Predicting antimicrobial resistance for precision medicine

Antibiotics are among medicine's greatest successes, but resistance evolution threatens their continued efficacy. Decades of research have deepened our understanding of the mechanisms and evolutionary dynamics of antimicrobial resistance. More recently, advances in machine learning (ML) and artificial intelligence (AI) show promise in predicting antimicrobial resistance in pathogens based on rapid whole-genome sequencing and other accessible data. In this perspective, we highlight advances in understanding the mechanisms and spread of antimicrobial resistance. We discuss how this knowledge, coupled with ML- and AI-based approaches, can inform the prediction of resistance and a precision-medicine strategy that targets pathogenic bacteria specifically, thereby limiting resistance evolution and collateral damage to the microbiome. These accurate predictions of bacterial vulnerabilities will enable the adaptation of classical antimicrobial treatments with adjuvants, as well as the use of novel, narrow-spectrum therapeutics. Implementing these strategies, while also identifying key challenges, will help bring this strategy into clinical practice.

10 July 2026

Read appraisal →
qualitativeEvidence: Weak
75CEBM

BMC public health

Motivating factors and possible barriers to participation in digital prevention courses of two statutory health insurance funds in Germany: a qualitative interview study

BACKGROUND: Digital prevention courses offered by statutory health insurance funds in Germany represent a low-threshold opportunity to support health-promoting behaviors through pre-recorded, video-based programs focusing on exercise, nutrition, or stress management. However, little is known about insured individuals' motivations and barriers to participation. This study aimed to identify relevant motivating and obstructive factors for participation in digital prevention courses and to explore differences compared to on-site formats. METHODS: We conducted semi-structured telephone interviews with 20 insured members of two large regional statutory health insurance funds in Germany (AOK North-East and AOK North-West) who had recently completed digital prevention courses. Participants were recruited via e-mail invitations sent by the health insurance funds following course participation. Five additional interviews were conducted with participants of on-site courses for contextual comparison. Interviews were conducted between February and December 2023, transcribed verbatim, and analyzed using qualitative content analysis with a deductive-inductive coding approach based on the study objectives and interview guide. RESULTS: The sample included 20 participants (mean age 41 years, range 25-72; 19 women and 1 man) who had completed digital prevention courses. Participants emphasized temporal flexibility, home-based access, and low preparation effort as key facilitators of participation. Additional motivators included health-related goals, free access, practical content, reminders, and the possibility to repeat sessions. Reported barriers included a lack of discipline, reduced social interaction and commitment compared to on-site formats, and work- and family-related demands. The availability of recorded sessions was perceived as a major advantage that facilitated course completion. Participants also valued competent instructors, but called for clearer differentiation of difficulty levels, broader course options, continued access to materials, and follow-up opportunities. The contextual comparison showed that on-site course participants particularly valued the sense of community and the possibility of individual corrections by instructors. CONCLUSIONS: Digital prevention courses are well accepted and can promote participation due to their flexibility and accessibility. However, challenges such as reduced social interaction and the need for self-discipline limit their full potential. To enhance their effectiveness and reach, digital prevention programs should integrate more tailored content, structured follow-up offers, and interactive or hybrid elements that foster engagement and social connectedness. TRIAL REGISTRATION: German Clinical Trial Register DRKS00029761; Registration date 27.07.2022; https://drks.de/search/de/trial/DRKS00029761 .

9 July 2026

Read appraisal →
otherEvidence: Strong
95CEBM

BMC medical research methodology

Variable selection for clinical prediction models in low-dimensional data - a simulation study comparing traditional regression and machine learning methods

PURPOSE: A wide range of methods exist for developing a clinical prediction model (CPM) and for performing variable selection. Our purpose was to develop a fair simulation study design and to investigate the properties, strengths, and weaknesses of different methods to predict a continuous outcome in low-dimensional data situations. METHODS: In this simulation study, we conducted a neutral comparison of traditional (linear regression with stepwise selection) and machine learning (regularized regression with elastic net, gradient boosting, random forest) variable selection strategies to derive a CPM. The generated datasets included a total of 15 variables, with 8 of those being predictor variables. Four data- and outcome-generating mechanisms with increasing complexity produced data structures typical for biomedicine covering linear associations and gradually introducing non-linear and non-additive elements into the data structure. RESULTS: All methods generally performed better with increasing sample size and less noise in the data. Gradient boosting with regression models and with trees as base learners, and the elastic net regularized regression included nearly all variables (i.e., both the predictor and non-predictor variables), especially with increasing sample size. The linear regression model with stepwise selection (LMSS) showed the best trade-off between correctly including the predictors and excluding the non-predictor variables in most of the scenarios, even when the functional form of continuous predictors deviated from linearity. In more complex data, variable selection using the Boruta or Hapfelmeier approach for random forest performed similar to LMSS. CONCLUSION: The sample size must be sufficiently large to enable the methods to reliably identify the predictor variables and to ensure that the developed CPMs are accurate and well-calibrated. LMSS revealed good properties and the random forest with the Boruta or Hapfelmeier approach are suitable alternatives if complex associations between predictors and outcomes are assumed.

9 July 2026

Read appraisal →
Randomised Controlled TrialEvidence: Weak
45CEBM

Physiological measurement

A comprehensive inference-time augmentation framework in physiological signals: application to PPG-based AF detection

Objective.Accurate classification of physiological signals in real-world deployments is challenged by sensor noise, motion artifacts, and distribution shifts between training and deployment data. Inference-time augmentation (ITA), which applies augmentations during inference rather than retraining, offers a simple, model-agnostic mechanism to improve robustness. However, ITA application to physiological signals has remained narrow in scope, relying on limited augmentation methods with fixed, unoptimized parameters. This work proposes a unified ITA framework to address that gap.Approach.The framework incorporates 13 augmentation methods spanning time-domain, amplitude-domain, frequency-domain, and artifact-injection transformations, with hyperparameters systematically optimized via Bayesian optimization. We evaluate the framework on atrial fibrillation (AF) detection from 30 s photoplethysmography (PPG) signals using two deep learning architectures, generative pre-trained transformer (GPT)-PPG (in two sizes) and ResNet, across five datasets comprising more than 400 patients and ∼9800 h of PPG recording. Two evaluation strategies are assessed: standard ITA applied to all inputs, and selective ITA applied to initially positive predictions.Main results.Standard ITA consistently improved the area under the receiver operating characteristic curve (AUROC, up to 8.5% for GPT-PPG and 0.7% for ResNet) and the area under the precision-recall curve (AUPRC, up to 10.6% for GPT-PPG and 0.8% for ResNet) across all model-dataset combinations. Selective ITA further reduced average FPR by up to 4.4% (GPT-PPG) and 1.3% (ResNet) on non-AF PPG datasets.Significance.These findings establish ITA as a practical, model-agnostic approach for improving the reliability of PPG-based AF classification in deployment settings where retraining is not feasible, with broader applicability to physiological signal analysis.

9 July 2026

Read appraisal →
Systematic ReviewEvidence: Weak
45CEBM

Nutrition journal

Mindfulness and sustainable diets: a meta-analysis and CO2 emission savings scenarios

BACKGROUND: Dietary habits are a major driver of both human health outcomes and global greenhouse gas emissions, with excessive meat consumption being a key contributor. Mindfulness, understood as purposeful and non-judgmental awareness of the present moment, has been suggested as a psychological factor that can foster more sustainable food choices. However, the empirical evidence on the link between mindfulness and sustainable diets remains fragmented. MAIN BODY: This study presents the first systematic review and meta-analysis of the association between mindfulness and sustainable dietary behaviors. Twelve articles with 13 studies, published between 2018 and 2025, were included, covering both observational and interventional designs with sample sizes ranging from 16 to 560 participants. A random-effects model revealed a small but significant positive association between mindfulness and sustainable diets (d = 0.28, 95% CI [0.13, 0.43]). Subgroup analyses indicated that observational studies tended to show stronger effects than interventions, and that cultural context played an important moderating role, with larger effects reported in Asian samples. Facet-specific analyses suggested that certain mindfulness dimensions - particularly observing, describing, and non-reactivity - were more strongly associated with sustainable food choices, while others showed weaker or non-significant links. A focused analysis on meat consumption revealed that mindfulness was positively associated with reduced meat intake or vegetarian dietary intentions (d = 0.25). Scenario modelling using national Life Cycle Assessment data from Germany, the UK, and the U.S. demonstrated that the potential reductions in meat intake could translate into substantial CO2 emission savings. A full adoption of the EAT-Lancet diet was projected to result in cumulative savings of up to 11% of Germany's total emissions budget by 2040, whereas incorporating the population's willingness to reduce meat intake yielded more conservative cumulative savings of 2.8% over the same period. CONCLUSION: This meta-analysis provides systematic evidence that mindfulness is positively associated with more sustainable dietary behavior and reduced meat consumption. While effects are modest and heterogeneous, the findings suggest that mindfulness could represent a promising lever for promoting dietary change aligned with planetary health goals. Future research should further investigate cultural differences, the role of specific mindfulness facets, and the feasibility of scaling mindfulness-based interventions to support both individual well-being and climate targets.

8 July 2026

Read appraisal →
observationalEvidence: Weak
55CEBM

BMJ open

The Cyber Paranoia and Fear Scale-Updated (CPFS-U): development and implications for digital health engagement

OBJECTIVES: To update and revalidate the Cyber Paranoia and Fear Scale to reflect current technological contexts and examine its relevance to digital health readiness and engagement. METHODS: Using an online community sample (n=433), exploratory factor analysis was conducted to examine the factor structure of the revised item pool. Items were refined through consultation with Patient and Public Involvement and Engagement groups to ensure contemporary relevance and clarity. RESULTS: Analysis supported a four-factor structure representing artificial intelligence (AI) and digital dependence, technological risk awareness, perceived data vulnerability and surveillance-related mistrust. The updated Cyber Paranoia and Fear Scale-Updated (CPFS-U) demonstrated good internal consistency and supported construct validity. Cyber-paranoia and fear were conceptually and empirically distinct from general paranoia and anxiety, highlighting the specific cognitive and emotional responses elicited by digital technologies. CONCLUSIONS: The CPFS-U offers a psychometrically robust, modernised measure for understanding individuals' responses to digital and AI-based technologies. Its application in digital health research and practice can inform risk communication, user engagement strategies and the design of trustworthy digital interventions. By identifying individuals who may disengage due to online mistrust, the CPFS-U has the potential to inform more inclusive and psychologically informed digital health systems.

8 July 2026

Read appraisal →
otherEvidence: Weak
55CEBM

Proceedings of the National Academy of Sciences of the United States of America

Analytical modeling for suction cup designs for skin-interfaced wearable devices

Stable mounting is a central requirement for skin-interfaced wearable biomedical devices, because accurate and long-term measurements with clinical utility typically demand intimate contact with the skin, whereas practical use also requires gentle removal to minimize skin irritation and damage. Existing mounting strategies often struggle to satisfy these competing requirements simultaneously, especially under prolonged wear or in the presence of sweat and moisture. Suction-based mounting has recently emerged as a promising alternative because it can provide strong, reversible, and adhesive-free attachment, yet its underlying mechanics remain insufficiently understood. Here, we establish analytical models for the deformation and force of suction cups in a fully explicit form, covering both the cone suction cup and an optimized ring suction cup design. Unlike previous approaches that rely on indirect quantities such as the pressure difference and contact radius, which are not available before experiments and therefore cannot serve as controllable design variables, the present framework yields direct relations between suction performance and geometry parameters, material properties, and loading conditions, including the maximum push down displacement and the subsequent pull up displacement. The resulting formulas agree closely with accurate numerical solutions and lead to compact scaling laws that clearly identify how geometry and material parameters govern suction performance. These results provide a quantitative and physically transparent foundation for the design of suction-based mounting strategies in wearable devices.

7 July 2026

Read appraisal →
observationalEvidence: Weak
70CEBM

Supportive care in cancer : official journal of the Multinational Association of Supportive Care in Cancer

Patients' expectations and barriers towards postoperative telehealth follow-up after oncologic breast surgery in France

BACKGROUND: Surgery is the most frequent treatment for breast cancer and requires postoperative follow-up to detect complications and monitor patient recovery. Postoperative telemonitoring has demonstrated benefits in improving rehabilitation and psychological well-being. However, patient expectations of and barriers to telehealth follow-up after oncologic breast surgery remain insufficiently explored. METHODS: A prospective survey was conducted among adult patients attending pre- or postoperative follow-up appointments for breast cancer in the Department of Breast and Reconstructive Surgery at the Civil Hospitals of Colmar between March and August 2025. The questionnaire comprised four sections: demographic data, treatments received, telemedicine accessibility and acceptability, and patient expectations of and barriers to telemonitoring. RESULTS: A total of 124 questionnaires were analyzed. Results demonstrated high accessibility to digital tools among participants. Although patients recognized several advantages of telemonitoring, most expressed a strong preference for in-person follow-up visits. The most expected features included the ability to report postoperative complications and to monitor recovery and pain. The main barrier was the fear of reduced human contact. Retired status was associated with lower acceptance of postoperative telemonitoring. CONCLUSION: These findings support the implementation of a hybrid follow-up model combining early postoperative telemonitoring with later in-person consultations, which could help reassure patients and shorten hospital stays while maintaining personalized care. The monocentric design may limit the external validity of the findings. Future studies should explore patients' perceptions following repeated use of telemonitoring tools and incorporate healthcare professionals' perspectives on their integration into clinical practice.

7 July 2026

Read appraisal →
Systematic ReviewEvidence: Moderate
65CEBM

Journal of medical Internet research

Patients' Experiences of Nurse-Led eHealth Interventions for Chronic Heart Failure: Qualitative Systematic Review and Meta-Synthesis

BACKGROUND: Chronic heart failure (CHF) is a major chronic condition in the context of global population aging and is associated with high prevalence, frequent rehospitalization, and high mortality. Self-management is widely recognized as an important factor influencing prognosis and quality of life among patients with CHF. Nurse-led eHealth interventions have been increasingly used in postdischarge CHF management and have shown potential for improving adherence and health-related behaviors. However, existing studies have predominantly focused on objective outcomes, and limited systematic synthesis has examined patients' subjective experiences, which may restrict the patient-centered optimization and sustainable implementation of these interventions. OBJECTIVE: This study aimed to systematically synthesize patients' experiences of nurse-led eHealth interventions for CHF, identify facilitators and barriers to engagement and implementation, and inform the patient-centered optimization of intervention design through a qualitative systematic review and meta-synthesis. METHODS: This qualitative systematic review and meta-synthesis was conducted in accordance with the Joanna Briggs Institute methodology for qualitative systematic reviews. A comprehensive search was performed in PubMed, Web of Science, Embase, the Cochrane Library, CINAHL, CNKI, WanFang, and VIP from database inception to April 30, 2026, supplemented by manual searches of the reference lists of the included studies. Following methodological quality appraisal, the findings of the included studies were extracted and synthesized using thematic synthesis. RESULTS: A total of 23 studies involving 424 patients with CHF were included, comprising 17 qualitative studies and 6 mixed methods studies. Four synthesized themes and 12 subthemes were generated: (1) patient empowerment and enhanced self-management, (2) sense of security and continuity of care under professional support, (3) variations in acceptance and emotional responses, and (4) barriers and challenges in implementing eHealth interventions. CONCLUSIONS: Nurse-led eHealth interventions may provide important support for CHF management from the patient perspective by promoting empowerment and extending care support beyond hospital settings. The findings suggest that future intervention design should better address patient heterogeneity, technological usability, and long-term sustainability. Further efforts are needed to strengthen individualized adaptation and technological optimization to enhance the accessibility, acceptability, and patient-centered implementation of these interventions.

7 July 2026

Read appraisal →
Systematic ReviewEvidence: Weak
50CEBM

Food & function

Apples and apple-based products in the modulation of cardiometabolic and functional markers: a systematic review of human intervention studies.

Apple is one of the most widely consumed fruits worldwide, yet its effects on human health remain the subject of ongoing debate. The aim of the present review was to examine evidence from human intervention studies evaluating the impact of apples and apple-based products on functional and metabolic health markers. A total of 38 studies were included: 13 postprandial interventions, 22 medium- or long-term interventions, and 3 assessing both postprandial and chronic effects. Postprandial studies predominantly investigated the effects of apple consumption on blood glucose levels and plasma antioxidant capacity, whereas medium- or long-term interventions assessed a broader range of biomarkers related to cardiometabolic health, oxidative stress, vascular function, inflammation, and gut function. Overall, the findings suggest that apples and apple-based products may beneficially modulate glycaemia, antioxidant capacity, and vascular endothelial function mainly in short-term interventions, while medium/long-term studies reported an apparent improvement in gut microbiota composition. However, the current evidence remains insufficient to draw definitive conclusions. Additionally, substantial heterogeneity in study design, populations, and products tested limits the ability to generalize results. Nonetheless, apple consumption, consistent with fruit intake in general, represents an important component of a healthy and balanced diet, providing valuable nutrients and bioactive compounds whose intake should be encouraged. Therefore, further well-designed intervention studies, particularly in populations with cardiometabolic risk factors, are warranted to better clarify the role of apples and apple-derived products in human health. Future research should also aim to identify the effective amounts and specific bioactive components responsible for the observed effects, as well as to determine whether these benefits may vary according to the health status of the target population.

6 July 2026

Read appraisal →
observationalEvidence: Moderate
75CEBM

Journal of medical Internet research

Attitudes and Needs of Health Care Providers Toward Artificial Intelligence-Assisted Pediatric Palliative Care: Mixed Methods Study

BACKGROUND: While artificial intelligence's (AI's) transformative potential in health care is widely acknowledged, its application in highly sensitive, humanistic domains like pediatric palliative care (PPC) remains largely unexplored. OBJECTIVE: This study aims to explore the attitudes and needs of health care providers on the PPC assisted by AI, with the goal of informing future development and implementation of AI systems in this field. METHODS: This was an explanatory sequential mixed methods study consisting of a nationwide cross-sectional questionnaire survey (March-April 2025) followed by qualitative semistructured interviews (August-October 2025). The quantitative study aimed to investigate PPC health care providers' experiences, attitudes, and needs for the application of AI. Participants included team members of all recognized PPC teams in mainland China. The qualitative study aimed to explore in greater depth the potential future roles of AI in this field, as well as the features of an ideal AI-assisted tool for PPC. Potential interviewees were recruited from the pool of quantitative survey respondents. RESULTS: Among 352 survey respondents, most (n=205, 58.24%) reported moderate familiarity with AI, with large language models being the most commonly used (n=280, 79.55%). Among large language model users, over half (161/280, 57.50%) reported using them for clinical purposes. Attitudes were generally positive: 67.05% (236/352) believed AI's benefits would outweigh drawbacks, and 75% (264/352) considered its implementation feasible. The most desired applications were patient and family education (276/352, 78.41%) and symptom management (257/352, 73.01%). Interviews with 17 providers revealed three themes: (1) clinical roles and boundaries, (2) elements for clinical integration, and (3) challenges in development and deployment. CONCLUSIONS: This study reveals that PPC providers express positive attitudes and strong demand for AI-assisted clinical work. Furthermore, the research clarifies appropriate roles for AI, outlines elements for clinical integration, and highlights potential challenges in development and integration. This study provides evidence for the feasibility of AI application in PPC and offers guidance for the future development and deployment of AI tools.

5 July 2026

Read appraisal →
Systematic ReviewEvidence: Weak
45CEBM

Supportive care in cancer : official journal of the Multinational Association of Supportive Care in Cancer

Interventions to improve the psychosocial outcomes of individuals with lymphedema: a systematic review

PURPOSE: Lymphedema is a common and enduring consequence of cancer treatment with substantial physical and psychosocial morbidity, yet research predominantly focuses on physical outcomes. This systematic review examines psychosocial interventions for individuals with lymphedema, describing the intervention characteristics, methodological quality, and strength of the evidence. METHODS: A systematic literature search was conducted of studies published through November 2025. Empirical interventions with psychosocial components and outcomes for individuals with lymphedema were eligible. Methodological quality was assessed. Strength of the evidence was determined by combining quality with intervention effectiveness. RESULTS: Ten studies met inclusion criteria, including mind-body (n = 3), education and support (n = 3), and physical activity (n = 4) intervention models, and most delivered following completion of decongestive therapy. Psychosocial outcomes included mental health (n = 8) and quality of life (n = 6). Methodological rigor was mixed, with MQRS scores ranging from 5-13 (of 14). Interventions delivered in combined in-clinic and remote settings demonstrated stronger evidence for improving mental health outcomes than either setting alone. CONCLUSION: Few empirical evaluations exist targeting the psychosocial needs of individuals with lymphedema, particularly during intensive phases of decongestive therapy. Heterogeneity in intervention models, outcome measures, and treatment timing limits cross-study comparisons; however, hybrid delivery approaches show promise for improving psychosocial outcomes. Greater integration of psychosocial care across the lymphatic care continuum, supported by standardized outcome measures, may strengthen supportive care for individuals living with lymphedema.

5 July 2026

Read appraisal →
Systematic ReviewEvidence: Weak
55CEBM

PloS one

Delayed health-seeking behavior and its associated factors among cancer patients in Ethiopia: A systematic review and meta-analysis, 2025

BACKGROUND: Delayed Health-seeking behavior among cancer patients is a major contributor to late diagnosis, poor prognosis, and high mortality, particularly in low-resource settings like Ethiopia. However, evidence on the magnitude and determinants of delayed care-seeking remains fragmented. OBJECTIVE: This systematic review and meta-analysis aimed to estimate the pooled prevalence of delayed Health-seeking behavior among cancer patients in Ethiopia and to identify associated factors influencing delays. METHODS: This study employed a systematic review and meta-analysis design to assess delayed Health-seeking behavior and its influencing factors among cancer patients in Ethiopia. A systematic search was conducted in PubMed, Scopus, Web of Science, CINAHL, AJOL, Google Scholar, and Ethiopian University repositories until April 27, 2025. The data were extracted from March 10-20 and analyzed from March 21-30, with report generation till April 27, 2025, using R software. Meta-analysis was performed using a random-effects model, with forest plots illustrating pooled prevalence and associated factors. Heterogeneity was assessed using the I² statistic, and study quality was evaluated using a validated tool. RESULTS: Seven studies conducted across multiple regions of Ethiopia were included in the final analysis with a total of 2,641 participants. The pooled prevalence of delayed Health-seeking behavior among cancer patients was 54% (95% CI: 39%-68%). Meta-analysis of associated factors showed that rural residence was significantly associated with delayed Health-seeking behavior, with patients residing in rural areas having more than threefold higher odds of delay (AOR = 3; 95% CI: 1.81-4.19), poor knowledge about cancer was strongly associated with delay, with nearly seven times higher odds among patients with poor knowledge compared to those with adequate knowledge (AOR = 6.63; 95% CI: 2.21-11.05), lack of cancer awareness was also a significant predictor of delayed Health-seeking behavior (AOR = 2.63; 95% CI: 1.75-3.51), and patients without pain were over three times more likely to delay Healthcare(AOR = 3.38; 95% CI: 2.44-4.67) were factors associated with delayed Health-seeking behavior. CONCLUSIONS: Our review showed that half of the cancer patients in Ethiopia experienced delayed health-seeking behavior. Delayed care-seeking was associated with rural residence, poor knowledge, limited awareness of cancer, and absence of pain symptoms. Targeted interventions, including public awareness campaigns, expansion of healthcare services in rural areas, and financial support initiatives, are urgently needed to reduce delays and improve early cancer diagnosis and outcomes. PROSPERO REGISTRATION NUMBER: CRD420251037845.

4 July 2026

Read appraisal →
otherEvidence: Moderate
55CEBM

Philosophy, ethics, and humanities in medicine : PEHM

Moral diversity and the challenge of responsibility in AI-CDSS

The increasing integration of artificial intelligence in clinical decision support systems (AI-CDSS) has fueled expectations of more personalized and effective diagnostics and therapies. By incorporating machine learning methods, AI-CDSS promise enhanced predictive accuracy, improved stratification, and innovative individualized care. However, this technological optimism is accompanied by complex ethical challenges, including issues of explainability, trust, autonomy, and data security. At the core of these debates lies the question of responsibility, which involves both its attribution and diffusion, as well as the underlying normative standards guiding moral action. In the context of healthcare practice, responsibility is further complicated by moral diversity-the coexistence of varying moral values, cultural beliefs, and ethical frameworks among healthcare professionals, patients, and institutional stakeholders. This plurality challenges the establishment of a unified normative standard necessary for ethically sound responsibility attribution. This paper offers an analysis of moral diversity and AI-CDSS as a challenge for responsibility in healthcare environments. Using a relational concept of responsibility the study examines key areas in which moral diversity affects responsibility in AI-mediated decision-making. This includes algorithmic bias, healthcare professional and patient interaction and the role of patients. Through these examples, the paper explains how different normative standards intensify ethical complexity in AI-supported clinical contexts. It argues that greater ethical sensitivity to moral diversity is essential-both in the development of AI-CDSS and in their application within morally value-laden healthcare situations.

4 July 2026

Read appraisal →
diagnosticEvidence: Weak
45CEBM

BMJ health & care informatics

Detection of cancer recurrence from Thai-English electronic medical records using sentence embeddings.

OBJECTIVE: This study developed and validated monolingual and bilingual sentence-bidirectional encoder representations from transformers (SBERT) models for detecting cancer recurrence within Thai-English electronic medical records (EMRs) from Thai cancer hospitals. METHOD: A multicentre dataset of 32 436 documents from 1250 patients was used for model development. External validation involved an independent dataset of 9244 documents from 384 patients across two Thai cancer hospitals. Performance was benchmarked against a fine-tuned PubMedBERT (MetBERT). RESULTS: The development dataset included breast (43.9%), colorectal (12.1%), cervical (28.0%) and head and neck (16.0%) cancers. MetBERT achieved the highest area under the precision-recall curve (AUPRC) for locoregional versus no recurrence (11.1%) and locoregional versus distant recurrence (91.7%), while monolingual-SBERT excelled at distant versus no recurrence (32.0%). External validation demonstrated MetBERT superiority for locoregional versus no recurrence (9.30%-21.50%). For distant versus no recurrence, bilingual-SBERT performed best with AUPRC 17.55%-24.39%. While MetBERT led in distinguishing locoregional versus distant recurrence (88.30%-94.70%), bilingual-SBERT demonstrated robust external validation performance (AUPRC 85.25%-91.80%). DISCUSSION: Low AUPRC values (9%-32%) reflect the extreme class imbalance in real-world data (~1% recurrence prevalence). Despite this, fine-tuned MetBERT achieved highest performance, while bilingual-SBERT demonstrated superior robustness during external validation. This validates sentence embedding models for handling mixed Thai-English medical records in multilingual clinical environments. CONCLUSION: Sentence embedding frameworks provide a practical, generalisable solution for detecting cancer recurrence within multilingual EMRs. Despite text-length constraints, these models are suitable for clinical integration as a screening tool for cancer registry workflows.

4 July 2026

Read appraisal →
otherEvidence: Weak
25CEBM

BMJ health & care informatics

Predicting health and disease: a conceptual framework for AI in preventive and precision medicine

The conventional medical approach of treating symptoms as they appear with restricted screening often limits intervention to slowing disease progression rather than fully reversing it. A new approach leveraging artificial intelligence (AI) and computational technologies across expanding multimodal biomedical datasets holds the promise to enable predicting actionable future changes in health before symptom onset. This article presents a conceptual framework for a preventive paradigm in precision medicine and healthcare, integrating recent advancements in AI and biomedical datasets. Key remaining challenges facing computational systems, real-world clinical validation and implementation, and preventive interventions together with recommendations and prioritised future directions are highlighted. Such an approach could pave the way for more proactive and preventive medical interventions to effectively address the growing burden of chronic disease.

4 July 2026

Read appraisal →
Systematic ReviewEvidence: Insufficient
5CEBM

Advanced materials (Deerfield Beach, Fla.)

Advances in Wearable Bioimaging

The development of wearable bioimaging platforms has accelerated in the era of precision medicine, addressing the limitations of conventional, facility-based systems, which can be costly and difficult to access. In this review, we first outline the operating principles of major imaging modalities, with an emphasis on target tissues and factors that contribute to image quality. Next, we survey advances in miniaturizing these modalities into wearable devices, from electrical impedance tomography belts to flexible ultrasound and photoacoustic patches, highlighting materials selection and fabrication strategies that enhance image quality, skin adhesion, and functional integration. Lastly, we discuss outstanding technical and translational challenges that must be addressed to realize the full potential of these platforms, including safety and long-term use. With sustained research and responsible commercialization, these wearable devices could reshape continuous physiological monitoring and enable more proactive, preventive health management.

3 July 2026

Read appraisal →
otherEvidence: Moderate
70CEBM

Behavior research methods

A validity-guided workflow for robust large language model research in psychology.

Large language models (LLMs) are rapidly being integrated into psychological and behavioral research as research tools, evaluation targets, human simulators, and cognitive models. Yet recent evidence reveals severe measurement unreliability: personality assessments degenerate under factor analysis, moral preferences reverse with punctuation changes, and theory-of-mind accuracy varies widely with trivial rephrasing. These "measurement phantoms"-statistical artifacts masquerading as psychological phenomena-threaten the validity of a growing body of research. Guided by the dual-validity framework that integrates psychometrics with causal inference, this article presents a six-stage workflow that scales validity requirements to research ambition-using LLMs to code text requires basic reliability and accuracy, whereas claims about psychological properties demand comprehensive construct validation. Researchers must (1) explicitly define their research goal and corresponding validity requirements, (2) develop and validate computational instruments through psychometric testing, (3) design experiments that control for computational confounds, (4) execute protocols transparently, (5) analyze data with methods appropriate for nonindependent observations, and (6) report findings within boundaries and use results to refine theory. The workflow is illustrated through an example of model evaluation-"LLM selfhood"-showing how systematic validation can distinguish genuine computational phenomena from measurement artifacts. By establishing validated computational instruments and transparent practices, this workflow provides a path toward building a robust empirical foundation for AI psychology research.

3 July 2026

Read appraisal →
otherEvidence: Weak
25CEBM

Journal of medical Internet research

Transformation Versus Innovation in Digital Health Care and the Future of Clinical AI

Digital innovation is frequently presented as the key to transforming health care. In this News and Perspectives article, JMIR Correspondent and academic physician Boon-How Chew reports on the distinction between and the direction of transformation and innovation, reflecting on the future of clinical AI and what lasting change in health care ultimately requires.

3 July 2026

Read appraisal →
Systematic ReviewEvidence: Weak
10CEBM

The Nigerian postgraduate medical journal

The Role of Digital Therapeutics and Artificial Intelligence in Chronic Disease Management: A Narrative Review

We conducted research showing that chronic disease management continued to challenge healthcare systems, payers, and patients. At the same time, digital therapeutics emerged as a promising and potentially transformative approach. They supported better management of long-term conditions, improved patient outcomes, and helped streamline healthcare delivery. We carried out a comprehensive literature search to identify both original reports and reviewed publications. The search covered multiple databases, including Google Scholar and PubMed, and we also gathered relevant information from credible online sources such as the World Health Organization and India's National Crime Records Bureau. Using these findings, the narrative explained how digital therapeutics reshaped chronic disease management by highlighting key benefits, the technologies that enabled these solutions, and the expected impacts on both clinical and economic outcomes. We also discussed the obstacles the sector encountered as it developed, and we considered the future prospects of digital therapeutics in chronic care. In conclusion, digital therapeutics offered personalised care, improved patient engagement and adherence, provided real-time monitoring and feedback, enhanced accessibility and convenience, and were cost-effective in managing chronic disease; however, challenges such as cybersecurity concerns, reliability of data, the digital divide, and a lack of extensive clinical validation needed to be addressed for widespread adoption. While the evidence to date suggested clear clinical and economic promise, realizing that promise required coordinated action stronger clinical trials to build robust evidence, clear regulatory pathways to ensure safety and efficacy, investment in secure interoperable infrastructure, and targeted efforts to close the digital divide so vulnerable populations were not left behind; policymakers, clinicians, payers, and technology developers had to collaborate to translate innovation into equitable, scalable care improvements.

3 July 2026

Read appraisal →
otherEvidence: Weak
50CEBM

Magnetic resonance in medical sciences : MRMS : an official journal of Japan Society of Magnetic Resonance in Medicine

Integrating Artificial Intelligence into Prostate MR Imaging: Technical Foundations, Clinical Applications, and Workflow Implications

Prostate MRI has become a cornerstone of contemporary prostate cancer diagnosis, enabling improved detection of clinically significant disease while reducing unnecessary biopsies and overtreatment. However, prostate MRI remains technically demanding, time-consuming, and subject to inter-reader variability, particularly as healthcare systems move toward abbreviated protocols such as non-contrast MRI (biparametric MRI). In this context, artificial intelligence (AI) has emerged as a promising tool to enhance image quality, diagnostic consistency, and workflow efficiency across the prostate MRI pathway. This non-systematic narrative review provides a comprehensive overview of the technical foundations, clinical applications, and workflow implications of AI integration into prostate MRI. It summarizes key concepts in machine learning and deep learning relevant to prostate imaging and reviews current evidence supporting AI-based solutions for image quality assessment and reconstruction, automated prostate segmentation, lesion detection, and risk stratification. Particular attention is given to human-AI collaboration models, the role of AI in supporting equivocal lesions, and the integration of imaging with clinical variables for personalized risk estimation. In addition, it discusses the impact of AI on reporting efficiency, training, and standardization, as well as the current landscape of commercially available AI tools. Despite encouraging results from large multicenter studies, important challenges remain, including heterogeneity in study design, limited prospective validation, generalizability across institutions, and ethical and regulatory considerations. Overall, AI should be regarded as a complementary decision-support technology rather than a replacement for radiologists. Thoughtful implementation, robust validation, and appropriate user training are essential to ensure that AI meaningfully enhances the quality, efficiency, and reliability of prostate MRI-based care.

2 July 2026

Read appraisal →
otherEvidence: Weak
25CEBM

Journal of medical Internet research

How Does That Large Language Model Make You Feel?

People are increasingly turning to commercially available large language models (LLMs) for emotional support. In this News and Perspectives article, JMIR Correspondent Simon Spichak reports on the role of LLMs in mental health, speaking with experts about safety concerns, research gaps, and next steps.

2 July 2026

Read appraisal →
Randomised Controlled TrialEvidence: Weak
40CEBM

Journal of the American Medical Informatics Association : JAMIA

TRUST: an large language model-based dialogue system for trauma understanding and structured assessments

OBJECTIVES: While large language models (LLMs) have been widely used to assist clinicians and support patients, no existing work has explored dialogue systems for standard diagnostic interviews and assessments. This study aims to bridge the gap in mental healthcare accessibility by developing an LLM-powered dialogue system that replicates clinician behavior. MATERIALS AND METHODS: We introduce TRUST, a framework of cooperative LLM modules capable of conducting formal diagnostic interviews and assessments for post-traumatic stress disorder (PTSD) following the Clinician-Administered PTSD Scale for DSM-5 (CAPS-5). To guide the generation of appropriate clinical responses, we propose a Dialogue Acts schema specifically designed for clinical interviews. Additionally, we develop a patient simulation approach based on real-life interview transcripts to replace time-consuming and costly manual testing by clinicians. RESULTS: A comprehensive set of evaluation metrics is designed to assess the dialogue system from both the agent and patient simulation perspectives. Expert evaluations by conversation and clinical specialists show that TRUST performs comparably to real-life clinical interviews. DISCUSSION: Our system performs with clinical quality approaching that of human clinicians, with room for future enhancements in communication styles and response appropriateness. CONCLUSIONS: Our TRUST framework shows its potential to facilitate mental healthcare availability.

2 July 2026

Read appraisal →
Systematic ReviewEvidence: Weak
20CEBM

Journal of diabetes science and technology

Artificial Intelligence for Diabetic Foot Screening Based on Digital Image Analysis: A Systematic Review

INTRODUCTION: Early detection of diabetic foot complications is essential for effective management and prevention of complications. Artificial intelligence (AI) technology based on digital image analysis offers a promising noninvasive method for diabetic foot screening. This systematic review aims to identify a study on the development of an AI model for diabetic foot screening using digital image analysis. METHODOLOGY: The review scrutinized articles published between 2018 and 2023, sourced from PubMed, ProQuest, and ScienceDirect. The keyword-based search resulted in 2214 relevant articles and nine articles that met the inclusion criteria. The article quality assessment was done through Quality Assessment of Diagnostic Accuracy Studies (QUADAS). Data were extracted and analyzed using NVivo. RESULTS: Thermal imagery or foot thermogram was the main data source, with plantar temperature distribution patterns as an important indicator. Deep learning methods, specifically artificial neural networks (ANNs) and convolutional neural networks (CNNs), are the most commonly used methods. The highest performance is demonstrated by the ANN model with MATLAB's Image Processing Toolbox that is able to classify each type of macula with 97.5% accuracy. The findings show the great potential of AI in improving the accuracy and efficiency of diabetic foot screening. CONCLUSION: This research provides important insights into the development of AI in digital image-based diabetic foot screening. Future studies need to focus on evaluating clinical applicability, including ethical aspects and patient data security, as well as developing more comprehensive data sets.

2 July 2026

Read appraisal →
Systematic ReviewEvidence: Weak
45CEBM

International journal of medical informatics

The role of artificial intelligence in virtual emergency care: a systematic review

BACKGROUND: The integration of artificial intelligence (AI) into virtual emergency care represents a potentially transformative approach to healthcare delivery, yet the evidence base remains poorly characterized. This systematic review comprehensively evaluates the current state of AI applications in virtual emergency care settings. METHODS: We systematically searched eight databases (Embase, PsycINFO, MEDLINE, PubMed, Scopus, Web of Science, CINAHL, Cochrane Library) from inception through March 2025. Of 7,098 records identified and 4,935 screened after deduplication using Covidence, 8 studies met inclusion criteria following exclusion of one study lacking AI components. Studies were assessed using PROBAST + AI for risk of bias and quality assessment, TRIPOD + AI for reporting quality, and GRADE for certainty of evidence. RESULTS: The eight included studies (total participants: approximately 0.5 million) evaluated diverse AI applications including decision trees, machine learning ensembles, and graph neural networks across multiple virtual emergency contexts. Performance varied widely (accuracy 77.5-100%, sensitivity 63-100%, specificity 60% in single study reporting). All clinical studies demonstrated serious risk of bias. TRIPOD + AI compliance averaged only 36.9% (range 30.9-48.1%). GRADE assessment revealed very low to low certainty evidence across all outcomes, with no studies measuring actual clinical outcomes. CONCLUSIONS: Current evidence is insufficient to support widespread clinical implementation of AI in virtual emergency care. While preliminary results suggest potential benefits in triage accuracy and resource efficiency, critical gaps exist in validation, clinical outcome assessment, and reporting standards. Future research must prioritize prospective controlled trials with real patient data, clinical outcome measurements, and adherence to reporting guidelines.

2 July 2026

Read appraisal →
otherEvidence: Weak
45CEBM

Physiology (Bethesda, Md.)

Wearable Sensing for Clinical Physiology Monitoring: Emerging Paradigms

Advances in wearable sensing technology are driving a new era of personalized health monitoring. In contrast to the hard, rigid form factors of conventional wearable sensors, these emerging skin-interfaced systems support high-quality physiological measurements across biophysical, biochemical, and kinematic signals of interest. These platforms enable continuous monitoring of complex physiological processes with unprecedented detail as a result of a seamless, conformal skin interface and advanced wireless communications capabilities. These platforms integrate flexible materials, miniaturized electronics, and wireless communication to provide detailed physiological data during daily activities. This review examines how skin-interfaced wearables are advancing patient care, remote monitoring, and large-scale health studies. We highlight critical barriers to clinical adoption including interpreting data, validating devices, and integration into health care systems. Key opportunities include sustainable manufacturing, point-of-care fabrication, and development of disease-specific digital biomarkers. By addressing these challenges through collaboration among engineers, clinicians, and data scientists, wearable sensors can expand patient access to advanced physiological monitoring and transform personalized medicine.

2 July 2026

Read appraisal →
otherEvidence: Weak
70CEBM

Advances in skin & wound care

Mobile Health App Needs Among Patients With Diabetic Foot Ulcers in China: A Qualitative Study From the Perceptions of Patients, Caregivers, and Health Care Professionals

OBJECTIVE: To explore the specific needs of patients with diabetic foot ulcers (DFUs), caregivers, and health care professionals (HCPs) for a mobile health (mHealth) app, aiming to inform the design and development of effective mHealth service solutions. METHODS: This descriptive qualitative study was conducted from June to September 2024 in the wound care clinics of 2 local hospitals. Participants included patients with DFUs, caregivers, and HCPs directly involved in their care. Interview data were analyzed, synthesized, and refined using content analysis. RESULTS: Five key themes emerged: the pressing need to implement mHealth app services, convenient and personalized access to information, continuous and specialized health guidance, a multidisciplinary approach to disease management, and free access alongside privacy and legal protections. CONCLUSIONS: This study provides valuable insights for the design and development of an mHealth app for DFU. During the development process, it is essential to consider user needs and ensure the app meets the expectations of patients and related groups for personalized, continuous, and specialized health guidance; free access; privacy protection; and multidisciplinary collaboration. A balance should be struck between convenience and security of the app's features, to encourage user engagement and enhance the app's value and effectiveness.

2 July 2026

Read appraisal →
Systematic ReviewEvidence: Weak
10CEBM

International journal of gynaecology and obstetrics: the official organ of the International Federation of Gynaecology and Obstetrics

Integrated artificial intelligence and omics for prediction and monitoring of pre-eclampsia

Pre-eclampsia is a difficult pregnancy condition that causes high blood pressure and can lead to health complications in both mother and newborn, resulting in a higher fatality rate. It presents with a wide range of symptoms and lacks specific indicators, as the contemporary diagnostic techniques, including proteinuria testing and blood pressure measurements, are not reliable. The current evolution in artificial intelligence (AI) technology tends to show a promising transformation of pre-eclampsia management. AI algorithms are applied to process larger sets of clinical, biochemical, and image data that facilitate timely medical interventions by bringing up the early-onset and severity of pre-eclampsia. By analyzing the red cell distribution width (blood test indicators for pre-eclampsia), it is recognized as a cost-effective way of detecting inflammation. The application of AI technology on non-invasive diagnostic (wearable) devices enables continuous monitoring with imaging techniques for the placenta and retina via cloud-based systems. These developments are not only applied for early detection of pre-eclampsia, but also assist decision making capabilities in both high- and low-resource environments. This article explains how the growing use of AI is changing the way that pre-eclampsia is understood and managed, with the aim of improving accuracy and offering more personalized care for pregnant women.

1 July 2026

Read appraisal →
otherEvidence: Weak
40CEBM

Joint Commission journal on quality and patient safety

Large Language Models as Physician Recommenders: Limitations, Alternatives, and the Path Toward Hybrid AI Systems

Digital tools are increasingly used by patients to access health information and navigate care, including choosing clinicians, but the evidence supporting "best doctor" recommendations varies widely across available methods. This manuscript compares the emerging use of consumer-facing large language models (LLMs) for this purpose with four alternative physician-selection strategies: patient networks, expert referrals, consumer rating/search platforms, and outcome-informed artificial intelligence / machine learning (AI/ML) recommender systems. Consumer rating platforms mainly reflect patient experiences rather than technical skills and often show weak or inconsistent links with objective performance measures. General-purpose LLMs may produce convincing rationales that highlight visibility and reputation, but they remain limited by hallucinations, prompt sensitivity, and uncertainty about whether their explanations truly reflect the underlying reasoning. Outcome-informed AI/ML recommender systems could, in theory, incorporate risk-adjusted clinical performance, but they face significant challenges, including unreliable physician-level measurement, incomplete data, residual confounding, and limited transparency, which can undermine trust and informed consent. We suggest that a hybrid system should be seen as a guiding concept and a testable approach rather than a final solution. Such a system could integrate proven, auditable performance measures with transparent patient-preference filters and verifiable, patient-facing explanations. We present key design principles, including clinical relevance, explainability, autonomy, equity, and governance, in line with current US Food and Drug Administration (FDA) guidance related to clinical decision support and adaptive AI life cycle management.

1 July 2026

Read appraisal →
otherEvidence: Moderate
70CEBM

Journal of medical Internet research

Operationalizing Digital Health Equity in Artificial Intelligence-Enabled Patient Decision Aids for Older Adults: Mixed Methods Study

BACKGROUND: Artificial intelligence-enabled patient decision aids (AI-PDAs) hold promise for supporting older adults with chronic diseases in accessing personalized health information, clarifying preferences, and engaging in shared decision-making. Achieving equity in their design requires attention to the complex health care and digital contexts in which these tools are used. While the Digital Health Equity Framework (DHEF) provides a conceptual foundation, practical strategies for its application remain limited. OBJECTIVE: This study aimed to identify equity-related determinants and generate actionable design strategies for applying the DHEF to AI-PDAs for older adults. METHODS: A mixed methods study was conducted. Semistructured interviews were conducted with older adults living with hypertension and/or diabetes, health care providers, and medical students to explore equity determinants relevant to AI-PDAs. In parallel, a review of reviews synthesized existing evidence on approaches to addressing these determinants. Interview findings and review findings were integrated through an iterative mapping process conducted by the research team and refined through multidisciplinary expert consultation involving medicine, public health, social services, and computer science. RESULTS: A total of 33 stakeholders were interviewed, including 15 older adults, 8 health care providers, and 10 medical students. Thirteen reviews were included in the umbrella review. The integrated synthesis identified equity determinants spanning individual, interpersonal, community, and societal levels across both the health care and digital environments, together with cross-level concerns related to algorithmic fairness. These findings informed 5 recommendations for equitable AI-PDA development: (1) co-design with end users to address their needs, (2) embrace relationship-centered design, (3) leverage community resources to improve support, (4) promote accessible and equitable artificial intelligence (AI) governance in society, and (5) enhance equitable AI through algorithmic fairness. Together, these recommendations provide practical guidance for design, pilot testing, implementation, and evaluation. CONCLUSIONS: By integrating stakeholder perspectives with synthesized review evidence, this study extends the DHEF from a primarily conceptual framework toward a more practice-oriented approach for AI-PDAs for older adults with chronic disease. Health care settings serve as a mediating sociotechnical context where AI tools may either support or constrain equitable care participation. The findings underscore the need for interdisciplinary collaboration to align technological innovation with equity-oriented design. Future work should focus on co-designed prototypes, real-world testing, and measurable equity outcomes.

30 June 2026

Read appraisal →
otherEvidence: Weak
65CEBM

Journal of medical Internet research

Using a Large Language Model to Support Thematic Analysis of Patient Experiences in Chronic Illness Management: Comparative Qualitative Study

BACKGROUND: Qualitative health research often focuses on how patients experience and manage chronic illnesses, a topic that has been extensively studied in the literature. With the emergence of large language models (LLMs), such as Claude (Anthropic PBC) and ChatGPT (OpenAI), new opportunities are arising to support and scale the thematic analysis of narrative health data. However, their role and added value compared to traditional human-led approaches remain underexplored, particularly in complex clinical contexts such as multimorbidity. OBJECTIVE: We aim to evaluate the methodological contribution of LLM-assisted analysis by examining its ability to replicate and extend established qualitative insights, in comparison with traditional thematic analysis. METHODS: Semistructured interviews were conducted with 30 individuals living with two or more chronic illnesses. Transcripts were analyzed using both manual thematic coding and Claude 3.5 Sonnet. A structured comparison was conducted to identify shared and unique themes across the two approaches. The analysis examined thematic overlap, differences in subtheme identification, and variation in the level of detail between the methods. RESULTS: Both approaches identified similar core themes related to the patient experience, including health care navigation and challenges, support systems and family dynamics, and emotional challenges and coping. Manual analysis produced more contextually detailed interpretations, while the LLM approach identified a larger number of subthemes. Each method also revealed distinct themes: the manual analysis included themes such as faith, caregiving roles, and a proactive mindset, whereas the LLM identified themes such as future planning and multiple health conditions. The findings show both similarities and differences between the two approaches. The LLM analysis also demonstrated efficiency in processing large volumes of qualitative data. CONCLUSIONS: A hybrid approach that integrates artificial intelligence-assisted and human-led thematic analysis can enhance both analytical depth and scalability. These findings support the use of LLMs as a complementary tool in qualitative research, while highlighting the importance of combining automated pattern detection with human interpretation.

30 June 2026

Read appraisal →
observationalEvidence: Weak
55CEBM

JMIR cancer

Large Language Models for Breast and Cervical Cancers Communication: Mixed Methods Evaluation Study Assessing Linguistic Quality, Safety, and Accessibility

BACKGROUND: Effective communication about breast and cervical cancers remains a public health challenge, with widespread misinformation and barriers to cancer-related language understanding. Large language models (LLMs) offer potential for scalable health communication, yet trade-offs between quality, safety, and accessibility of general-purpose and medical-domain LLMs remain underexplored. OBJECTIVE: This study aimed to propose a comprehensive evaluation framework and systematically assess the performance of LLMs in generating breast and cervical cancer information, with a focus on linguistic quality, safety and trustworthiness, and communication accessibility and affectiveness. METHODS: This mixed methods evaluation study assessed outputs from 5 general-purpose and 3 medical LLMs using real-world breast and cervical cancer-related questions curated from publicly available medical datasets. LLM-generated responses were evaluated in a controlled offline setting. Primary outcomes included linguistic quality (fluency, coherence, and accuracy), safety and trustworthiness (toxicity, bias, and harm potential), and communication accessibility and affectiveness (readability, empathy, and clarity). Qualitative ratings were performed by domain experts, while quantitative metrics were compared across models. Statistical analyses included Welch ANOVA to detect differences in metric scores, Games-Howell tests for pairwise comparisons, and Hedges g to assess effect sizes. RESULTS: General-purpose LLMs, particularly Llama 3 and Gemma, demonstrated superior linguistic quality and affectiveness but often produced complex outputs that may limit accessibility. In contrast, medical LLMs (eg, MedAlpaca and BioMistral) generated simpler content suitable for broader audiences but scored lower in safety and empathy due to higher levels of hallucination, bias, and toxicity. CONCLUSIONS: While LLMs show promise for improving digital cancer communication, our findings reveal a trade-off between domain specialization and overall communication quality and safety. Future development of health-focused LLMs should prioritize hybrid modeling strategies to enhance trust, clarity, and clinical relevance in patient-facing tools.

29 June 2026

Read appraisal →
Randomised Controlled TrialEvidence: Weak
45CEBM

Journal of medical Internet research

Effectiveness of a Home-Based and Group-Based Tele-Exercise Program for Breast Cancer Survivors: Pilot Randomized Controlled Trial

BACKGROUND: Breast cancer remains the most prevalent cancer among women globally. Adjuvant therapies can cause adverse effects that compromise physical and mental health. Exercise may mitigate these effects; however, many breast cancer survivors remain insufficiently active. OBJECTIVE: This pilot study aimed to test the effectiveness of a 7-week theory-informed tele-exercise intervention for breast cancer survivors in Hong Kong. METHODS: We developed a 12-week theory-informed tele-exercise intervention for breast cancer survivors in Hong Kong; this pilot tested a 7-week abbreviated version. In this 2-group randomized controlled pilot trial, 34 individuals were assessed for eligibility, and 27 were randomized; 24 completed baseline and were included in the modified intention-to-treat analyses (12 per group). The intervention comprised a progressively supervised group-based tele-exercise intervention transitioning to unsupervised sessions, combined with psychological counseling. Outcomes were guided by the RE-AIM (reach, effectiveness, adoption, implementation, and maintenance) framework, with emphasis on feasibility, acceptability, and preliminary effects. RESULTS: Recruitment was 79.4% (27/34) and baseline-to-post retention was 100% (24/24), with satisfactory attendance (87.7%) and compliance (85.3%) in the intervention group. Acceptability was high. Preliminary signals of improvement were observed in cardiorespiratory fitness (primary outcome), lower-extremity strength, balance, affected-side shoulder range of motion, and health-related quality of life. CONCLUSIONS: This pilot supports the feasibility and acceptability of tele-exercise for breast cancer rehabilitation and provides preliminary evidence to justify a larger trial with longer follow-up to assess sustained effects and broader applicability.

29 June 2026

Read appraisal →
otherEvidence: Insufficient
45CEBM

The Cochrane database of systematic reviews

Views and experiences of weight management for people living with mobility-limiting conditions, intellectual disabilities or severe mental illness: a qualitative evidence synthesis

This is a protocol for a Cochrane Review (qualitative). The objectives are as follows: This qualitative evidence synthesis (QES) aims to address the question: What is known from qualitative evidence about behavioural weight management for people living with mobility-limiting conditions, intellectual disabilities, or severe mental illness? The five review objectives are as follows. To explore individuals' experiences with weight and behavioural weight-management programmes. This includes understanding how individuals living with mobility-limiting conditions, intellectual disabilities, severe mental illness and their carers experience managing their weight and how they perceive and engage with behavioural weight-management programmes, including how accessible, supportive, and relevant they find the programme elements. To identify barriers to and facilitators of managing a healthy weight and accessing behavioural weight-management programmes. Focusing on the unique obstacles encountered by each population group, we will examine issues such as physical accessibility, communication challenges, financial constraints, and gaps in tailored support that may prevent equitable participation. We will also include people's perspectives on structural and system-level barriers, as well as the influence of intersectional factors (e.g. gender and ethnicity), on the challenges of maintaining weight and accessing and benefiting from support. To understand the impact of stigma and discrimination on programme engagement and outcomes. This objective seeks to uncover how stigma related to weight, disability, or mental health influences individuals' willingness to participate, their inclusion and sense of belonging with others within the programmes, and overall programme satisfaction. It also examines the role of stigma as a potential barrier to success. To capture the perspectives of programme providers and healthcare practitioners. Gathering insights from healthcare providers and programme facilitators, this objective assesses which programme elements are considered successful and which present ongoing challenges. This includes evaluating aspects like accessibility, engagement strategies, and support mechanisms that may require improvement. To produce an inclusive piece of research through incorporating the insight of people with lived experience and professional expertise in interpreting and considering the implications of the evidence. By synthesising this qualitative evidence, the review aims to provide a comprehensive understanding of how to design and implement behavioural weight-management programmes that are more inclusive, effective, and responsive to the needs of these populations.

29 June 2026

Read appraisal →
Systematic ReviewEvidence: Weak
40CEBM

European journal of pediatrics

Congenital cytomegalovirus screening by dried blood spot: a systematic review

UNLABELLED: Congenital human cytomegalovirus (cCMV) infection is a leading cause of sensorineural hearing loss and neurological sequelae in children. Although newborn screening strategies remain controversial, dried blood spots (DBS) collected for routine metabolic screening have been proposed as a low-cost method for large-scale detection. This systematic review assessed studies that evaluated the accuracy of DBS testing for screening newborns for cCMV infection. A systematic review was conducted following the Cochrane Handbook for Systematic Reviews of Diagnostic Test Accuracy and reported in accordance with PRISMA-DTA. The PIRT strategy included the following: Population (newborns undergoing screening for cCMV), Index test (detection of cCMV using DBS samples), Reference standard (confirmatory testing using saliva or urine PCR collected within 21 days of life), Target condition (cCMV infection). Searches were performed in PubMed/MEDLINE, Scopus, Bireme, Embase, LILACS, SciELO, and CENTRAL. The search identified 365 articles; after removing 210 duplicates, 155 were screened by title and abstract, and 9 studies were included. The included studies tested between 1174 and 551,034 DBS samples each. Confirmatory testing used PCR in urine (seven studies) or saliva (two studies). Three studies used paired samples for screening; sensitivity ranged from 28.3% to 79.3%, specificity from 99.9% to 100%. Positive predictive value (PPV) ranged from 39 to 100%, and negative predictive value (NPV) from 99.6% to 99.9%. CONCLUSION: Neonatal screening for congenital cytomegalovirus using DBS samples is promising, since it would allow early identification for timely interventions to reduce sequelae. WHAT IS KNOWN: • Congenital CMV is a leading cause of sensorineural hearing loss and neurodevelopmental sequelae, with many cases asymptomatic at birth. Diagnosis relies on early viral detection, and while saliva/urine PCR is sensitive, newborn screening lacks consensus and feasibility is unknown in many settings. WHAT IS NEW: • DBS PCR emerges as a scalable, feasible low-cost screening alternative for cCMV screening. A few states in the USA and provinces in Canada have implemented newborn cCMV screening programs using DBS, but the sensitivity appears lower than in research studies. Improving DNA extraction and PCR methods are required as well as assessing feasibility of using other specimens (e.g., urine or saliva) for screening.

28 June 2026

Read appraisal →
Systematic ReviewEvidence: Weak
40CEBM

Journal of medical Internet research

Applications, Challenges, and Future Directions of Large Language Models in Health Care Communication: Scoping Review

BACKGROUND: Effective health care communication is crucial in the medical field. However, effective communication in clinical practice still faces numerous obstacles, and large language models (LLMs) offer various possibilities for improving the quality of medical communication. To date, there are no published reviews on the use of LLMs in health care communication. OBJECTIVE: This review sought to summarize the applications and challenges of LLMs in health care communication and to identify directions for future research. METHODS: A comprehensive literature search was conducted in PubMed, Embase, Web of Science, and the Cochrane Library from January 2018 to November 2025. The search and selection process followed the PRISMA-ScR (Preferred Reporting Items for Systematic Reviews and Meta-Analyses extension for Scoping Reviews) guideline and the PRISMA-S (Preferred Reporting Items for Systematic Reviews and Meta-Analyses literature search extension) checklist. Eligible studies used LLMs to facilitate health care communication among the public, patients, and clinicians. Following rigorous data extraction and cross-checking, we conducted a quantitative analysis of characteristics of the included literature. Furthermore, using communication accommodation theory as a framework, we identified application patterns of LLMs in health care communication and summarized current challenges and future directions. RESULTS: Ninety-six studies were included in this review, all published between 2023 and 2025, summarizing 4 patterns of LLM application in health care communication: transforming medical information (n=30), facilitating dynamic interaction (n=38), empowering communication capabilities (n=10), and optimizing clinical workflows (n=18). The role of LLMs in health care communication is undergoing a paradigm shift from "static information processing" to "dynamic intelligent interaction." Although they show great promise for practical applications, current evaluation methods and dimensions exhibit significant heterogeneity. Furthermore, LLMs still face multiple challenges in their practical application in health care communication, including technical reliability issues, social trust and adoption, interaction and access barriers, and clinical integration challenges. CONCLUSIONS: Unlike previous studies that merely touched upon the challenges and future directions, this scoping review uses communication accommodation theory to systematically map the application patterns and developmental landscape of LLM-mediated health care communication. Health care communication powered by LLMs holds significant innovation potential and is currently still in the early stages of rapid development. Future research should focus on optimizing model performance, strengthening ethical governance frameworks, enhancing human-machine collaboration models, and ensuring responsible application of LLMs in health care through rigorous empirical validation.

28 June 2026

Read appraisal →
Systematic ReviewEvidence: Weak
55CEBM

Calcified tissue international

Artificial Intelligence Approaches for Osteoporotic Fracture Risk Prediction Using Administrative Health Data: A Systematic Review

To systematically review Machine Learning (ML) fracture risk prediction models developed solely using administrative data, evaluating their development, performance, risk of bias, and concerns regarding applicability. A systematic search was conducted in PubMed, Embase, IEEE, and Web of Science following PRISMA guidelines (up to November 2025). We included studies developing or validating ML models for osteoporotic fracture risk in adults using administrative data without clinical measurements. Risk of bias and concerns for applicability were assessed using the PROBAST tool. Seven studies were included from 3435 initial records. A range of of ML models were utilized, including Random Forests, XGBoost, LASSO-regularized Logistic Regression and Neural Networks. Discriminative performance was moderate-to-good, with best Area Under the Curve (AUC) values ranging from 0.818 for osteoporotic fractures in general to 0.905 for hip fractures only. While five studies showed low risk of bias, five raised applicability concerns due to using specific subpopulations or features with database characteristics. Only two studies used external validation and none directly evaluated clinical utility of the models. ML models using solely administrative data show some promise as scalable, automated tools for fracture risk prediction without requiring manual clinical input. However, clinical implementation is hindered by limited external validation and a lack of formal utility evaluation of the models. Future research should prioritize more rigorous external validation and recalibration of the existing models, more so than the development of novel models, ensuring they are robust, interpretable, and integrable into future clinical workflows.

28 June 2026

Read appraisal →
Systematic ReviewEvidence: Insufficient
5CEBM

ACS sensors

Flexible Wearable Closed-Loop Diagnostic and Therapeutic System: A New Paradigm for Skin Health Management

As the largest external barrier of the human body, the skin is essential to overall health and quality of life. Conventional skin health management faces limitations such as insufficient real-time monitoring and lack of personalization. Flexible wearable devices have recently emerged as a transformative technology, enabling non-invasive sensing and targeted therapy with high skin compatibility. This review comprehensively summarizes recent advances in flexible wearable closed-loop diagnostic and therapeutic systems for skin health management. It outlines the core technological components, including skin bioinformation collection, biosignal monitoring, signal transmission and data analysis, and targeted intervention. Key applications in physiological and pathological monitoring, wound healing, drug delivery, and physical therapy are discussed. Furthermore, the innovation trajectory of closed-loop systems is clarified, key challenges are analyzed, and targeted solutions are proposed. Future directions are prospected from cross-disciplinary views including digital twins and personalized medicine, establishing a new paradigm for real-time, personalized, and continuous skin health management.

28 June 2026

Read appraisal →
observationalEvidence: Weak
40CEBM

Science advances

Smart contact lens-trained digital twin for device-free personalized uric acid prediction

Tears contain valuable biomarkers and offer potential for noninvasive disease monitoring. However, the lack of correlation analysis between serum uric acid (SUA) and tear uric acid (TUA) has limited their clinical application in personalized medicine. Here, we present a wireless smart contact lens capable of real-time, noninvasive monitoring of TUA as an alternative to blood-based UA testing. We validated the device through correlation analysis in rabbits and human participants, including individuals with hyperuricemia and gout. Daily-life monitoring enabled personalized characterization of TUA fluctuations in response to food intake and physical activity, together with individualized lag time profiling. A strong linear relationship allowed development of a regression model to derive estimated SUA from TUA. Building on these temporal profiles, lifestyle-informed digital twins were constructed to predict daily uric acid dynamics without continuous lens wear. This digital twin-enabled, device-free framework provides a practical route toward unobtrusive and personalized metabolic health monitoring.

27 June 2026

Read appraisal →
observationalEvidence: Weak
40CEBM

PloS one

Few-shot skin lesion classification with Adaptive Multi-Scale Convolutional Attention Network

Computer-aided diagnosis of skin lesions faces core challenges, including scale diversity, blurred boundaries, intra-class morphological variations, and sparse data. Existing methods often rely on fixed receptive fields or generic attention mechanisms, struggling to fully adapt to the unique characteristics of skin lesions. To address this, we propose an Adaptive Multi-scale Convolutional Attention Network (AMCANet), which aims to achieve accurate and robust classification of skin lesions with limited data. AMCANet comprises three core modules: the adaptive multi-scale convolution module dynamically adjusts the receptive field to accommodate lesions of varying sizes; the hierarchical channel attention module integrates multi-level semantic information across different resolutions; and the skin spatial attention module leverages image gradient information to enhance lesion boundaries and local texture features. Extensive few-shot experiments on the HAM10000 and PAD-UFES-20 public datasets demonstrate that AMCANet significantly outperforms existing baseline models across multiple metrics, exhibiting promising generalization capabilities on the evaluated datasets. Qualitative and visual analyses further validate the model's ability to extract discriminative features and effectively focus on lesion regions. This study proposes a deep learning model, which demonstrates certain effectiveness in classifying skin lesions even with a small number of samples, providing a potential direction for future research.

27 June 2026

Read appraisal →
observationalEvidence: Weak
55CEBM

Placenta

Deep learning-based early prediction of gestational diabetes mellitus through first-trimester placental texture analysis

INTRODUCTION: This study aimed to develop a multi-parameter fusion model for early GDM risk prediction and validate its performance through external multicenter testing. METHODS: A total of 628 pregnant women at 11+0-13+6 weeks were enrolled from two medical centers. The Center I cohort was divided into training (n = 356) and testing sets (n = 153). Radiomic features (1,289) and deep learning features (2,048) were extracted from placental ultrasound images. Feature-level fusion resulted in 3337 features, which were selected using Spearman correlation, mRMR, and LASSO. Five models were built: Rad Model, DTL Model, DLR Model, Clinic Model, and Combined Model. Performance was assessed using ROC analysis, DCA, and calibration curves. RESULTS: The Combined Model achieved the best overall performance, with an area under the ROC curve (AUC) of 0.879 in the internal validation, significantly outperforming any single-modality model (P < 0.05). DCA demonstrated that the fusion-based model provided higher net clinical benefit across a wide range of threshold probabilities compared with both "treat-all" and "treat-none" strategies. The calibration curve showed excellent agreement between predicted and observed probabilities (Hosmer-Lemeshow test, P > 0.05). DISCUSSION: The multimodal fusion model enhanced early GDM prediction by detecting subtle placental changes in first-trimester, enabling timely intervention and personalized decision-making.

27 June 2026

Read appraisal →
observationalEvidence: Weak
25CEBM

Journal of health organization and management

Unlocking potential: impact of mHealth affordance realization on patient well-being

PURPOSE: This study investigates how realized affordances of mobile health technologies influence the relationship between technology use, patient empowerment and well-being in diabetes self-management. DESIGN/METHODOLOGY/APPROACH: The research proposes a technology use-empowerment-well-being (TEW) model based on affordance theory. Data were collected through a survey of 257 diabetes patients using mobile technologies for self-management and analyzed using structural equation modeling with SmartPLS. FINDINGS: Results indicate that patient empowerment fully mediates the relationship between mobile technology use and patient well-being. Importantly, realized affordances moderately strengthen the relationship between technology use and patient empowerment, demonstrating that patients derive greater benefits when they actively utilize technology's capabilities for diabetes management. RESEARCH LIMITATIONS/IMPLICATIONS: Reliance on self-reported data and the cross-sectional design limit causal inferences about the evolution of technology-empowerment-well-being relationships over time. PRACTICAL IMPLICATIONS: Healthcare technology designers should focus on creating clear, tangible benefits for patients that support specific self-management behaviors. Healthcare providers should educate patients about the benefits of technology and support them in realizing these affordances through training and ongoing assistance. ORIGINALITY/VALUE: This study integrates affordance theory with patient empowerment concepts, providing empirical evidence that realized affordances - not merely technology availability - are critical for improving self-management outcomes. The TEW model introduces a novel framework for understanding how technology empowers patients and enhances well-being in chronic disease management.

27 June 2026

Read appraisal →
observationalEvidence: Weak
55CEBM

Familial cancer

When screentime fails: initiative to improve completion of hereditary cancer genetic testing after telemedicine counseling.

Telemedicine has broadened access to genetic counseling while maintaining high levels of patient satisfaction. However, emerging evidence suggests lower completion rates of genetic testing following virtual visits compared with in-person encounters. This study evaluated genetic testing completion rates among patients who underwent cancer risk assessment via telemedicine and assessed the effect of a structured follow-up telephone reminder on genetic testing completion. We conducted a retrospective review of a quality improvement initiative - Genetic Engagement via Navigation and Encouragement (GENE CALL) - designed to address incomplete genetic testing following telemedicine counseling. The study included all patients who underwent virtual cancer risk assessment between September 2023 and July 2024 at a single urban academic cancer genetics program. Patients who expressed interest in genetic testing but had not scheduled specimen collection were contacted by telephone as a reminder. During the call, patients were asked about ongoing interest in testing and barriers to scheduling; they were also offered assistance completing specimen collection. Among 647 patients evaluated via telemedicine, 469 (72.5%) expressed interest in genetic testing. Of these, 290 (61.8%) had scheduled or completed testing, while 179 (38.2%) had not. Among those who had not scheduled testing, 120 (67.0%) were successfully reached by telephone, and 98 (81.7%) confirmed ongoing interest in completing testing. At 3-month follow-up, 48 (40.0%) of the 120 contacted patients had completed testing, compared with 13 (22.0%) of 59 patients who were not reached (p = 0.017). Telephone follow-up is an effective strategy to improve genetic testing completion after virtual counseling.

27 June 2026

Read appraisal →
Randomised Controlled TrialEvidence: Weak
55CEBM

The Keio journal of medicine

Comparison of Clinical Development of Digital Therapeutics among Japan, the United States, and Germany

Digital therapeutics (DTx) have demonstrated promising potential as a novel therapeutic approach, with their clinical development gaining momentum. Comparison of DTx development between Japan, a relative newcomer, and Germany and the U.S., known for favorable circumstances, is a topic of interest. DTx developed in Japan as of November 2023 were identified on the Pharmaceuticals and Medical Devices Agency (PMDA) website and in the literature. Similar DTx available in the U.S. and Germany were identified in the databases maintained by the U.S. Food and Drug Administration (USFDA) and Federal Institute for Drugs and Medical Devices (BfArM). Data on their clinical trials and regulatory status were obtained from these databases and their respective national clinical trial registries. The data obtained were compared and analyzed. By November 2023, Japan had developed DTx, encompassing 11 therapeutic areas. Seven DTx had reached or completed the confirmatory study stage, whereas five DTx were in the exploratory phase. Twenty DTx in the therapeutic areas were selected from the U.S. and Germany. A total of 27 DTx were reviewed regarding the designs of confirmatory studies and regulatory actions taken. Trial designs demonstrated more similarities than differences across the countries and therapeutic areas. Placebo trial design-associated difficulties were conspicuous. Although regulatory actions to place DTx in the market differed across countries, they effectively ensured the efficacy and safety of DTx that were suitable for marketing considering the most current science. The regulatory measures in the three countries seem to have positively impacted DTx development.

27 June 2026

Read appraisal →
qualitativeEvidence: Weak
40CEBM

JMIR formative research

Supporting Student Mental Health With the Safespace Generative AI Chatbot: Mixed Methods Feasibility Study

BACKGROUND: Generative artificial intelligence (GenAI) chatbots have the potential to provide personalized mental health support to individuals at scale. OBJECTIVE: This study evaluates the feasibility and usage patterns of the Safespace GenAI chatbot, an artificial intelligence (AI)-driven smartphone app that offers a large language model-powered interactive chatbot to support mental health. METHODS: Using a mixed methods approach, we explored baseline attitudes toward GenAI chatbots and chatbot usage patterns, conducted a qualitative content analysis of participants' experiences, and descriptively assessed patterns related to preintervention depressive symptoms. The study included an initial sample of 42 university students, 20 of whom actively used the chatbot over 2 to 4 weeks, generating 286 user-chatbot interactions. RESULTS: Preintervention surveys indicated that the majority of participants anticipated that the chatbot would be helpful (27/42, 64%) and that they trusted its privacy safeguards (39/42, 93%). Usage patterns suggested that the highest levels of interaction occurred early in the morning and late at night, when peer and professional support may be inaccessible. The qualitative analysis indicated that participants appreciated using the chatbot for reflection as a blended-care tool between their counseling sessions, while also naming technical barriers and specific design needs required to sustain engagement. In addition, our exploratory analyses descriptively showed that participants with elevated depression scores engaged in emotional disclosure during 99% (38 sessions with 8 participants) of their sessions, compared to 84% (26 sessions of 12 participants) of those with low symptoms. Due to the small sample size, future adequately powered studies are needed to inferentially examine these observed patterns. CONCLUSIONS: These findings provide initial insights into the usage and engagement dynamics of the Safespace GenAI chatbot and highlight directions for future research to optimize GenAI-driven mental health interventions.

27 June 2026

Read appraisal →
Systematic ReviewEvidence: Weak
45CEBM

PloS one

Artificial Intelligence in emergency department triage: A scoping review

BACKGROUND: Triage in emergency departments (ED) is a critical process for prioritizing care and ensuring clinical safety. However, current triage systems often exhibit vulnerabilities that compromise the efficiency and quality of healthcare delivery. Artificial Intelligence (AI) has emerged as a promising innovation to support decision-making and optimize patient flow in these high-pressure environments. OBJECTIVE: To map the available evidence regarding the implementation and performance of artificial intelligence in emergency department triage. METHOD: This scoping review followed the Joanna Briggs Institute (JBI) methodology and the PRISMA-ScR guidelines. A comprehensive search was conducted across 13 databases (CINAHL, Cochrane Library, PubMed Central, SciELO, Web of Science, SCOPUS, Science Direct, VHL, Embase, and several regional dissertation repositories), with no language or time restrictions. Two independent reviewers performed the selection process using the Rayyan platform, with discrepancies resolved by a third evaluator. Data were synthesized using the PAGER framework, categorizing findings into Patterns, Advances, Gaps, Evidence for practice, and Recommendations for research. RESULTS: Nineteen studies met the inclusion criteria. AI was primarily implemented through Machine Learning (ML) algorithms, including Deep Learning architectures. Natural Language Processing (NLP) was frequently employed to process unstructured clinical data, with recent studies exploring the potential of Large Language Models (LLMs). Overall, ML-based models consistently outperformed traditional triage systems in predictive accuracy. These techniques were mainly utilized for automated classification, predicting clinical severity, and enhancing patient prioritization by integrating both objective and subjective assessment data. CONCLUSIONS: The findings indicate that AI has significant potential to enhance emergency triage by streamlining service flows and providing robust clinical decision support. However, the current evidence remains heterogeneous and largely exploratory. Key challenges include variability in model performance, a lack of external validation, and studies often limited to specific populations. Consequently, many current tools still lack the necessary reliability for safe, large-scale clinical implementation.

27 June 2026

Read appraisal →
otherEvidence: Weak
55CEBM

Journal of medical Internet research

Personal Health Large Language Models and the Negotiation of Medical Authority in Clinical Care: Opportunities, Risks, and Governance

Personal health large language models (PH-LLMs) are patient-facing conversational systems that synthesize user-entered information, patient-generated health data, wearable data, and selected personal health records-where users choose to connect them-into personalized, longitudinal, action-oriented health narratives. Unlike generic health chatbots that mainly provide one-off responses to isolated questions, PH-LLMs may generate continuing interpretations, priorities, and candidate next steps that patients bring into clinical encounters. In contrast to electronic health record-tethered clinical artificial intelligence (AI), they often originate outside institutional oversight and may be selected, used, or trusted by patients before professional review. This viewpoint examines how PH-LLMs may reshape the negotiation of medical authority by contributing to a shift from the traditional dyadic clinician-patient relationship toward a triadic model of negotiated authority, in which clinicians may increasingly need to mediate among clinical evidence, patient values, and algorithmic narratives. PH-LLMs may support patient participation by organizing symptoms, contextualizing home-monitoring and wearable data, improving health literacy, assisting chronic disease self-management, and preparing patients for more collaborative visits. Patient-facing AI narratives may also introduce distinct risks. At the individual level, these include inaccurate or incomplete responses arising from imprecise queries or missing context, misinterpretation of otherwise accurate information in the absence of clinical context, and contextually biased or poorly matched advice across demographic, cultural, linguistic, disability-related, or socioeconomic contexts. At the system level, they include authority conflict when AI recommendations diverge from clinical judgment, fragmentation of clinical truth, privacy and data-governance concerns, diffusion of accountability when harm results from advice produced outside clinical governance, and inequitable access to premium tools and continuous monitoring devices. To address these challenges, we propose a 3-layer clinical governance framework for patient-brought PH-LLM narratives. The first layer, evidence and provenance, makes AI-generated narratives epistemically legible by clarifying platform identity, data sources, temporal anchoring, uncertainty, and privacy-relevant data-use and retention conditions. The second layer, clinical arbitration and workflow integration, uses risk-stratified intake, proportionate documentation, escalation triggers, and equity-preserving workflows to embed PH-LLM outputs into routine care. The third layer, competence and accountability, defines the communication competencies, AI literacy supports, institutional responsibilities, vendor accountability, and risk-proportionate verification duties needed for triadic care. This framework is a conceptual and governance-oriented proposal rather than a validated clinical protocol. Future empirical work should evaluate its feasibility, documentation burden, equity effects, clinical safety impact, and acceptability among patients, clinicians, and health systems. Governed through these interdependent layers, PH-LLMs may serve as supporting infrastructure for safer, person-centered longitudinal care.

26 June 2026

Read appraisal →
Systematic ReviewEvidence: Weak
55CEBM

Journal of medical Internet research

Accessibility of Digital Financial Applications for People With Visual Impairment: Scoping Review

BACKGROUND: Routine financial activities are now conducted primarily through digital channels. Many such systems remain inaccessible to more than 2.2 billion people globally living with vision impairment, limiting independent financial management. Constrained access can create financial strain and social disadvantage, reducing access to health-enabling resources, and contributing to avoidable health inequities. OBJECTIVE: This scoping review maps evidence on the accessibility of digital financial services for individuals with visual impairment (VI) as a digital determinant of health. We synthesized barriers and facilitators, characterized study designs, settings, and populations, and identified evidence gaps to inform inclusive design, digital health research priorities, and policy. METHODS: A scoping review was conducted using the Joanna Briggs Institute framework and reported in line with PRISMA-ScR (Preferred Reporting Items for Systematic Reviews and Meta-Analyses extension for Scoping Reviews) guidelines. Eight databases (PubMed, MEDLINE, CINAHL, Scopus, Web of Science, Business Source Complete, ProQuest, and IEEE Xplore) were searched for peer-reviewed papers in English published between 1995 and 2026. Searches featured controlled vocabulary and free-text terms structured in 3 conceptual blocks (VI, digital financial services, and accessibility or usability). A random sample of 20% of titles, abstracts, full texts, and included studies was independently screened or charted by 2 reviewers to calibrate decisions; the remainder were screened and charted by a single reviewer. Data were charted using a standardized extraction form, and results were synthesized descriptively and thematically. RESULTS: Twenty-three studies met the inclusion criteria. Studies were conducted across 12 countries, with the largest number from India (n=7), Indonesia (n=2), Thailand (n=2), and the United States (n=2). Study designs included qualitative studies (n=6), mixed methods studies (n=1), cross-sectional studies (n=4), nonrandomized experimental studies (n=2), and technical or design-focused evaluations (n=6). One study was a large population survey (n=19,136), and the remaining studies with human participants had sample sizes ranging from 4 to 36 participants. Accessibility barriers were reported across all platform types, with authentication-related barriers described in 18 studies and screen reader incompatibility in 17 studies. Reported barriers included reliance on sighted assistance for tasks such as login, verification, and payments, compromising privacy and independence. Facilitators included assistive technology support, logical navigation order, nonvisual feedback mechanisms, and accessible authentication alternatives. Evidence mapping revealed recurrent barrier patterns across Android, iOS, and web platforms. No longitudinal or intervention-based evaluations were identified. CONCLUSIONS: This review provides a focused synthesis of accessibility evidence at the intersection of digital financial services and VI, a domain addressed by neither prior digital accessibility reviews nor financial inclusion for people with disabilities. Authentication methods, interface labeling, and navigation were identified as persistent cross-platform accessibility barriers. The findings carry implications for financial technology developers, accessibility auditors, and policymakers implementing accessibility legislation and extend the digital determinants of health framework by demonstrating how inaccessible financial technology may compound health inequities.

25 June 2026

Read appraisal →
Systematic ReviewEvidence: Weak
20CEBM

JMIR medical informatics

Blockchain-Based Dynamic and Revocable Consent for Secondary Health Data Use: Systematic Review

BACKGROUND: The secondary use of health data holds substantial potential for advancing biomedical research, strengthening population health analytics, and enabling artificial intelligence-driven decision-making support. Yet, ensuring that such reuse respects patient autonomy, privacy, and regulatory obligations remains a major challenge. Conventional consent mechanisms are typically static, difficult to revoke, and offer limited transparency or accountability after data disclosure. OBJECTIVE: This review aimed to systematically examine blockchain-based frameworks that enable dynamic, auditable, and revocable consent for the secondary use of health data. METHODS: A structured literature search was conducted in PubMed, Scopus, and Web of Science covering the period 2020 to 2025. Following PRISMA (Preferred Reporting Items for Systematic Reviews and Meta-Analyses) guidelines, 55 peer-reviewed studies meeting predefined inclusion criteria were analyzed. Data extraction focused on four dimensions: (1) consent life cycle management, (2) auditability and traceability, (3) usability and patient empowerment, and (4) legal and ethical alignment. RESULTS: Findings indicate that blockchain technologies provide a robust foundation for automating consent life cycles, ensuring immutable auditability, and enabling decentralized patient control. Most frameworks used smart contracts, decentralized identifiers, and verifiable credentials to implement programmable and verifiable consent processes. Nevertheless, key challenges persist, including limited usability testing, complexities in real-time revocation propagation, interoperability gaps with clinical systems, and tensions with regulatory requirements such as the General Data Protection Regulation right to erasure. Only a small subset of studies reported real-world deployments or user-centered evaluations. CONCLUSIONS: Blockchain offers substantial promise for improving the trustworthiness, transparency, and accountability of consent management for secondary health data use. However, wider adoption requires human-centered design approaches, stronger interoperability through standards such as Fast Healthcare Interoperability Resources, verifiable credentials, and consent receipts, and clearer legal guidance for compliance. Future research should prioritize integrating blockchain-enabled consent infrastructures into national and cross-border digital health ecosystems such as the European Health Data Space to support secure, patient-controlled, and ethically governed secondary data use.

24 June 2026

Read appraisal →
otherEvidence: Weak
30CEBM

Journal of medical Internet research

Solidarity or Segregation? ChatGPT Health and US Health Care Disparities.

ChatGPT Health, an artificial intelligence feature launched by OpenAI in January 2026, integrates personal medical records and consumer health data into a large language model-based chatbot. It positions itself as a digital health tool to help users navigate a fragmented US health care system marked by inaccessible insurance, medical deserts, and workforce shortages. We argue that, rather than closing these gaps, ChatGPT Health risks widening what we call the solidarity gap in health care. By appearing to address unmet needs, it may obscure and entrench the structural conditions that produce health care disparities. We identify specific risks to patient safety, such as the legitimization of dangerous self-rationing and self-medication behaviors, inaccurate triage of clinical emergencies, erosion of effective health communication, and the reinforcement of confirmation and anchoring bias. These risks are particularly acute for individuals facing structural barriers to care. To prevent ChatGPT Health and similar large language model-based health tools from deepening inequities, we propose solidarity-based policy interventions such as upstream reforms that rebuild the institutional foundations of access and health equity and downstream safeguards that embed artificial intelligence tools within accountable, equitable health care infrastructures as complements to rather than substitutes for human care.

24 June 2026

Read appraisal →
otherEvidence: Moderate
70CEBM

JMIR formative research

Clinician Decision-Making Around Offering Home Video Telehealth Visits: Qualitative Study

BACKGROUND: Use of in-home video telehealth rapidly expanded in response to the COVID-19 pandemic, including at Veterans Affairs (VA), a forerunner in telehealth. Despite this uptick, differences in use by patient age and rurality created a digital divide that persists to this day. While clinicians frequently cite patients' older age and lack of technical skills as barriers to in-home video telehealth, it remains unclear how clinicians decide whether to offer video visits to patients and to what extent these beliefs may hinder offering video visits to older adults. Gathering perspectives from clinician users of in-home video telehealth may illuminate opportunities to ensure continued access to care through solutions such as telehealth. OBJECTIVE: This study aimed to examine clinician decision-making around the offer of in-home video telehealth to understand how interprofessional clinicians determine to whom they offer in-home video telehealth and what factors (organizational, personal, or attitudinal) influence their decision. METHODS: We conducted a qualitative study by using semistructured interviews. Participants were interprofessional clinicians (N=16) employed by 11 different VA hospitals and included 1 clinical pharmacist, 6 medical doctors, 2 nurse practitioners, 1 occupational therapist, 3 psychologists, 1 physical therapist, 1 speech-language pathologist, and 1 social worker. All the participants had at least some experience using in-home video telehealth from locations across VA (the largest integrated health care system in the United States) and were interviewed over a 6-month period. Interviews focused on clinicians' use of video telehealth and the decision-making process involved in offering in-home video telehealth. We used directed content analysis with a rapid analytic approach, given the time-pressured nature of our project. RESULTS: This study revealed that clinician decision-making around offering in-home video visits is complex and influenced by several domains, namely, (1) clinician factors, including experience with video and perceived benefits of video; (2) appointment factors, including the visit's clinical goal; (3) clinician-reported patient factors, including age and willingness to try video; (4) patient social context, including caregiver availability; (5) geographical factors, such as availability of reliable high-speed internet and patient distance from the medical center; and (6) health system factors, including technical support and clinician ability to work from home. Access to in-home video telehealth may be facilitated by clinician familiarity and confidence with telehealth technology, strategies to improve patients' technical skills, and support for caregivers. Infrastructure also plays an important role, including device availability, broadband reliability, and clear protocols for matching services to video visits. CONCLUSIONS: Findings highlight the importance of clinician competency in telehealth, patient and caregiver digital readiness, and a supportive technology infrastructure to equitable in-home video care. In addition, improved guidance for specific clinical services and awareness of potential biases may enable consistent, accessible telehealth delivery for older adults and medically complex populations.

20 June 2026

Read appraisal →
Randomised Controlled TrialEvidence: Weak
55CEBM

Lasers in medical science

Clinical dosimetry and efficacy of LED photobiomodulation for chronic lower-limb wound healing: a systematic review of randomized trials

Chronic lower-limb wounds represent a major clinical and socioeconomic burden due to delayed healing and high recurrence rates. Light-emitting diode (LED) photobiomodulation has been proposed as a noninvasive and low-cost adjunctive therapy; however, clinical evidence remains inconsistent. This systematic review aimed to synthesize evidence from randomized clinical trials investigating the efficacy of LED photobiomodulation for chronic lower-limb wound healing and to identify irradiation parameters associated with improved outcomes. A comprehensive search was conducted in PubMed, Scopus, Web of Science, and Embase up to May 2026. Six randomized clinical trials met the inclusion criteria. Wavelengths ranged from 620 to 950 nm, and energy densities varied between 2.4 and 126 J/cm². The findings suggest that LED photobiomodulation may promote wound area reduction, improve wound bed quality, and increase microcirculation, particularly in diabetic foot ulcers. However, one study using a high energy density (126 J/cm²) did not demonstrate beneficial effects, suggesting a possible dose-dependent response. The overall certainty of the evidence, assessed using the GRADE approach, was classified as very low due to inconsistency and indirectness among the included studies. Although LED photobiomodulation appears to be safe and demonstrates therapeutic potential, substantial heterogeneity in irradiation parameters, small sample sizes, and methodological limitations preclude definitive conclusions regarding its clinical efficacy. Well-designed randomized controlled trials with standardized protocols and dose-response investigations are needed to establish optimal therapeutic parameters and confirm clinical efficacy.

20 June 2026

Read appraisal →
Systematic ReviewEvidence: Weak
40CEBM

Journal of medical Internet research

Ethical Considerations in Personal Health Large Language Models

Personal health large language models (PH-LLMs) have rapidly evolved from research prototypes into consumer-facing, data-linked systems that support symptom triage, medication questions, mental health check-ins, and longitudinal self-management. Their direct-to-consumer use without clinical oversight creates a distinct ethical risk profile that general artificial intelligence governance frameworks do not fully address. This viewpoint focuses on text-based, platform-mediated PH-LLMs and synthesizes PH-LLM-specific challenges across 6 domains: privacy, accuracy, equity, transparency, human-artificial intelligence interaction, and regulatory governance. These risks may be amplified by health literacy gaps, longitudinal data aggregation, persuasive conversational design, and fragmented oversight across the consumer-clinical boundary. Grounded in the 4 principles of biomedical ethics, we propose a governance framework that operationalizes beneficence, nonmaleficence, autonomy, and justice through design and deployment controls, including health literacy-aligned communication, crisis and pharmacological safeguards, hallucination mitigation, role disclosure, granular consent, fairness auditing, and accessible design. We further outline implementation mechanisms, including risk-tiered certification, tiered accountability, and postdeployment oversight through adverse-event reporting, transparency reporting, and independent safety evaluation. This framework is intended as an evidence-informed but partly anticipatory approach to governing PH-LLMs in personal health management.

19 June 2026

Read appraisal →