Evidence-Based Medicine

Research Appraisals

Evidence-based critical appraisals of the latest medical research, systematically evaluated using Oxford CEBM methodology.

Showing 14 appraisals

otherEvidence: Moderate
65CEBM

Rheumatic diseases clinics of North America

Sources of Bias in Clinical Artificial Intelligence and Applications in Rheumatology

Rheumatology machine-learning models are limited by preexisting, technical, and emergent biases; the interaction of data constraints, design choices, and real-world clinical workflows, rather than from isolated technical errors. Across the model lifecycle, optimization objectives can encode patterns of care, access, and documentation, producing hidden subgroup failures that are obscured by aggregate performance metrics. Given the heterogeneity of rheumatic disease and current disparities in care delivery, addressing bias requires deliberate design choices before, during, and after a model is built, as well as a commitment to transparency, and sustained oversight.

2 Aug 2026

Read appraisal →
Systematic ReviewEvidence: Weak
45CEBM

Preventing chronic disease

Reducing Rates of Cesarean Delivery in Rural US Communities: A Systematic Review of Interventions and Approaches to Care

INTRODUCTION: Cesarean deliveries are the most common major surgery in the US, with rates rising disproportionately in rural communities. While sometimes medically necessary, unnecessary cesarean births increase risks for maternal death, long-term complications, and intergenerational health effects that contribute to the burden of chronic disease. This systematic review examined studies describing interventions and approaches to care that reported outcomes related to reducing cesarean delivery rates in rural US communities. METHODS: We searched 4 databases in September 2025. Studies were eligible if they were conducted in rural US settings and reported cesarean-related outcomes associated with an intervention or care approach. We categorized interventions as patient level, provider level, or health system level. Two reviewers independently screened articles for inclusion, extracted data, and assessed study quality using the Newcastle-Ottawa Scale and the Joanna Briggs Institute checklist. RESULTS: Nine studies met inclusion criteria. Of the 2 patient-level interventions, psychosocial education was associated with lower cesarean delivery rates (21% in intervention vs 40% in control), whereas a mobile health application showed only a marginal difference (27.1% among application users vs 27.7% among nonusers). None were categorized at the provider level. Seven interventions tested system-level models, primarily comparing hospitals with different staffing patterns; family medicine-led hospitals had lower rates of low-risk nulliparous, term, singleton, vertex cesarean delivery than hospitals staffed by both family medicine physicians and obstetricians (23% vs 28%), certified nurse-midwife-managed births had lower cesarean delivery rates than family medicine physician-managed births (8% vs 14%), and collaborative maternity care models integrating midwives, nurses, and obstetricians were associated with cesarean delivery rates declining from 26.2% to 11.2%. CONCLUSION: System-level approaches, particularly those that restructure maternity care teams, emphasize family medicine physician-led models, and integrate midwifery and culturally grounded childbirth practices, are more consistently associated with lower cesarean delivery rates in rural US settings than patient-level interventions alone. Future efforts to reduce unnecessary cesarean deliveries should prioritize strategies tailored to the variability of rural care capacity.

31 July 2026

Read appraisal →
observationalEvidence: Weak
30CEBM

Journal of medical Internet research

Initial-Visit Specialty Triage in Rare Diseases Using Large Language Models: Retrospective Benchmarking Study

BACKGROUND: Specialty triage at first contact is an overlooked step in early diagnostic pathways for rare diseases. Patients often present with overlapping, multisystem, and atypical manifestations, making first-visit specialty selection challenging and potentially prolonging diagnostic pathways. OBJECTIVE: The aim of this study is to evaluate the accuracy, response time, and consistency of large language models (LLMs) for initial-visit specialty triage in rare diseases across multiple datasets, and to compare their performance with registered nurses and nonmedical participants. METHODS: In this retrospective benchmarking study, we used 5 rare disease datasets: a publication-derived case set, 3 RareBench-derived datasets, and a Facial phenotype-Gene-Disease Dataset-derived set. Fourteen LLMs were evaluated over 5 independent runs per case. Performance was assessed using accuracy, response time, and consistency, with subgroup analyses by model accessibility, reasoning mode, parameter scale, and phenotype count. Human comparison was conducted on the publication-derived case set using registered nurses and nonmedical participants. RESULTS: Across datasets, model accuracy ranged from 0.4378 to 0.7141. Claude-opus-4-5 achieved the highest accuracy (0.7141) and consistency (0.9653), averaging 10.79 seconds per case. GPT-5.1 had the shortest response time (3.39 s/case) and high accuracy (0.6948). Proprietary models had numerically higher average accuracy than open-weight models (0.6973 vs 0.6365). Nonthinking models achieved higher average accuracy than thinking models (0.6789 vs 0.5826) and had shorter response times, although this exploratory comparison was based on a small number of thinking models. Accuracy varied by phenotype count, with higher performance in cases with 1 to 2 or more than 14 phenotypes. On the publication-derived case set, LLMs achieved higher average accuracy than registered nurses and nonmedical participants (0.5978 vs 0.4914 and 0.4573). CONCLUSIONS: LLMs showed potential as assistive tools for initial-visit specialty triage in rare diseases. Model choice, reasoning mode, and phenotype information density influenced performance, but subgroup findings should be interpreted cautiously. Future work should evaluate LLM-based specialty triage in prospective clinical settings and develop clinician-supervised workflows with traceable evidence support.

24 July 2026

Read appraisal →
otherEvidence: Weak
50CEBM

Journal of medical Internet research

Model and Task-Aware Test-Time Scaling Strategies for Large Language and Vision-Language Models in Medicine: Evaluation Study

BACKGROUND: Test-time scaling has emerged as a promising method to enhance the reasoning capabilities of large language models (LLMs) and vision-language models (VLMs) during inference without additional training. While foundational studies established scaling paradigms in general domains, their applicability to the unique complexities of medical AI remains underexplored. OBJECTIVE: This study aims to conduct a comprehensive investigation of test-time scaling in the medical domain. We evaluate the impact of scaling across different model sizes and task complexities. Furthermore, we seek to identify domain-specific bottlenecks and assess model robustness against user-driven perturbations, such as misleading clinical authority. METHODS: This study evaluated a diverse set of general and medical-specific LLMs and VLMs. Experiments used five textual medical benchmarks comprising over 5500 questions and two multimodal benchmarks comprising 7000 samples. Performance was measured under three scaling conditions: increasing token budgets, iterative sequential scaling, and parallel scaling. Robustness was tested by embedding misleading hints with varying tones and levels of simulated clinical expertise into prompts. RESULTS: For nonreasoning LLMs, accuracy saturated quickly, with token usage often remaining under 500 tokens regardless of budget increases. Reasoning models demonstrated significant performance gains on complex tasks as token budgets increased. Notably, we identified distinct domain-specific behaviors. First, current VLMs showed a structural bottleneck in integrating visual clues and experienced limited benefit from token expansion. Second, medically fine-tuned LLMs excelled in clinical question answering but exhibited degraded scaling efficiency on calculation tasks compared to general-domain models. This reflects a disparity between qualitative clinical alignment and procedural logic. Third, while optimal scaling improved robustness, models exhibited a cognitive vulnerability by readily abandoning correct reasoning when confronted with misleading expert physician hints. Regarding scaling strategies, parallel scaling outperformed sequential scaling on easier tasks. Conversely, extended sequential scaling or increased budgets proved essential for complex problem-solving. CONCLUSIONS: Test-time scaling rules from general domains do not perfectly translate to medical AI. Longer reasoning is not universally beneficial. Concise reasoning with parallel scaling is optimal for simpler tasks. An extended chain of thought via sequential scaling or increased budgets is required for complex problems. Furthermore, safe clinical deployment requires addressing fundamental vision-language alignment, balancing clinical and procedural reasoning, and mitigating vulnerabilities to perceived clinical authority.

24 July 2026

Read appraisal →
otherEvidence: Weak
35CEBM

Journal of the American College of Surgeons

Defining Quality Metrics for Telemedicine in Surgery: A Critical Examination

The rapid adoption of telemedicine has transformed healthcare delivery in the US, enhancing access and communication for surgical patients across wide geographic areas. However, a critical challenge persists: the need to define and standardize quality metrics specifically for telemedicine in surgery. Existing studies suggest that telemedicine can yield equivalent or improved outcomes compared with in-person visits. Nevertheless, literature remains limited in evaluating patient satisfaction, cost-effectiveness, and diagnostic accuracy. Through expert consensus within the American College of Surgeons Board of Governors Telehealth Pillar, we propose a structured framework for defining and implementing quality metrics for surgical telemedicine. This is based on the Donabedian model, addressing structure, process, and outcome, while also considering the distinct phases of surgical care: preoperative, intraoperative, and postoperative. This framework identifies unique domains of telemedicine in surgical care, emphasizing hospital and organizational structure, patient and provider readiness, and policy alignment. A comparative analysis with existing AHRQ and WHO frameworks highlights gaps in surgical applicability. Finally, we propose an implementation roadmap prioritizing immediate, feasible metrics while identifying areas for future validation. Collaboration among researchers, clinicians, and policymakers will be essential to establish these metrics and ensure that telemedicine delivers on its potential to improve surgical care delivery while addressing disparities in access and outcomes.

18 July 2026

Read appraisal →
otherEvidence: Weak
65CEBM

BMJ open

Implementation determinants of a planned machine learning-enabled surgical scheduling system in a high-volume orthopaedic centre in Canada: qualitative findings.

OBJECTIVES: Elective non-emergent surgical wait times have increased across countries such as Canada, straining operating room (OR) resources and affecting patient outcomes and healthcare spending. Manual scheduling systems in Ontario orthopaedic centres create wide variations in wait times, with recent declines in meeting benchmark targets despite increased procedure volumes. Challenges stem from fragmented referral processes, outdated scheduling methods and resource constraints. Artificial intelligence and machine learning (ML) offer potential solutions for optimising scheduling; however, their implementation remains inconsistent. This study aims to identify determinants affecting the rollout of a new ML-driven automated scheduling system at a high-volume elective orthopaedic surgery centre. DESIGN: A qualitative description approach supported by implementation science frameworks. SETTING: A high-volume elective orthopaedic surgery unit at a Canadian tertiary care centre. PARTICIPANTS: 17 individuals from clinical, administrative and leadership roles who were directly involved in surgical scheduling. INTERVENTIONS: A new ML-driven automated surgical scheduling system. OUTCOMES: Perceptions of the proposed new surgical scheduling system (barriers and enablers of implementation, recommendations for improvement). RESULTS: Three main themes were identified, capturing challenges and enablers in the existing scheduling system: system functionality, process-related factors and resource constraints.Participants described substantial inefficiencies in the existing manual scheduling system, including outdated software, fragmented information systems, inconsistent communication and resource constraints. Across interest-holder groups, there was broad but variable perceived support for a planned ML-enabled scheduling system, particularly for improving duration prediction, access to scheduling data and reporting, alongside concerns about system complexity, workflow fit, training and resource implications. Interest-holders emphasised the importance of user-friendly design, interoperability, responsive training, phased implementation and ongoing feedback. CONCLUSIONS: This pre-implementation qualitative study identified significant process and resource limitations in manual orthopaedic surgical scheduling, but interest-holder support for a well-designed ML-driven system is strong. While participants anticipated potential benefits for scheduling accuracy, throughput and resource allocation, these perceived advantages will require meaningful user engagement, robust training, phased rollout and evaluation in subsequent implementation and outcome studies.

8 July 2026

Read appraisal →
Systematic ReviewEvidence: Moderate
75CEBM

Transactions of the Royal Society of Tropical Medicine and Hygiene

Extreme weather effects on health services and communities in low and lower-middle income countries: a thematic systematic review

Most previous research about the dangers of extreme weather events was applicable to populations in high-income countries. Data summarising harms related to extreme weather events in low-income settings are lacking. A systematic review thematically summarising evidence about weather event-linked harms and responses in low- and lower-middle-income countries was conducted. Peer-reviewed and grey literature was systematically searched and selected. Data were extracted about harms, responses and outcomes relevant to six WHO building blocks of healthcare systems. Framework analysis was used to identify predominant themes related to harms, responses and the WHO building blocks. In total, 183 reports were included. Flooding and high winds were the most common types of extreme weather events documented. The main community experience themes identified were the displacement of populations and disruption. The main themes identified for health service delivery were vulnerability, disruption and resilience. Documented examples of resilience or recovery were far fewer for all six WHO healthcare system building blocks than descriptions of vulnerability and disruption. Extreme weather events can be highly disruptive and harmful to healthcare systems and communities in LMIC settings that are often already highly vulnerable.

8 July 2026

Read appraisal →
observationalEvidence: Weak
60CEBM

Journal of hospital medicine

Gender differences in secure chat and electronic health records use among hospital-based physicians

BACKGROUND: Digital messaging within electronic health records (EHRs) is central to inpatient communication. While intended to enhance efficiency, these tools may contribute to unequal digital workloads. OBJECTIVE: To evaluate gender-based differences in EHR and secure messaging use among hospital-based physicians. METHODS: Retrospective observational study at a single academic tertiary care hospital using EHR metadata from July 2023 to June 2024. Participants included a total of 205 internal medicine and medicine-pediatrics physicians (108 senior residents, 97 faculty) serving as primary clinicians on hospitalist shifts. Measures included daily number of messages sent/received, time spent in EHR and secure chat, and hours worked, stratified by gender and role. RESULTS: Physicians worked a median of 9.4 h/day, spending 40.7% of their time in the EHR and 6.3% in secure chat. Women and men work similar hours (9.5 vs. 9.3, p = .40) but women spent more time in the EHR (248 vs. 222 min/day; p < .001) and secure chat (38 vs. 34 min, p < .001) and exchanged more daily messages (62 vs. 53 sent; 56 vs. 48 received, both p < .001). Patterns were consistent across residents and faculty. Among faculty, gender differences persisted after a new scheduling model, which reduced overall work hours. CONCLUSIONS: Women physicians engaged more in digital communication than men, despite similar hours worked, and these differences persisted after a workflow change. These findings underscore the need for equity-informed strategies that both mitigate excess burden and recognize the potential value of proactive communication in hospital settings.

3 July 2026

Read appraisal →
otherEvidence: Weak
25CEBM

Journal of medical Internet research

Transformation Versus Innovation in Digital Health Care and the Future of Clinical AI

Digital innovation is frequently presented as the key to transforming health care. In this News and Perspectives article, JMIR Correspondent and academic physician Boon-How Chew reports on the distinction between and the direction of transformation and innovation, reflecting on the future of clinical AI and what lasting change in health care ultimately requires.

3 July 2026

Read appraisal →
Systematic ReviewEvidence: Moderate
40CEBM

Telemedicine journal and e-health : the official journal of the American Telemedicine Association

Implementation and Evaluation of Virtual Care in Canadian Health Care Systems: A Scoping Review

OBJECTIVE: This scoping review examined available evidence in implementation and evaluation of virtual care in Canada. Virtual care saw recent uptake due to the COVID-19 pandemic; however, to ensure quality of care, rigorous implementation and evaluation frameworks are needed. METHODS: Peer-reviewed and gray literature were searched to determine extent, range, and nature of evidence surrounding implementation and evaluation of virtual care based on the guidelines of the Joanna Briggs Institute. Although virtual care can encompass synchronous and asynchronous modalities, this review focused on synchronous virtual care, defined as real-time interactions between patients and providers via videoconferencing or telephone. Search included MEDLINE, EMBASE, Psych Info, and CINAHL databases and national and provincial health system, professional organization, and regulatory websites. Inclusion criteria included videoconferencing or telephone and English and French Canadian sources. Citations were screened by two researchers at title, abstract, and full-text levels. RESULTS: Two hundred and eight (208) manuscripts were included for analysis. High numbers of studies on patient satisfaction, process outcomes, and barriers were identified, with underrepresentation of health and systems outcomes and impact evaluations. There were very few studies examining hybrid care, planetary health, and use of virtual care with equity-deserving groups. DISCUSSION: This scoping review identified areas of importance for future research, including the use of virtual care in rural and remote regions, inpatient, long-term, and emergency settings, hybrid care, economic and planetary health impacts, and artificial intelligence. As well, enhancing standardization of implementation and evaluation guidelines will optimize quality of care and best practice.

3 July 2026

Read appraisal →
Systematic ReviewEvidence: Weak
20CEBM

Current allergy and asthma reports

Harnessing Machine Learning and Electronic Health Record Data to Improve Asthma Management

PURPOSE OF REVIEW: The review examines the application of machine learning (ML) and large language models (LLMs) to asthma management. We sought to identify clinically relevant applications, and particularly those that harness electronic health record data. We review methodological challenges and future directions for translating these tools into meaningful improvements in asthma care. RECENT FINDINGS: ML applied to electronic health record data has been utilized across several domains of asthma management: predicting medication response to inhaled corticosteroids and biologics, improving inhaler adherence through digital inhaler systems, and predicting exacerbation risk with moderate-to-high accuracy. Tools for patient education include clinician-guided chatbots, as well as publicly available LLMs. Limitations include accuracy, hallucinations, and patient health literacy. ML and LLMs offer promising pathways towards personalized, data-driven asthma management. However, harnessing this potential will require rigorous external validation, transparent model design, equitable implementation, and adaptive clinician oversight. Collaboration among clinicians, data scientists, and policymakers will be essential to implement these tools for patient-centered asthma care.

30 June 2026

Read appraisal →
Systematic ReviewEvidence: Weak
45CEBM

PloS one

Artificial Intelligence in emergency department triage: A scoping review

BACKGROUND: Triage in emergency departments (ED) is a critical process for prioritizing care and ensuring clinical safety. However, current triage systems often exhibit vulnerabilities that compromise the efficiency and quality of healthcare delivery. Artificial Intelligence (AI) has emerged as a promising innovation to support decision-making and optimize patient flow in these high-pressure environments. OBJECTIVE: To map the available evidence regarding the implementation and performance of artificial intelligence in emergency department triage. METHOD: This scoping review followed the Joanna Briggs Institute (JBI) methodology and the PRISMA-ScR guidelines. A comprehensive search was conducted across 13 databases (CINAHL, Cochrane Library, PubMed Central, SciELO, Web of Science, SCOPUS, Science Direct, VHL, Embase, and several regional dissertation repositories), with no language or time restrictions. Two independent reviewers performed the selection process using the Rayyan platform, with discrepancies resolved by a third evaluator. Data were synthesized using the PAGER framework, categorizing findings into Patterns, Advances, Gaps, Evidence for practice, and Recommendations for research. RESULTS: Nineteen studies met the inclusion criteria. AI was primarily implemented through Machine Learning (ML) algorithms, including Deep Learning architectures. Natural Language Processing (NLP) was frequently employed to process unstructured clinical data, with recent studies exploring the potential of Large Language Models (LLMs). Overall, ML-based models consistently outperformed traditional triage systems in predictive accuracy. These techniques were mainly utilized for automated classification, predicting clinical severity, and enhancing patient prioritization by integrating both objective and subjective assessment data. CONCLUSIONS: The findings indicate that AI has significant potential to enhance emergency triage by streamlining service flows and providing robust clinical decision support. However, the current evidence remains heterogeneous and largely exploratory. Key challenges include variability in model performance, a lack of external validation, and studies often limited to specific populations. Consequently, many current tools still lack the necessary reliability for safe, large-scale clinical implementation.

27 June 2026

Read appraisal →
observationalEvidence: Moderate
65CEBM

Proceedings of the National Academy of Sciences of the United States of America

AI agents are sensitive to nudges.

Large language models (LLMs) are increasingly deployed as autonomous agents that make choices and use tools on behalf of users. Yet, we have limited evidence about how their decisions are shaped by their environment. We adapt a human decision-making task to test leading LLMs under four forms of choice architecture: defaults, suggestions, information highlighting, and "optimal" nudges derived from a resource-rational model of human choice. We treat human behavior as a baseline for predictable sensitivity to such interventions. Across models and prompting strategies, LLMs often depart substantially from this baseline. They sometimes pay excessive costs to acquire information, sometimes ignore available information, and, most crucially, are far more responsive to nudges than humans, such that weak cues that slightly shift human behavior have larger effects on model choices, toward both better and worse payoff outcomes. Chain-of-thought prompting and in-context human data do not reliably stabilize behavior. Recent reasoning-optimized LLMs can, in some configurations, restore more human-level sensitivity to nudges, but do so inconsistently and at substantial computational cost. These results point to an important and largely neglected safety concern: LLM agents can be behaviorally brittle under subtle changes in choice architecture, even in the absence of adversarial settings.

23 June 2026

Read appraisal →
Systematic ReviewEvidence: Weak
35CEBM

The Journal of nursing administration

Virtual Nursing Programs in Acute Care Settings: A Scoping Review of Patient, Nurse, and System-Level Outcomes

OBJECTIVE: To synthesize the literature on the influence of virtual nursing (VN) on patient, nurse, and system-level outcomes in acute care. BACKGROUND: Persistent nursing workforce challenges, including staff shortages and turnover, have been intensified by the COVID-19 pandemic, rising patient acuity, and increasing documentation demands. Many health systems have adopted VN programs, where remote nurses support bedside staff using audiovisual technology. Although these programs are rapidly expanding, evidence of their effectiveness remains limited. METHODS: These authors conducted a scoping review following PRISMA guidelines. Eleven studies reporting patient, nurse, or system-level outcomes were included. RESULTS: Most studies were cross-sectional pilots. Evidence was strongest for nurse and patient satisfaction, with reports of improved discharge efficiency, reduced administrative burden, and higher patient satisfaction. Findings for other outcomes, including safety indicators and financial metrics, were inconsistently reported. CONCLUSIONS: Virtual nursing shows promise for enhancing patient satisfaction and workflow efficiency, but cannot replace investments in sufficient bedside staff and resources.

21 June 2026

Read appraisal →