Research Appraisals
Evidence-based critical appraisals of the latest medical research, systematically evaluated using Oxford CEBM methodology.
Showing 20 appraisals
International dental journal
Accuracy of Large Language Models in Answering Dental Examination Questions: A Systematic Review and Meta-Analysis
INTRODUCTION: Large language models (LLMs), including OpenAI's GPT family accessed via interfaces such as ChatGPT and Microsoft Copilot, as well as non-GPT systems such as Google Gemini, are increasingly applied in healthcare and dental education. However, the accuracy of these systems in specialized tasks such as answering dental examination questions remains unclear. METHODS: This systematic review and meta-analysis evaluated LLM performance in answering dental questions. Databases searched were PubMed, Embase, Scopus, and Web of Science. Data on question type and number, LLM versions, and accuracy rates were extracted. Pooled accuracy was estimated using a random-effects model; heterogeneity and publication bias were assessed. RESULTS: A total of 39 studies were included, with ChatGPT-4 being the most frequently evaluated model. The pooled accuracy for LLMs was 63.7% (95% CI: 60.3%-67.1%), with high heterogeneity (I² = 91.5%). Subgroup analysis revealed ChatGPT-4 and Copilot (a GPT-based interface) achieved the highest pooled accuracies (∼73% and ∼75%, respectively). Direct comparisons confirmed ChatGPT-4 significantly outperformed earlier versions and some competitor models. Sensitivity analyses supported the robustness of findings. CONCLUSION: LLMs demonstrate moderate accuracy in answering dental examination questions and are currently insufficient for autonomous clinical decision-making. When their limitations are explicitly recognized, however, these systems may serve as valuable adjuncts in dental education and examination preparation. Methodological strategies such as structured prompting and retrieval-augmented approaches warrant further investigation but were not the primary focus of the present analysis.
2 Aug 2026
Read appraisal →Bulletin of the World Health Organization
Digital epidemiology investments, Saudi Arabia
PROBLEM: Traditional epidemiological surveillance methods are often limited by delays in reporting and fragmented data systems. Saudi Arabia faces additional public health challenges from mass gatherings during Hajj and Umrah, an increasing burden of noncommunicable diseases and rapid urbanization, highlighting the need for investing in digital epidemiology. APPROACH: Saudi Arabia has accelerated digital transformation in health care through Vision 2030 initiatives. The strategies include health information exchange platforms, analytics driven by artificial intelligence, telemedicine services and digital monitoring systems used during Hajj. We review current initiatives to invest in digital epidemiology in Saudi Arabia, implementation challenges and policy priorities. LOCAL SETTING: Saudi Arabia's health system operates under a predominantly public model. The health ministry is the main provider, regulator and finance provider of most health-care services. Health-care coverage is nearly universal, with citizens receiving services free of charge through the public system. Ongoing reforms aim to gradually decentralize certain functions. RELEVANT CHANGES: The initiatives under Vision 2030 have supported disease surveillance, data integration and public health response capacities. Existing digital health reforms have created a foundation for integrating digital epidemiology into routine public health practice. However, challenges remain, including fragmented interoperability between institutions, workforce shortages, unequal digital access, and concerns about data governance, privacy and algorithmic bias. LESSONS LEARNT: Saudi Arabia's experience suggests that digital epidemiology is more effective when integrated within broader digital health reforms. Successful implementation requires not only digital infrastructure, but also workforce development, ethical governance, transparency and mechanisms for integrating digital data into public health decision-making. PROBLÈME: Les méthodes traditionnelles de surveillance épidémiologique sont souvent limitées par des retards dans la transmission des données et par la fragmentation des systèmes de données. L’Arabie saoudite est confrontée à des défis supplémentaires en matière de santé publique liés aux rassemblements de masse lors du Hajj et de la Omra, au fardeau croissant des maladies non transmissibles et à l’urbanisation rapide, ce qui souligne la nécessité d’investir dans l’épidémiologie numérique. APPROCHE: L’Arabie saoudite a accéléré la transformation numérique dans le domaine des soins de santé grâce aux initiatives de la Vision 2030. Ces stratégies comprennent des plateformes d’échange d’informations de santé, des analyses alimentées par l’intelligence artificielle, des services de télémédecine et des systèmes de surveillance numérique employés pendant le Hajj. La présente étude passe en revue les initiatives actuelles visant à investir dans l’épidémiologie numérique en Arabie saoudite, les défis liés à leur mise en œuvre et les priorités politiques. ENVIRONNEMENT LOCAL: Le système de santé saoudien fonctionne selon un modèle majoritairement public. Le ministère de la Santé est le principal prestataire, régulateur et bailleur de fonds de la plupart des services de santé. La couverture des soins de santé est pratiquement universelle, les citoyens bénéficiant de services gratuits au travers du système public. Les réformes en cours visent à décentraliser progressivement certaines fonctions. CHANGEMENTS SIGNIFICATIFS: Les initiatives menées dans le cadre de la Vision 2030 ont favorisé la surveillance des maladies, l’intégration des données et le renforcement des capacités d’intervention de santé publique. Les réformes existantes en matière de santé numérique ont jeté les bases nécessaires à l’intégration de l’épidémiologie numérique dans les pratiques courantes de santé publique. Toutefois, des défis subsistent, notamment le manque d’interopérabilité entre les établissements, la pénurie de main-d’œuvre, les inégalités d’accès au numérique, ainsi que des préoccupations en matière de gouvernance des données, de confidentialité et de biais dû aux algorithmes. LEÇONS TIRÉES: L’expérience de l’Arabie saoudite suggère une plus grande efficacité de l’épidémiologie numérique lorsque celle-ci s’inscrit dans le cadre de réformes plus larges en matière de santé numérique. La réussite de la mise en œuvre nécessite non seulement une infrastructure numérique, mais aussi le développement des ressources humaines, une gouvernance éthique, la transparence ainsi que des mécanismes d’intégration des données numériques dans la prise de décision en matière de santé publique. SITUACIÓN: Los métodos tradicionales de vigilancia epidemiológica suelen verse limitados por retrasos en la notificación y por sistemas de datos fragmentados. Arabia Saudí se enfrenta además a desafíos de salud pública derivados de las concentraciones masivas de personas durante el Hach y la Umrah, del aumento de la carga de las enfermedades no transmisibles y de la rápida urbanización, lo que destaca la necesidad de invertir en epidemiología digital. ENFOQUE: Arabia Saudí ha acelerado la transformación digital de la atención sanitaria mediante las iniciativas de la Visión 2030. Las estrategias incluyen plataformas de intercambio de información sanitaria, análisis basados en inteligencia artificial, servicios de telemedicina y sistemas digitales de vigilancia utilizados durante el Hach. Se examinan las iniciativas actuales de inversión en epidemiología digital en Arabia Saudí, los desafíos para su implementación y las prioridades en materia de políticas. MARCO REGIONAL: El sistema sanitario de Arabia Saudí funciona bajo un modelo predominantemente público. El ministerio de salud es el principal proveedor, regulador y financiador de la mayoría de los servicios sanitarios. La cobertura sanitaria es prácticamente universal y la ciudadanía recibe servicios gratuitos a través del sistema público. Las reformas en curso tienen por objeto descentralizar gradualmente determinadas funciones. CAMBIOS IMPORTANTES: Las iniciativas de la Visión 2030 han fortalecido la vigilancia de enfermedades, la integración de datos y la capacidad de respuesta en salud pública. Las reformas existentes en materia de salud digital han creado una base para integrar la epidemiología digital en la práctica habitual de la salud pública. No obstante, persisten desafíos, entre ellos la interoperabilidad fragmentada entre instituciones, la escasez de personal, las desigualdades en el acceso digital y las preocupaciones relacionadas con la gobernanza de los datos, la privacidad y los sesgos algorítmicos. LECCIONES APRENDIDAS: La experiencia de Arabia Saudí sugiere que la epidemiología digital es más eficaz cuando se integra en reformas más amplias de salud digital. Su implementación satisfactoria requiere no solo infraestructura digital, sino también desarrollo de capacidades del personal, gobernanza ética, transparencia y mecanismos para incorporar los datos digitales a la toma de decisiones en salud pública. المشكلة: غالبًا ما تكون طرق الرصد الوبائي التقليدية محدودة بسبب التأخيرات في الإبلاغ وأنظمة البيانات المشتتة. وتواجه المملكة العربية السعودية تحديات إضافية في مجال الصحة العامة نتيجة للتجمعات الكبيرة خلال موسمي الحج والعمرة، وتزايد عبء الأمراض غير المعدية، والتوسع الحضري السريع، مما يُبرز الحاجة إلى الاستثمار في علم الأوبئة الرقمي. الأسلوب: سارعت المملكة العربية السعودية في التحول الرقمي في مجال الرعاية الصحية من خلال مبادرات رؤية 2030. وتشمل هذه الاستراتيجيات منصات تبادل المعلومات الصحية، والتحليلات المدعومة بالذكاء الاصطناعي، وخدمات العلاج الطبي عن بُعد، وأنظمة المراقبة الرقمية المستخدمة خلال موسم الحج. نستعرض في هذا البحث المبادرات الحالية للاستثمار في علم الأوبئة الرقمي في المملكة العربية السعودية، وتحديات التنفيذ، وأولويات السياسة الحكومية. المواقع المحلية: يعمل النظام الصحي في المملكة العربية السعودية بموجب نموذج حكومي بشكل رئيسي. وتُعد وزارة الصحة هي المزود الرئيسي، والجهة المنظمة، والممول لمعظم خدمات الرعاية الصحية. وتعد تغطية الرعاية الصحية شاملة تقريبًا، حيث يحصل المواطنون على الخدمات مجانًا من خلال النظام الحكومي. وتهدف الإصلاحات الجارية إلى تطبيق اللامركزية تدريجيًا على وظائف محددة. التغيّرات ذات الصلة: إن المبادرات بموجب رؤية 2030 تدعم رصد الأمراض، وتكامل البيانات، وقدرات الاستجابة في مجال الصحة العامة. لقد أدت إصلاحات الصحة الرقمية إلى إنشاء أساس لدمج علم الأوبئة الرقمي في ممارسات الصحة العامة الروتينية. ومع ذلك، لا تزال هناك تحديات، تشمل التشغيل المشترك المشتت بين المؤسسات، ونقص الكوادر، والوصول غير العادل إلى البيانات الرقمية، والمخاوف المتعلقة بحوكمة البيانات، والخصوصية، والتحيز الخوارزمي. الدروس المستفادة: تشير تجربة المملكة العربية السعودية إلى أن علم الأوبئة الرقمي يكون أكثر فعالية عند دمجه ضمن إصلاحات أوسع نطاقًا في مجال الصحة الرقمية. ويتطلب التنفيذ الناجح ليس فقط بنية تحتية رقمية، بل أيضًا تطوير القوى العاملة، والحوكمة الأخلاقية، والشفافية، وآليات دمج البيانات الرقمية في عملية صنع القرار في مجال الصحة العامة. 问题: 传统流行病学监测手段常受制于报告滞后与数据系统碎片化的问题。沙特阿拉伯还面临朝觐与副朝期间大型聚众活动、非传染性疾病负担攀升以及快速城市化带来的额外公共卫生难题,凸显出投资发展数字流行病学的必要性。. 方法: 沙特阿拉伯依托《2030 愿景》(Vision 2030) 相关举措,加快推进医疗健康领域数字化转型。相关举措包括医疗信息交互平台、人工智能驱动的数据分析、远程医疗服务以及朝觐期间启用的数字化监测系统。本文梳理了沙特阿拉伯的当前数字流行病学投资项目、项目实施方面的挑战及政策优先发展方向。. 当地状况: 沙特阿拉伯的医疗卫生体系以公立运营模式为主。卫生部是大多数医疗护理服务的主要提供方、监管机构和出资主体。医疗卫生保障基本实现全民覆盖,本国公民可通过公立医疗体系免费就医。现行改革旨在逐步下放部分职能。. 相关变化: 《2030 愿景》框架下的各项举措,助力完善疾病监测、数据整合与公共卫生应急处置能力建设。现有的数字医疗改革为将数字流行病学融入常规公共卫生工作奠定了基础。然而,仍存在一些挑战:医疗机构间数据互通割裂、人才短缺、数字资源获取不均衡,以及数据治理、隐私安全与算法偏见等隐患问题。. 经验教训: 沙特阿拉伯的实践经验表明,将数字流行病学融入大范围数字医疗改革体系中,其实施成效更佳。实施成功不仅需要数字化基础设施,还需人才队伍建设、合规伦理治理、信息公开制度,以及把数字化数据纳入公共卫生决策的配套机制。. ПРОБЛЕМА: Традиционные методы эпидемиологического надзора часто ограничены задержками в предоставлении отчетности и фрагментированностью систем данных. Саудовская Аравия сталкивается с дополнительными проблемами общественного здравоохранения, связанными с массовыми скоплениями людей во время хаджа и умры, растущим бременем неинфекционных заболеваний и стремительной урбанизацией. Все это подчеркивает необходимость инвестиций в цифровую эпидемиологию. ПОДХОД: Саудовская Аравия ускорила цифровую трансформацию здравоохранения в рамках инициатив программы Vision 2030. Среди ключевых направлений – платформы для обмена медицинской информацией, аналитические системы на основе искусственного интеллекта, телемедицинские сервисы и цифровые системы мониторинга, используемые во время хаджа. Авторы рассматривают текущие инвестиционные инициативы в области цифровой эпидемиологии в Саудовской Аравии, проблемы внедрения и приоритетные направления политики. МЕСТНЫЕ УСЛОВИЯ: Система здравоохранения Саудовской Аравии опирается преимущественно на государственную модель. Министерство здравоохранения является основным поставщиком медицинских услуг, регулятором и главным источником финансирования большинства медико-санитарных услуг. Охват медицинской помощью практически всеобщий: граждане получают услуги бесплатно через государственную систему. Текущие реформы направлены на постепенную децентрализацию отдельных функций. ОСУЩЕСТВЛЕННЫЕ ПЕРЕМЕНЫ: Инициативы в рамках программы Vision 2030 способствовали развитию эпидемиологического надзора, интеграции данных и укреплению потенциала реагирования системы общественного здравоохранения. Проводимые реформы в области цифрового здравоохранения заложили фундамент для внедрения цифровой эпидемиологии в повседневную практику общественного здравоохранения. Однако сохраняются проблемы, включая недостаточную совместимость информационных систем разных учреждений, нехватку кадров, неравномерный доступ к цифровым технологиям, а также вопросы управления данными, конфиденциальности и алгоритмической предвзятости. ВЫВОДЫ: Опыт Саудовской Аравии показывает, что цифровая эпидемиология наиболее эффективна, когда интегрирована в более широкие реформы цифрового здравоохранения. Для успешного внедрения необходима не только цифровая инфраструктура, но и развитие кадрового потенциала, этическое регулирование, прозрачность и механизмы интеграции цифровых данных в процессы принятия решений в сфере общественного здравоохранения.
2 Aug 2026
Read appraisal →JMIR research protocols
Large Language Models in German Continuing Medical Education Assessments: Protocol for a Fully Crossed Experimental Study
BACKGROUND: Continuing medical education (CME) is a legal and ethical obligation for physicians in Germany. The rapid rise of large language models (LLMs) such as ChatGPT, Gemini, Claude, and Grok raises concerns about the integrity of CME assessments, as LLMs can already pass German CME tests. OBJECTIVE: This study aims to determine whether the choice of document format (searchable PDF, protected PDF, raster PDF, or vector PDF) and LLM influences the ability of LLMs to solve CME test questions at rates exceeding the passing threshold specified for each CME module (typically 70%). METHODS: In a fully crossed within-subjects repeated-measures design, 18 expired CME articles from 3 major German publishers across 6 specialties will be converted into 3 cheating-impeding PDF formats and processed alongside the original PDF files by 4 current LLMs (GPT-5, Claude Sonnet 4, Grok-4, and Gemini 3). This results in 16 model-format combinations. Each model will answer every article 3 times per file-format condition, with outcomes derived from aggregated run-level results. The primary outcome is the proportion of correctly answered questions; the secondary outcome is the pass/fail rate. RESULTS: The study has been approved by the Witten/Herdecke University Ethics Committee (S-260/2025; dated August 10, 2025) and is preregistered at the Open Science Framework. The study is supported by internal departmental resources only, and no external funding was received. Because this protocol evaluates LLMs using expired CME materials, no human participants are being recruited. Data collection is planned to begin in June 2026 and is expected to last approximately 4 weeks. At the time of manuscript submission, no data have been collected or analyzed. Results are expected to be available after the completion of data collection and statistical analysis in 2026. The analyses will quantify performance differences across document formats; these findings may inform the feasibility of nonsearchable document formats as a temporary measure to reduce LLM-enabled cheating risks in CME contexts. CONCLUSIONS: By quantifying how document format constrains LLM performance, this study aims to evaluate simple technical safeguards that may reduce artificial intelligence-assisted manipulation of CME tests and inform regulators and CME providers about how to balance assessment validity, accessibility, and responsible LLM integration into postgraduate medical education.
30 July 2026
Read appraisal →Proceedings of the National Academy of Sciences of the United States of America
Legal infrastructure for transformative AI governance
Most of our AI governance efforts focus on substance: What rules do we want in place? What limits or checks do we want to impose on AI development and deployment? But a key role for law is not only to establish substantive rules but also to establish legal and regulatory infrastructure to generate and implement rules. The transformative nature of AI calls especially for attention to building legal and regulatory frameworks. In this Perspective, I review three examples: the creation of registration regimes for frontier models; the creation of registration and identification regimes for autonomous agents; and the design of regulatory markets to facilitate a role for private companies to innovate and deliver AI regulatory services.
29 July 2026
Read appraisal →Proceedings of the National Academy of Sciences of the United States of America
The backfiring effect of weak AI safety regulation
Recent policy proposals aim to improve the safety of general-purpose AI, but there is little understanding of the efficacy of different regulatory approaches. We present a strategic model that explores interactions between safety regulation, general-purpose AI technology creators, and domain specialists-those who adapt the technology for specific applications. Our analysis examines how regulatory measures targeting different parts of the AI development chain affect the outcome of this game. Our model assumes AI technology is characterized by two key attributes: safety and performance. The regulator first sets a minimum safety requirement that applies to one or both players. The general-purpose creator then invests in the technology, establishing its initial safety and performance levels. Next, domain specialists refine the AI for their use cases, updating the safety and performance levels and taking the product to market. Resulting revenue is shared between the specialist and generalist. Our analysis reveals two insights: first, weak safety regulation imposed predominantly on domain specialists can backfire. While it might seem logical to regulate AI use cases, our analysis shows that weak regulations targeting domain specialists alone can reduce safety in a large class of parameterizations. Second, in contrast to the previous finding, we observe that stronger, well-placed regulation can mutually benefit all players. When regulators impose appropriate safety standards on both general-purpose AI creators and domain specialists, the regulation can function as a commitment device, leading to safety and performance gains, surpassing what is achievable under no regulation or regulating only one player.
28 July 2026
Read appraisal →Medicine
A bibliometric analysis of global trends in AI-driven digital health technologies for diabetes management
BACKGROUND: Digital health technologies are increasingly applied in diabetes care, enabling continuous monitoring, personalized support and remote interventions. Meanwhile, artificial intelligence (AI) is enhancing the precision and effectiveness of these tools. This study aims to map global research trends and thematic developments in AI-driven digital health technologies for diabetes management and to explore their future directions. METHODS: We collected data from the Web of Science Core Collection, including articles and reviews published up to July 12, 2025, using CiteSpace, VOSviewer, and Microsoft Excel to analyze countries/regions, institutions, journals, references, authors, and keywords. RESULTS: A total of 673 publications were included in the analysis. Global publications on AI-driven digital health technologies for diabetes increased steadily, with the USA leading in output. The University of London ranked as the most productive institution. Sensors and diabetes care were the most frequently published and cited journals in this field. Herrero P was among the most prolific authors. The most cited article was "Development and Validation of a Deep Learning Algorithm for Detection of Diabetic Retinopathy in Retinal Fundus Photographs." "diabetes" was the most frequently occurring keyword. Keyword cluster analysis identified 3 primary research hotspots: AI-enabled monitoring, digital health interventions, and AI-based diabetic retinopathy screening. CONCLUSIONS: This study summarizes the evolution of AI-driven digital health technologies in diabetes care. Although challenges remain in data security, standardization and validation, these technologies hold increasing potential for accurate diagnosis, real-time monitoring and personalized care.
26 July 2026
Read appraisal →International journal for equity in health
A conceptual framework for measuring AI health equity
BACKGROUND: Artificial intelligence (AI) is increasingly embedded in health systems globally and has the potential to improve efficiency, diagnostic accuracy, and decision support. However, its benefits remain unevenly distributed, particularly in low- and middle-income countries (LMICs). Models developed using datasets from specific populations may perform poorly in other settings, reinforcing structural inequities rather than correcting them. OBJECTIVE: This viewpoint proposes a composite framework, the AI in Healthcare Equity Index (AIHEI), to support measurable assessment of equity in health AI systems. FRAMEWORK: The AIHEI is designed to assess equity across five domains: data representation, algorithmic fairness, transparency and explainability, governance and oversight, and community impact and benefit sharing. By generating a standardised score, the index could enable comparisons across technologies, incentivise improvement, and support regulation, procurement, publication, and funding decisions. IMPLICATIONS: Pilots across diverse health domains and geographic settings are needed to assess feasibility, refine domain weighting, and evaluate reliability, reproducibility, and validity. Important challenges include contextual definitions of fairness, data sovereignty, post-deployment monitoring, and the risk of metric gaming. CONCLUSIONS: Quantifying equity in health AI is essential to ensure that AI does not create, widen, or exacerbate existing disparities by neglecting underserved populations. A common, objective measure of AI-related health equity can help move the field from ethical aspiration toward measurable accountability, monitoring, and enforcement.
25 July 2026
Read appraisal →Current biology : CB
Looking to the brain to improve energy efficiency of AI
Modern artificial intelligence (AI) systems have achieved remarkable capabilities, but at an extraordinary energy cost. Training and running large-scale models can consume vast resources, posing environmental, economic, and societal challenges. In contrast, biological brains perform lifelong learning, adaptive control, and flexible reasoning using orders of magnitude less energy for learning and adaptation over a lifetime. What accounts for this difference - and how can it guide future AI development? In this review, we identify key biological principles that support energy-efficient capacities in biological brains, and consider how they might inform the design of more sustainable artificial systems. We organize our analysis around three domains: architectural constraints, signaling strategies, and learning algorithms. In each domain, we discuss concrete observations from biology, from cell to circuit to cognitive level, and describe how current and emerging AI systems mirror or diverge from these motifs. One striking feature of biological energy optimization is often overlooked: that brains are remarkably stable in their energy usage across heterogeneous modes, suggesting they may minimize energy needs during active environmental processing through maximizing the utility of 'rest-like' background processes. Overall, rather than advocating for biomimicry for its own sake, we argue for biologically informed engineering. Understanding how natural systems minimize energetic cost while maximizing flexibility may help us build AI that is not only powerful, but also efficient, equitable, and environmentally responsible.
21 July 2026
Read appraisal →Science (New York, N.Y.)
AI in scientific publishing: Slower, worse, and more expensive.
There's a saying in the management world, popularized by NASA administrator Daniel Goldin in the 1990s, that the goal of technological improvements is to make products faster, better, and cheaper. Although this strategy had some success in the aerospace industry, the zealots of artificial intelligence (AI) have been making the same argument regarding how it will transform work, claiming that so little human effort will be required that humanity will enter an era of radical abundance, free from disease, drudgery, and danger, among other benefits, leaving society with more time for creative pursuits. But history tells a different story. When machines began to increase productivity during the second industrial revolution, American engineer Frederick Winslow Taylor's The Principles of Scientific Management encouraged corporations to use surveillance to get employees to work harder and longer, an approach that exhausted and discouraged workers and led to the transfer of knowledge and any decision-making from workers to management, while enriching the profits for only those at the top. Yet, it remains foundational to the American economic enterprise. Indeed, scientific publishing is starting to experience some Taylorism with the insertion of AI. Rigorous human checking of AI-generated research papers is creating bottlenecks as publishers strive to maintain the integrity of the scientific record. The challenge is requiring even more human effort, making the whole endeavor slower and more expensive.
17 July 2026
Read appraisal →Scientific reports
Differentially private federated learning for localized control of infectious disease dynamics
In times of epidemics, swift reaction is necessary to mitigate epidemic spreading. For this reaction, localized approaches have several advantages, limiting necessary resources and reducing the impact of interventions on a larger scale. However, training a separate machine learning (ML) model on a local scale is often not feasible due to limited available data. Centralizing the data is also challenging because of its high sensitivity and privacy constraints. In this study, we consider a localized strategy based on the German counties and communities managed by the related local health authorities (LHA). For the preservation of privacy to not oppose the availability of detailed situational data, we propose a privacy-preserving forecasting method that can assist public health experts and decision makers. ML methods with federated learning (FL) train a shared model without centralizing raw data. Considering the counties, communities or LHAs as clients and finding a balance between utility and privacy, we study a FL framework with client-level differential privacy (DP). We train a shared multilayer perceptron on sliding windows of recent case counts to forecast the number of cases in the future, while clients exchange only norm-clipped updates and the server aggregates updates with DP noise. We evaluate the approach on COVID-19 data on county-level during two phases: November 2020 and March 2022 (Omicron). As expected, very strict privacy ([Formula: see text]) yields unstable, unusable forecasts. At a moderately strong but still privacy-preserving level ([Formula: see text]), the DP model closely approaches the non-DP model: [Formula: see text] (vs. 0.96) and mean absolute percentage error (MAPE) [Formula: see text] in November 2020; [Formula: see text] (vs. 0.90) and MAPE [Formula: see text] in March 2022. Overall, our results support the feasibility of privacy-preserving collaboration among health authorities for local forecasting. In the evaluated COVID-19 phases, client-level DP-FL delivered useful county-level predictions with formal privacy guarantees under the stated threat model. The appropriate privacy budget should nevertheless be re-evaluated for other epidemic phases and applications.
8 July 2026
Read appraisal →Journal of medical Internet research
Exploring the Narratives of Patients With Cancer Using Large Language Models: Topic Modeling and Social Network Analysis
BACKGROUND: Patients with cancer often experience diverse psychosocial stressors that profoundly affect disease trajectories, treatment adherence, and overall quality of life. Understanding how patients experience and articulate these issues is critical for designing patient-centered interventions. Conventional data collection methods, such as surveys and interviews, provide depth but are constrained by recall bias and scalability and may overlook sensitive or underreported concerns. Patient-authored narratives in online health communities present a valuable opportunity to identify prevalent and underserved issues. However, critical analytic challenges remain in generating coherent and interpretable insights due to their unstructured and large-scale nature. OBJECTIVE: This study aims to leverage TopicGPT, a prompt-based topic modeling framework powered by large language models (LLMs), in combination with network analysis for interpretable topic discovery and interrelationship analysis in the narratives of patients with cancer. METHODS: Patient-authored posts describing psychosocial challenges about cancer experience were collected from 4 online health communities. Eligible posts were preprocessed and analyzed using TopicGPT, wherein topics were generated hierarchically and mapped at the sentence level. Comparison analyses were conducted among 3 state-of-the-art LLMs through cosine similarity and manual evaluation. Results from the best-performing LLM were further compared with 2 conventional topic models through topic diversity and were used to construct the network subsequently. Topic co-occurrence was examined using the pointwise mutual information algorithm and centrality metrics to reveal influential topics and thematic interconnections across narratives. RESULTS: A total of 11,306 posts were collected from Reddit, Macmillan, Mijian, and Douban between December 6, 2006, and September 24, 2025. Of these, 3169 posts were retained for topic modeling and network analysis. DeepSeek-V3.2 consistently outperformed Gemini-2.5-Flash and GPT-4o, with similarity scores of 0.6295, 0.5342, and 0.5247, respectively. TopicGPT maintained consistently high topic diversity across languages. "Fear of cancer recurrence" and "Psychological distress" emerged as both most frequent and bridging topics across a hierarchy comprising 42 top-level and 58 subtopics. Strong connections were observed among "Sexual health concerns," "Reproductive concerns," and "Quality of life impact"; "Family communication concerns" frequently co-occurred with "Employment concerns," "Diagnostic delays and misdiagnosis," and "Social support." CONCLUSIONS: This study demonstrates the potential of LLM-based topic modeling for large-scale, context-sensitive analysis of patient-authored narratives. The proposed integrated, domain-adaptable pipeline enables the identification of high-fidelity topics and their interrelationships, offering a scalable and interpretable approach to qualitative data in health care. Importantly, our findings reveal substantial concerns and unmet needs among patients with cancer, with potential to support patient-centered research and inform future clinical assessment and supportive care strategies.
7 July 2026
Read appraisal →Philosophy, ethics, and humanities in medicine : PEHM
Moral diversity and the challenge of responsibility in AI-CDSS
The increasing integration of artificial intelligence in clinical decision support systems (AI-CDSS) has fueled expectations of more personalized and effective diagnostics and therapies. By incorporating machine learning methods, AI-CDSS promise enhanced predictive accuracy, improved stratification, and innovative individualized care. However, this technological optimism is accompanied by complex ethical challenges, including issues of explainability, trust, autonomy, and data security. At the core of these debates lies the question of responsibility, which involves both its attribution and diffusion, as well as the underlying normative standards guiding moral action. In the context of healthcare practice, responsibility is further complicated by moral diversity-the coexistence of varying moral values, cultural beliefs, and ethical frameworks among healthcare professionals, patients, and institutional stakeholders. This plurality challenges the establishment of a unified normative standard necessary for ethically sound responsibility attribution. This paper offers an analysis of moral diversity and AI-CDSS as a challenge for responsibility in healthcare environments. Using a relational concept of responsibility the study examines key areas in which moral diversity affects responsibility in AI-mediated decision-making. This includes algorithmic bias, healthcare professional and patient interaction and the role of patients. Through these examples, the paper explains how different normative standards intensify ethical complexity in AI-supported clinical contexts. It argues that greater ethical sensitivity to moral diversity is essential-both in the development of AI-CDSS and in their application within morally value-laden healthcare situations.
4 July 2026
Read appraisal →Behavior research methods
A validity-guided workflow for robust large language model research in psychology.
Large language models (LLMs) are rapidly being integrated into psychological and behavioral research as research tools, evaluation targets, human simulators, and cognitive models. Yet recent evidence reveals severe measurement unreliability: personality assessments degenerate under factor analysis, moral preferences reverse with punctuation changes, and theory-of-mind accuracy varies widely with trivial rephrasing. These "measurement phantoms"-statistical artifacts masquerading as psychological phenomena-threaten the validity of a growing body of research. Guided by the dual-validity framework that integrates psychometrics with causal inference, this article presents a six-stage workflow that scales validity requirements to research ambition-using LLMs to code text requires basic reliability and accuracy, whereas claims about psychological properties demand comprehensive construct validation. Researchers must (1) explicitly define their research goal and corresponding validity requirements, (2) develop and validate computational instruments through psychometric testing, (3) design experiments that control for computational confounds, (4) execute protocols transparently, (5) analyze data with methods appropriate for nonindependent observations, and (6) report findings within boundaries and use results to refine theory. The workflow is illustrated through an example of model evaluation-"LLM selfhood"-showing how systematic validation can distinguish genuine computational phenomena from measurement artifacts. By establishing validated computational instruments and transparent practices, this workflow provides a path toward building a robust empirical foundation for AI psychology research.
3 July 2026
Read appraisal →Magnetic resonance in medical sciences : MRMS : an official journal of Japan Society of Magnetic Resonance in Medicine
Integrating Artificial Intelligence into Prostate MR Imaging: Technical Foundations, Clinical Applications, and Workflow Implications
Prostate MRI has become a cornerstone of contemporary prostate cancer diagnosis, enabling improved detection of clinically significant disease while reducing unnecessary biopsies and overtreatment. However, prostate MRI remains technically demanding, time-consuming, and subject to inter-reader variability, particularly as healthcare systems move toward abbreviated protocols such as non-contrast MRI (biparametric MRI). In this context, artificial intelligence (AI) has emerged as a promising tool to enhance image quality, diagnostic consistency, and workflow efficiency across the prostate MRI pathway. This non-systematic narrative review provides a comprehensive overview of the technical foundations, clinical applications, and workflow implications of AI integration into prostate MRI. It summarizes key concepts in machine learning and deep learning relevant to prostate imaging and reviews current evidence supporting AI-based solutions for image quality assessment and reconstruction, automated prostate segmentation, lesion detection, and risk stratification. Particular attention is given to human-AI collaboration models, the role of AI in supporting equivocal lesions, and the integration of imaging with clinical variables for personalized risk estimation. In addition, it discusses the impact of AI on reporting efficiency, training, and standardization, as well as the current landscape of commercially available AI tools. Despite encouraging results from large multicenter studies, important challenges remain, including heterogeneity in study design, limited prospective validation, generalizability across institutions, and ethical and regulatory considerations. Overall, AI should be regarded as a complementary decision-support technology rather than a replacement for radiologists. Thoughtful implementation, robust validation, and appropriate user training are essential to ensure that AI meaningfully enhances the quality, efficiency, and reliability of prostate MRI-based care.
2 July 2026
Read appraisal →Journal of medical Internet research
Operationalizing Digital Health Equity in Artificial Intelligence-Enabled Patient Decision Aids for Older Adults: Mixed Methods Study
BACKGROUND: Artificial intelligence-enabled patient decision aids (AI-PDAs) hold promise for supporting older adults with chronic diseases in accessing personalized health information, clarifying preferences, and engaging in shared decision-making. Achieving equity in their design requires attention to the complex health care and digital contexts in which these tools are used. While the Digital Health Equity Framework (DHEF) provides a conceptual foundation, practical strategies for its application remain limited. OBJECTIVE: This study aimed to identify equity-related determinants and generate actionable design strategies for applying the DHEF to AI-PDAs for older adults. METHODS: A mixed methods study was conducted. Semistructured interviews were conducted with older adults living with hypertension and/or diabetes, health care providers, and medical students to explore equity determinants relevant to AI-PDAs. In parallel, a review of reviews synthesized existing evidence on approaches to addressing these determinants. Interview findings and review findings were integrated through an iterative mapping process conducted by the research team and refined through multidisciplinary expert consultation involving medicine, public health, social services, and computer science. RESULTS: A total of 33 stakeholders were interviewed, including 15 older adults, 8 health care providers, and 10 medical students. Thirteen reviews were included in the umbrella review. The integrated synthesis identified equity determinants spanning individual, interpersonal, community, and societal levels across both the health care and digital environments, together with cross-level concerns related to algorithmic fairness. These findings informed 5 recommendations for equitable AI-PDA development: (1) co-design with end users to address their needs, (2) embrace relationship-centered design, (3) leverage community resources to improve support, (4) promote accessible and equitable artificial intelligence (AI) governance in society, and (5) enhance equitable AI through algorithmic fairness. Together, these recommendations provide practical guidance for design, pilot testing, implementation, and evaluation. CONCLUSIONS: By integrating stakeholder perspectives with synthesized review evidence, this study extends the DHEF from a primarily conceptual framework toward a more practice-oriented approach for AI-PDAs for older adults with chronic disease. Health care settings serve as a mediating sociotechnical context where AI tools may either support or constrain equitable care participation. The findings underscore the need for interdisciplinary collaboration to align technological innovation with equity-oriented design. Future work should focus on co-designed prototypes, real-world testing, and measurable equity outcomes.
30 June 2026
Read appraisal →PloS one
Forecasting COVID-19 new cases using NBEATS deep learning and mobility data
COVID-19 is a highly contagious disease transmitted primarily through human contact. Therefore, understanding population mobility is essential for predicting COVID-19 case trends. In this paper, we propose a novel deep learning approach for forecasting new COVID-19 cases using a neural architecture called Neural Basis Expansion Analysis for Interpretable Time Series (N-BEATS). The N-BEATS model effectively handles long input sequences and large output horizons without information loss or increased computational complexity. We compare the performance of N-BEATS with a state-of-the-art benchmark model, LSTM-Markov, across four major countries: the United States, the United Kingdom, Russia, and Brazil. Three distinct COVID-19 datasets from Google, Apple, and Our World in Data (OWID) were used in this study. Incorporating Google and Apple mobility data as covariates enhances both the accuracy and interpretability of the N-BEATS model. Our results show that N-BEATS consistently outperforms LSTM-Markov across all datasets and countries, consistently yielding lower Root Mean Squared Error (RMSE) and Mean Absolute Percentage Error (MAPE). Furthermore, the N-BEATS model with covariates outperforms its counterpart without covariates, indicating that mobility data provide substantial value for forecasting new COVID-19 cases. Overall, this study demonstrates the effectiveness of the N-BEATS architecture in capturing pandemic dynamics and offers valuable insights for policymakers and public health officials in managing future outbreaks.
30 June 2026
Read appraisal →Human genetics
AI in variant analysis: fast track to genetic diagnoses
While falling costs have expanded access to genomic sequencing, clinical utility is frequently hindered by the challenge of interpreting complex genetic data. Variant analysis for rare disease patients especially requires significant time and expertise, creating a bottleneck that delays diagnostics. Although advances in genetic variant classification have improved diagnostic precision, they have also increased the identification of variants of uncertain significance (VUSs), widening the interpretation gap between data generation and clinical actionability. The high prevalence of VUSs can lead to false reassurance or psychological distress by misinterpretting inconclusive results. We propose that artificial intelligence (AI) is a critical clinical decision-support tool for bridging this gap, offering a scalable framework to optimize variant interpretation and shorten the diagnostic odyssey. While reclassification ultimately requires biological evidence that AI cannot replace, these tools serve as essential aggregators and prioritizers, especially as guidelines transition toward the upcoming quantitative ACMG v4 framework. We advocate integrating AI throughout the genetic diagnostic workflow-from initial phenotyping to variant prioritization-to facilitate data-driven, personalized treatment. We outline current AI-assisted approaches and discuss anticipated challenges in this pursuit, such as privacy, training data bias and quality, model explainability, and the necessity of a total product life cycle for validation. To address these challenges, we provide recommendations for "human-in-the-loop" design and intuitive workflow integration to ensure AI tools meet the highest standards of precision, reproducibility, and transparency to maximize adoption. By standardizing AI across the variant analysis pipeline, we can fast-track the path to genetic diagnoses, effectively bridging the interpretation gap and enabling rapid delivery of personalized medical interventions.
28 June 2026
Read appraisal →Scientific reports
Cross-domain transfer learning strategy enhances interpretability of deep learning model explanations
Clinical decision-making increasingly relies on deep neural networks (DNNs), yet their deployment in practice requires transparent and interpretable predictions. Explainable artificial intelligence (xAI) methods can identify input regions relevant to a model's decision, but their clinical interpretability remains limited. In this study, we investigated whether inductive transfer learning (TL) can reinforce domain-specific feature separation in xECGArch, a two-branch convolutional neural network for atrial fibrillation (AF) detection from electrocardiograms (ECGs). Each branch was pre-trained on a task aligned with its designated feature domain, P wave detection for morphology and RR interval variability prediction for rhythm, then fine-tuned on binary AF classification using an iterative layer freezing schedule. Deep Taylor decomposition (DTD) was applied to analyze explanations across all configurations. Fine-tuning accuracy ranged from 85.70% to 95.23%, remaining comparable to the original xECGArch architecture and previous TL-based approaches. However, DTD analysis demonstrated that morphology pre-training directed relevance toward P waves, whereas rhythm pre-training concentrated explanations on R peaks, with domain specificity increasing as more layers were frozen. These findings suggest that inductive TL can encourage domain-specific feature attribution in DNNs, improving the alignment of post-hoc explanations with clinically meaningful ECG regions.
26 June 2026
Read appraisal →Journal of medical systems
Beyond Validation: Operationalising Post-Deployment Surveillance of AI Medical Devices in Clinical Practice
Artificial intelligence medical devices are increasingly deployed in clinical practice, yet practical approaches to post-deployment monitoring remain poorly defined. We present a structured, decision-oriented approach to monitoring within healthcare institutions, grounded in our own deployment experience. By framing surveillance as a set of interdependent decisions, this model supports effective performance assessment and governance-linked corrective action, enabling safer and more accountable integration of AI into routine clinical care.
21 June 2026
Read appraisal →Journal of medical Internet research
The Open Syndrome Definition as a Machine-Readable Standard for Public Health: Design and Implementation Study
BACKGROUND: Case definitions are essential for effectively communicating public health threats. However, the absence of a standardized, machine-readable format poses significant challenges to interoperability, epidemiological research, data sharing, and the application of computational methods, including artificial intelligence. These barriers complicate collaboration across regions and organizations and hinder technological progress in public health. OBJECTIVE: This study aims to propose and release the first open, machine-readable format for representing case and syndrome definitions, together with tools and resources that enable their standardized and scalable use. METHODS: We developed the Open Syndrome Definition, a structured, machine-readable schema for representing case and syndrome definitions. We compiled official public health case definitions from multiple institutions and converted them into standardized, machine-readable representations using open-source tools. These tools, available through GitHub under the Massachusetts Institute of Technology license, automate the translation of narrative definitions into structured data. We also created a platform for browsing, analyzing, and contributing new definitions on our initiative website. RESULTS: The Open Syndrome Definition format enabled consistent, automated representation of case definitions across different diseases and jurisdictions. The conversion tools achieved high semantic fidelity, as assessed by qualitative expert review, between narrative and structured representations, supporting human verification and automated analysis. The dataset and accompanying tools demonstrated structural and semantic interoperability by standardizing definitions from various health systems into a unified format and integrating existing medical ontologies through JSON for Linked Data. To further illustrate practical applicability and downstream usage, we introduced a data filtering prototype that allows users to upload their own datasets and verify the results against the standardized definitions. CONCLUSIONS: The Open Syndrome Definition establishes a foundation for consistent and machine-readable public health definitions, facilitating reproducible research and interoperability at scale. By enabling systematic data exchange and artificial intelligence-driven analysis, it strengthens public health preparedness and supports more rapid, coordinated responses to emerging health threats.
19 June 2026
Read appraisal →