Research Appraisals
Evidence-based critical appraisals of the latest medical research, systematically evaluated using Oxford CEBM methodology.
Showing 11 appraisals
Scientific reports
Safety design guidelines for clinician-AI interaction in computer-aided diagnosis systems using system-theoretic framework with explainability validation
Ensuring the safety of Artificial Intelligence-enabled Computer-Aided Diagnosis systems is critical because diagnostic errors can have serious consequences for patient care. However, existing regulatory and risk management frameworks often do not sufficiently address the complex socio-technical interactions between clinicians and Artificial Intelligent systems, leaving key human-centered safety challenges underexplored. This paper presents a systematic and human-centered approach to deriving safety design guidelines for clinician-Artificial Intelligence interaction in Computer-Aided Diagnosis systems using System-Theoretic Process Analysis. Through this analysis, we identify critical hazards associated with clinician-Artificial Intelligence collaboration, including automation bias on system recommendations, and misinterpretation of explanations. Based on the identified unsafe control actions, we formulate a set of actionable and traceable safety design guidelines that promote transparency, coherent explanations, and calibrated trust in Artificial Intelligence-assisted decision-making. To bridge safety analysis and system design, the proposed guidelines are operationalized within a Computer-Aided Diagnosis interaction framework. The framework includes a safety-oriented Graphical User Interface that integrates multiple explanation methods and interactive mechanisms to promote clinician engagement. Furthermore, we introduce a safety-oriented evaluation approach that uses consistency across multiple explanation methods as a quantitative indicator of potentially unreliable or ambiguous explanations. By linking System-Theoretic Process Analysis, interaction design, and explainability evaluation, this work provides a unified and reusable framework for improving the safety and reliability of Artificial Intelligence-driven Computer-Aided Diagnosis systems.
31 July 2026
Read appraisal →JMIR research protocols
Evaluation and Comparison of Latent Health Risk Prediction Models for Clinical Triage: Protocol for a Mixed Methods Study
BACKGROUND: Clinical triage requires integrating multiple information sources to identify patients at risk of deterioration. Tools capturing global health assessments beyond disease-specific scores are being developed using either bottom-up aggregation of simple indicators or top-down machine learning from large datasets. Their alignment with expert clinical judgment remains poorly characterized. OBJECTIVE: This study evaluates 2 latent health measurement approaches: Frailty Index-laboratory, a transparent bottom-up tool aggregating laboratory abnormalities via deficit accumulation theory, and ETHOS-ARES (Enhanced Transformer for Health Outcome Simulation-Adaptive Risk Estimation System), a transformer-based foundation model generating multidimensional patient representations from electronic health records. We assess whether each tool's severity rankings align with clinical consensus and whether they offer utility in triage decisions. METHODS: In this 3-phase mixed methods study, at least 30 clinicians across hospital specialties reviewed 20 emergency department presentations derived from Medical Information Mart for Intensive Care IV-Emergency Department. Phase 1 compared unaided clinician severity and urgency judgments against model outputs using Spearman rank correlation, with a Turing-inspired indistinguishability test assessing whether model rankings fell within the distribution of clinician assessments. Phase 2 allocated clinicians to receive Frailty Index-laboratory or ETHOS-ARES outputs, measuring anchoring effects via within-person pre-post comparisons and exploring clinical utility through semistructured interviews analyzed using the Framework Method. RESULTS: Ethics approval was granted in June 2025 (KCL Research Ethics Office; MRSP-24/25-48707). Recruitment began in October 2025 (32 clinicians recruited as of manuscript submission), with data collection expected to be completed in January 2026 and analysis planned for March or April 2026. CONCLUSIONS: This study will quantify model-clinician agreement, measure anchoring effects, and generate qualitative insights on utility, trust, and adoption. The findings will inform the implementation of latent health measurement tools in clinical practice and provide a framework for the early-stage evaluation of artificial intelligence-based clinical decision support systems.
5 July 2026
Read appraisal →Medical image analysis
Medical hierarchical image classification via dual-geometry image-text learning
Hierarchical image classification is a fundamental challenge in medical image analysis, as tree-structured taxonomies inherently reflect biological and clinical relationships, spanning the general categorisation of disease entities and fine-grained cellular distinctions. Existing approaches primarily rely on multi-task learning and fine-grained detection, often requiring intricate model design and complex training strategies. In this paper, we aim to exploit the negative curvature property of hyperbolic space, which allows efficient representation of hierarchical structures. We propose a dual-geometry image-text framework, termed H2CL. Specifically, we introduce a lightweight classifier head on top of image backbones to extract both Euclidean and hyperbolic features, which are then combined to simultaneously preserve taxonomic consistency from an etiological perspective and enhance instance discrimination from a morphological perspective. Furthermore, a text branch is incorporated to integrate label semantics, where an entailment loss is employed to jointly model image-text alignment and inter-sample relationships. Extensive experiments on cervical cell, skin lesion, and gallbladder disease datasets demonstrate that our framework consistently outperforms advanced methods. Compared to the standard Swin Transformer, H2CL achieves an average accuracy improvement of 7% across all three datasets at the fine-grained level, with similarly consistent gains observed when integrated with other backbone models. The source code is publicly available at https://github.com/MCPathology/H2CL.
4 July 2026
Read appraisal →Medical physics
Gamma Knife treatment planning using knowledge-based reinforcement learning
BACKGROUND: Inverse planning is often used for Gamma Knife radiosurgery, allowing clinicians to mathematically specify desired clinical objectives and dose limits. The objectives are controlled by weights that are manually tuned to find the desired trade-off, which varies from case to case. Automation of this process can reduce clinical workload and improve consistency in plan quality. PURPOSE: To train a deep reinforcement learning agent using a reward function that incorporates the clinical metrics from past plans into its scoring criteria. The metric trade-off from the clinical plan is scored higher than all others, guiding the agent to produce plans with similar trade-offs. METHODS: An agent was trained to adjust the two priority weights (i.e., digital slider bars) in the clinical inverse planner. The agent consists of a neural network that receives the metrics and dose distribution of the current plan and the target and organ-at-risk masks as inputs. These methods were demonstrated on a dataset of 204 single-target metastases and a dataset of 71 acoustic neuroma cases. The cases were split into training, validation, and testing sets of size 123/41/40 and 42/14/15 for the metastases and acoustic neuromas, respectively. RESULTS: On the metastases test dataset, the agent achieved a significantly higher (p = 0.0136) average plan score (3.925 ± 0.130) compared to the default slider plans (3.874 ± 0.147). On the acoustic neuromas test dataset, the agent achieved a higher (p = 0.4493) average plan score (4.035 ± 0.177) compared to the default slider plans (3.995 ± 0.365). The higher plan scores are reflected in the four plan quality metrics: the agent's plans, on average, had metrics more similar to the clinical plans, compared to the default slider plans, for both test datasets. CONCLUSIONS: The proposed reward function enabled the agent to learn to find plans that aligned with historical planning decisions. Future work will investigate providing the agent with additional inputs that can explain the variability in planning decisions, which would further improve its performance.
4 July 2026
Read appraisal →BMJ health & care informatics
Detection of cancer recurrence from Thai-English electronic medical records using sentence embeddings.
OBJECTIVE: This study developed and validated monolingual and bilingual sentence-bidirectional encoder representations from transformers (SBERT) models for detecting cancer recurrence within Thai-English electronic medical records (EMRs) from Thai cancer hospitals. METHOD: A multicentre dataset of 32 436 documents from 1250 patients was used for model development. External validation involved an independent dataset of 9244 documents from 384 patients across two Thai cancer hospitals. Performance was benchmarked against a fine-tuned PubMedBERT (MetBERT). RESULTS: The development dataset included breast (43.9%), colorectal (12.1%), cervical (28.0%) and head and neck (16.0%) cancers. MetBERT achieved the highest area under the precision-recall curve (AUPRC) for locoregional versus no recurrence (11.1%) and locoregional versus distant recurrence (91.7%), while monolingual-SBERT excelled at distant versus no recurrence (32.0%). External validation demonstrated MetBERT superiority for locoregional versus no recurrence (9.30%-21.50%). For distant versus no recurrence, bilingual-SBERT performed best with AUPRC 17.55%-24.39%. While MetBERT led in distinguishing locoregional versus distant recurrence (88.30%-94.70%), bilingual-SBERT demonstrated robust external validation performance (AUPRC 85.25%-91.80%). DISCUSSION: Low AUPRC values (9%-32%) reflect the extreme class imbalance in real-world data (~1% recurrence prevalence). Despite this, fine-tuned MetBERT achieved highest performance, while bilingual-SBERT demonstrated superior robustness during external validation. This validates sentence embedding models for handling mixed Thai-English medical records in multilingual clinical environments. CONCLUSION: Sentence embedding frameworks provide a practical, generalisable solution for detecting cancer recurrence within multilingual EMRs. Despite text-length constraints, these models are suitable for clinical integration as a screening tool for cancer registry workflows.
4 July 2026
Read appraisal →Medical image analysis
STAGE challenge: Structural-Functional Transition in Glaucoma Assessment
Glaucoma is a leading cause of irreversible yet preventable blindness in working-age populations. Clinical diagnos is currently relies on functional visual field (VF) examinations to evaluate visual function and monitor disease progression, but these tests are time-consuming and require close cooperation with ophthalmologists, limiting their practicality. Although deep learning has shown promise for glaucoma diagnosis, most models have focused on structural changes in fundus and OCT images without linking them to functional VF outcomes. To bridge this gap, the HDMI Laboratory, in collaboration with the Zhongshan Ophthalmic Center of Sun Yat-sen University, organized the STAGE Challenge: Structural-Functional Transition in Glaucoma Assessment. The challenge explores predicting functional VF indicators-mean deviation, sensitivity maps, and pattern deviation probability maps-directly from structural OCT images. A dataset of 401 OCT volumes (each with 256 cross-sectional images) is released with corresponding VF labels and demographic data, along with a standardized evaluation framework to ensure fair comparison. This paper summarizes the methods of the seven finalist teams and analyzes their results. All teams employed dual-branch architectures integrating OCT with tabular data, and those using task-specific OCT models achieved the highest performance. These findings highlight the importance of tailored deep learning strategies for linking structural imaging to functional outcomes. The STAGE Challenge thus establishes the first standardized benchmark with a large, curated dataset for this task, enabling systematic evaluation of algorithms for structure-function analysis in glaucoma, and providing a foundation for future research. Details of the competition are available at http://hdmilab.cn/competition/stage.
3 July 2026
Read appraisal →Journal of dental research
AI in Oral Health Surveillance: Critical Review
Artificial intelligence (AI) holds transformative potential for advancing oral health surveillance by streamlining data collection, integration, and dissemination. This review critically synthesizes AI applications in oral health surveillance, highlighting its roles in 1) mapping population-level trends and oral health inequities using machine learning on epidemiological data; 2) enabling remote screening of oral diseases/conditions, including caries, oral hygiene, gingivitis, oral cancer, and malocclusion from intraoral images via computer vision models; and 3) integrating multimodal data through emerging large language models (LLMs) to enhance precision public health. We clarify the comparative strengths of distinct AI modeling for processing the primary data types in surveillance: structured clinical records, unstructured images, and integrated multimodal data. Traditional machine learning methods have been effectively applied to map population-level oral health disparities and identify risk factors but are constrained to structured data. Computer vision methods excel in individual-level diagnostics using intraoral photographs. To translate such capability into scalable surveillance, it is recommended to establish standardized imaging protocols for nonclinical settings, develop scalable models for fine-grained feature extraction, and implement reliable evaluation. These steps are essential to address pervasive challenges, including inconsistent image quality, domain shift, prevalence imbalance, and cost-effectiveness constraints. The future of AI-driven oral health surveillance lies in developing dental-adapted multimodal LLMs (MLLMs). Such MLLMs are uniquely capable of synthesizing disparate data streams, from structured clinical data to heterogeneous imaging modalities (e.g., intraoral photographs and radiographs), or even biomolecular data. This integration capacity facilitates a paradigm shift, moving current applications in dental consultation and clinical decision support toward a novel, tiered system for population-level monitoring. Such a system would provide actionable insights for public health policymaking via spatiotemporal analysis and causal inference. Next-generation AI-driven oral health surveillance systems can only succeed when built on a strong foundation of rigorous ethical principles and safeguards.
2 July 2026
Read appraisal →Journal of pediatric surgery
Artificial intelligence in rare pediatric solid tumor research and clinical care: A scoping review
BACKGROUND: Clinical and research advances for children with solid tumors are limited by their rare nature. Artificial intelligence (AI) and machine learning (ML) hold potential for advancing diagnosis, risk stratification, and treatment in rare diseases where individual patients contribute high-dimensional data. This review characterizes current AI/ML applications in rare pediatric solid tumors. METHODS: PubMed, Embase, and Web-of-Science Core Collection databases were queried for relevant articles published before February 10, 2025. Eligible studies applied AI/ML methods to study rare tumors in pediatric populations (≤19 years). Articles were evaluated for tumor diagnosis, AI/ML methods, clinical applications, comparators, effectiveness, and interpretability. RESULTS: Twenty-three studies (2009-2025) were included. Hepatoblastoma (12/23) and pediatric thyroid cancers (4/23) were most frequently studied. Supervised learning predominated (20/23), followed by unsupervised (9/23), deep learning (4/23), and hybrid/ensemble models (6/23). Applications included diagnosis (12/23), prognosis (10/23), and risk stratification (9/23). Twenty studies reported effectiveness measures, with many models achieving AUCs >0.85. In comparative analysis (17/23), AI/ML often equaled or exceeded expert consensus, traditional models, or alternative algorithms. Eight studies reported external validation. CONCLUSIONS: Current AI/ML research in rare pediatric extra-cranial solid tumors focuses on diagnosis, risk stratification, and prognosis, often outperforming traditional methods. Future work should prioritize external validation and clinical applicability.
1 July 2026
Read appraisal →Biomedical physics & engineering express
Transformer imputation in CTG: a length-dependent evaluation of reconstruction methods
Objective.Computer and artificial intelligence (AI) analyses are being increasingly used in intrapartum cardiotocography (CTG). However, fetal heart rate (FHR) signal loss, which frequently occurs in clinical practice, hinders visual interpretation and reduces accuracy. Although the impact is well recognized, there is no consensus on the maximum continuous gap length that can be reliably reconstructed under clinical conditions. Therefore, we aim to identify suitable imputation methods and clarify the clinical limits of valid missing-segment lengths.Methods.Using an open FHR dataset (CTU-UHB), we extracted continuous segments and artificially introduced Removed data of varying lengths. Using performance metrics such as difference and similarity, we compared the performance among a Transformer-based model and linear and spline interpolation. Additionally, we quantified the similarity between the Pre-impute and removed data to assess task difficulty.Results.We analyzed 2727 segments from 552 cases across multiple gap lengths. In terms of numerical accuracy root mean square error (RMSE), spline consistently performed significantly worse than others. The Transformer generally maintained a better mean accuracy than linear interpolation, although significant differences were observed only under specific conditions. Conversely, for waveform preservation (correlation), the Transformer consistently outperformed linear interpolation. Notably, in highly complex imputation tasks, the Transformer proved most robust, yielding the lowest RMSE and highest correlation. However, performance systematically degraded for all methods as gap lengths increased.Conclusion.The Transformer provides an effective baseline for FHR imputation under clinical conditions, achieving a favorable balance between waveform and numerical accuracy. By clarifying the clinical limits of valid missing-segment lengths-specifically the decline in reliability beyond 30 s-this study provides guidance for standardizing preprocessing in future CTG AI research and clinical implementation.Significance.For imputing intrapartum FHR data, the Transformer generally improves waveform reproducibility over linear interpolation for short-to-moderate gaps. Defining the reliability limit provides a crucial baseline for future CTG AI.
29 June 2026
Read appraisal →Physics in medicine and biology
Hybrid learning: a combination of self-supervised and supervised learning for joint MRI reconstruction and denoising in low-field MRI
Objective.Deep learning has demonstrated strong potential for magnetic resonance imaging (MRI) reconstruction. However, conventional supervised learning requires high-quality, high-signal-to-noise-ratio (SNR) reference data for network training, which are often difficult or impossible to obtain, particularly in low-field MRI. Self-supervised learning (SSL) eliminates the need for reference training data but may suffer from degraded performance under low-SNR conditions. To address these limitations, we propose hybrid learning, a new training framework that integrates self-supervised and supervised learning for joint MRI reconstruction and denoising when only low-SNR training data are available.Approach.Hybrid learning is implemented in two sequential stages. In the first stage, SSL is applied to fully sampled low-SNR data to generate higher-quality pseudo-references. In the second stage, these pseudo-references are then used as targets for supervised learning to reconstruct and denoise undersampled, noisy data. The proposed method was evaluated in four experiments using simulated and real noisy MRI data of the breast, lung, and brain across different field strengths (0.3 T to 3 T), sampling trajectories (Cartesian, spiral, and radial), noise levels, and undersampling ratios.Main Results.Hybrid learning consistently improved reconstruction quality relative to both supervised and self-supervised baselines under different acceleration rates, noise levels, and sampling patterns in all experiments. Compared with standard supervised learning using noisy references, it achieved up to 167.70% higher structural similarity index measure (SSIM), 95.41% lower normalized mean squared error (NMSE), and 90.70% lower high-frequency error norm (HFEN). Compared with standard SSL, it achieved up to 23.88% higher SSIM, 60.85% lower NMSE, and 49.13% lower HFEN.Significance.Hybrid learning enables improved MRI reconstruction under low-SNR imaging conditions by jointly addressing noise and undersampling. It provides a practical solution for robust deep learning-based reconstruction and is particularly well suited for applications such as low-field MRI, where image quality is limited by reduced SNR.
26 June 2026
Read appraisal →Physics in medicine and biology
Domain-specific adaptation for MR image synthesis with text-guided diffusion
Objective.Deep learning in medical imaging is severely constrained by data scarcity. Data synthesis offers a promising solution, but existing generative models have difficulty in restoring pathological texture features when trained on small-scale datasets. To address this, we propose a domain-specific, partition-based parallel text-guided latent diffusion model (LDM) for medical image synthesis.Approach.Each LDM operates on a defined image domain and is fine-tuned to reproduce specific texture characteristics. Diseased regions are identified from segmentation masks, while healthy regions are further subdivided using Voronoi-grayscale adaptation, enabling localized texture preservation. The fine-tuned LDMs independently synthesize corresponding image partitions, which are subsequently merged and denoised to form complete synthetic images with paired segmentation masks.Main results.We evaluated the approach on glioma MRI data, achieving a Fréchet Inception Distance of 13.65, demonstrating high perceptual realism. Texture fidelity was further supported by SSIM of 0.9674 and radiomic feature distribution analysis, both confirming close alignment between real and synthetic images. In a blinded visual Turing test, three radiologists achieved an average sensitivity of only 25.5% when identifying synthetic MRI slices, resulting in a 74.5% deception rate, and 41% of the synthetic samples were universally misclassified as real by all experts. In downstream experiments, U-Net trained on the synthetic-augmented dataset improved DSC by 14% on average.Significance.These results demonstrate that the proposed domain-specific adaptation framework can generate perceptually plausible, structure-preserving synthetic MRI slices in data-constrained environments, while improving downstream segmentation performance. The method therefore shows potential as an augmentation-oriented tool for AI model development, clinical teaching, assisted diagnosis, and rare-disease research.
25 June 2026
Read appraisal →