Research Appraisals
Evidence-based critical appraisals of the latest medical research, systematically evaluated using Oxford CEBM methodology.
Showing 5 appraisals
Seminars in fetal & neonatal medicine
Intraventricular hemorrhage in preterm infants: A systematic review of risk- and outcome-prediction models
Intraventricular hemorrhage (IVH) is a major complication of prematurity and one of the top causes of mortality and neurodevelopmental impairment. We conducted a systematic review of PubMed, Scopus, and Web of Science (From 1980 to 2025), identifying 40 studies evaluating prediction models (regression and machine learning) for risk of IVH occurrence and short- and long-term outcomes in preterm infants with IVH. Across these published studies, IVH risk prediction models included combinations of perinatal clinical variables, physiologic and hemodynamic indices, and serum biomarkers. Outcome-prediction models likewise varied. IVH grade was commonly included, with varying inclusion of comorbidities and neuroimaging-based injury markers. Most models performed acceptably, with machine-learning models showing better discrimination in larger datasets. However, the generalizability of the models is limited by heterogeneity in predictors, limited sample sizes, and inconsistent timing of predictor measurements. IVH severity remained the most consistent predictor across all outcome-prediction models. Despite promising performance assessments, clinical relevance is limited due to the infrequent reporting of model calibration, internal validation, and external validation. Future research that standardizes predictor definitions, leverages multicenter cohorts, and ensures thorough validation can advance early IVH risk and outcome-prediction models, thereby meaningfully improving clinical relevance and neonatal care.
2 Aug 2026
Read appraisal →Archives of toxicology
A review of machine learning in toxicology: current practices and reporting gaps
In recent years, machine learning and artificial intelligence approaches have been increasingly applied in the context of toxicological risk assessment. Many published overview, review, and comment papers discuss advantages, disadvantages, success stories, and open challenges for the application of machine learning models in toxicology. Machine learning methods using information from in vitro experiments can help to avoid animal experiments, thus allowing for larger numbers of experiments to be conducted. Drawbacks of machine learning models are the lack of mechanistic interpretability and the need for large amounts of high-quality data. In this work, we present a literature review of papers indexed in PubMed or published in the journal Computational Toxicology in the years 2022 to 2024, to assess the usage of machine learning methods in toxicology as well as the practices in reporting of methods and corresponding results. We do not address the suitability or the performance of methods, which is impossible to assess objectively without reanalysis on raw data, but focus on common practices and gaps in reporting. Major results are that many different machine learning methods are used in toxicology, often with appropriate internal validation. However, in only half of the cases, interpretation methods are used to address the problem that these models often make predictions as a black box. Moreover, there are very frequent gaps in reporting, in particular related to handling of missing values, and availability of data and code. Thus, this review can serve as a starting point for further tailored methodological research and guidance.
2 Aug 2026
Read appraisal →JMIR research protocols
Evaluation and Comparison of Latent Health Risk Prediction Models for Clinical Triage: Protocol for a Mixed Methods Study
BACKGROUND: Clinical triage requires integrating multiple information sources to identify patients at risk of deterioration. Tools capturing global health assessments beyond disease-specific scores are being developed using either bottom-up aggregation of simple indicators or top-down machine learning from large datasets. Their alignment with expert clinical judgment remains poorly characterized. OBJECTIVE: This study evaluates 2 latent health measurement approaches: Frailty Index-laboratory, a transparent bottom-up tool aggregating laboratory abnormalities via deficit accumulation theory, and ETHOS-ARES (Enhanced Transformer for Health Outcome Simulation-Adaptive Risk Estimation System), a transformer-based foundation model generating multidimensional patient representations from electronic health records. We assess whether each tool's severity rankings align with clinical consensus and whether they offer utility in triage decisions. METHODS: In this 3-phase mixed methods study, at least 30 clinicians across hospital specialties reviewed 20 emergency department presentations derived from Medical Information Mart for Intensive Care IV-Emergency Department. Phase 1 compared unaided clinician severity and urgency judgments against model outputs using Spearman rank correlation, with a Turing-inspired indistinguishability test assessing whether model rankings fell within the distribution of clinician assessments. Phase 2 allocated clinicians to receive Frailty Index-laboratory or ETHOS-ARES outputs, measuring anchoring effects via within-person pre-post comparisons and exploring clinical utility through semistructured interviews analyzed using the Framework Method. RESULTS: Ethics approval was granted in June 2025 (KCL Research Ethics Office; MRSP-24/25-48707). Recruitment began in October 2025 (32 clinicians recruited as of manuscript submission), with data collection expected to be completed in January 2026 and analysis planned for March or April 2026. CONCLUSIONS: This study will quantify model-clinician agreement, measure anchoring effects, and generate qualitative insights on utility, trust, and adoption. The findings will inform the implementation of latent health measurement tools in clinical practice and provide a framework for the early-stage evaluation of artificial intelligence-based clinical decision support systems.
5 July 2026
Read appraisal →Forensic science international
Is my smartwatch a valid witness? A systematic review and meta-analysis.
Wearable devices are being increasingly used not only in traditional research fields (e.g. clinical, health) but also in innovative ones such as digital forensics. Modern commercial smartwatches, equipped with sensors like accelerometers, gyroscopes, and photoplethysmographs, can record activity, heart rate, sleep, and other physiological parameters that may support forensic investigations, given that their data is valid. To perform a first step in studying such validity, this article presents a systematic review and meta-analysis summarizing the available studies validating data collected from smartwatches, focusing on their potential use in forensic investigations. The review examined studies evaluating smartwatches from four major brands, comparing their outputs, such as activity classification, step count, distance, heart rate, energy expenditure, and sleep metrics, against established gold-standard measures. We found that not all brands of smartwatches and not all outcomes were equally studied, with varying results in terms of accuracy, protocols, and reporting of results. Heart rate resulted the most studied (and most accurate) measure. Most studies validated smartwatches in healthy populations and performed validation in laboratory settings. Overall, while smartwatches show promise for enhancing digital forensic analyses, their current validation evidence is heterogeneous. The findings highlight the need for more extensive and standardized validation efforts, larger and more diverse study populations, and improved transparency from manufacturers. These conclusions are not only relevant for forensic applications but also extend to broader domains, such as clinical and health research, where the reliability of wearable-derived data is equally critical.
1 July 2026
Read appraisal →Physics in medicine and biology
Interlayer-aware postoperative facial appearance prediction in orthognathic surgery with bio-geometric guidance
Objective.Accurate prediction of postoperative facial appearance is essential for orthognathic surgical planning, yet remains challenging due to the nonlinear biomechanical coupling between bone and soft tissue. While deep learning methods offer a faster alternative to traditional biomechanical simulation, they typically map bony displacements directly to facial surface deformation, overlooking the intervening soft-tissue layers with distinct biomechanical properties through which displacement is progressively transmitted. Moreover, the many-to-one effect of regional bony movements on each soft-tissue point remains insufficiently captured.Approach.We propose a novel network that introduces an interlayer mixture-of-experts mechanism to decouple bone-to-surface deformation propagation into three biomechanically inspired proxy representations at different tissue depth levels. Since the contribution of each depth level varies across individuals and facial regions, a gated routing network adaptively weights each layer's contribution, providing a data-driven approximation of spatially heterogeneous, patient-specific deformation transmission. Additionally, a bio-geometric convolution module captures regional bony influences through elastic-weighted neighborhood aggregation.Main results.Evaluations on 88 orthognathic surgery patients demonstrate that the proposed method achieves a mean whole-face error of 1.39 mm, below the 2 mm clinical threshold, with statistically significant improvements over all baselines (p<0.05) and the highest clinician acceptance rate (94%).Significance.By incorporating bio-geometric priors into a deep learning framework, our approach enables a more physically grounded and interpretable prediction paradigm, supporting efficient and clinically reliable surgical planning.
21 June 2026
Read appraisal →