Research Appraisals
Evidence-based critical appraisals of the latest medical research, systematically evaluated using Oxford CEBM methodology.
Showing 22 appraisals
European journal of radiology
A CT-based deep learning model to differentiate between benign and malignant adrenal lesions
OBJECTIVE: To develop a deep learning model to differentiate benign from malignant adrenal lesions. MATERIALS AND METHODS: A total of 380 patients with 385 pathologically confirmed adrenal lesions (101 malignant, 284 benign) were retrospectively included. Adrenal lesions were manually segmented on CT images and analyzed in a deep learning pipeline aimed at differentiating benign from malignant lesions. Four predictive models were developed that incorporated combinations of radiological data (tumor size, and spontaneous attenuation) and non-radiological data (i.e., medical history and laboratory results). Data of 267 patients were used as a training set and those of 113 patients for the test set. The diagnostic capabilities of the four models were estimated using sensitivity, specificity, accuracy, and areas under the receiver operating characteristic curves (AUC) using histopathological findings as the gold standard. The reproducibility of manual segmentation was estimated using the Dice similarity coefficient after blinded resegmentation of 40 adrenal lesions by an independent radiologist. RESULTS: Segmentation reproducibility achieved a mean Dice similarity coefficient of 0.92 ± 0.03 (range: 0.72-0.97). The most accurate model, which combined clinical, biological, and radiological data, achieved 84.2% accuracy (95% confidence interval: 79.9, 88.6) and an AUC of 0.93 (95% confidence interval: 89.9, 97.9) in the test set for diagnosis of malignant adrenal lesion. CONCLUSION: A deep learning model integrating preoperative clinical, biological, and radiological features demonstrates high capabilities in differentiating benign from malignant adrenal lesions on initial CT examination.
3 Aug 2026
Read appraisal →Telemedicine journal and e-health : the official journal of the American Telemedicine Association
Application of Telehealth and Artificial Intelligence in Military Health Care: A Promising Combination for the Future
BACKGROUND: Military and peace-support operations increasingly depend on medical capability delivered under difficult circumstances: distance, limited number of specialists, disrupted connectivity, and mass-casualty risk. Consequently, military health care systems operate during missions in uniquely constrained and high-risk environments, ranging from remote bases and naval platforms to active combat zones and humanitarian missions. These mission contexts demand rapid, reliable, and scalable medical solutions that can function despite limited personnel, infrastructure, and connectivity. METHOD: By reviewing existing literature and considering several different programs for military health care using telehealth and artificial intelligence (AI), the key performance indicators are explored to evaluate the synergy of telehealth and AI, following its implementation. RESULTS: Telehealth has become an increasingly important component of modern health care and holds the promise of increasing access to (special) health care without moving the patient but the information. Telehealth also enables continuity of care on missions and has already demonstrated its value in extending medical expertise across distance, while AI is evolving rapidly as a powerful enabler of decision support, automation, and predictive analytics in the medical field. As AI capabilities mature, several medical specialties may benefit, particularly in medical imaging, through faster triage, more consistent interpretation, and better prioritization of scarce specialist time. It also contributes to "decision support "while making proposals based on big data, machine learning and AI (e.g., in a Trauma registry). Telehealth provides the communication backbone that preserves these interactions and enables safe use of AI-enabled support in dispersed operations. CONCLUSIONS: The convergence of telehealth and AI in military health care will lead to a strategic capability with increased efficiency and enhance readiness, resilience, access to medical care, and quality of care for military personnel, including faster and more accurate diagnoses and better patient monitoring. Both technologies (AI and telehealth) are expected to continue to advance and play an even larger role in optimizing health care in the military.
3 Aug 2026
Read appraisal →Rheumatic diseases clinics of North America
Demystifying Artificial Intelligence: Key Concepts with Examples in Rheumatology
Artificial intelligence (AI) refers to a broad class of computational methods to perform tasks that typically require human intelligence, such as learning patterns, reasoning, and problem solving. AI is increasingly applied across rheumatology research and clinical practice to analyze complex data and support clinical decision-making, yet its underlying concepts are often perceived as opaque or inaccessible. This aim of this article is to demystify foundational AI concepts by defining core terminology and describing major learning paradigms and analytical methods, highlighted through clinically relevant examples from rheumatology. We explain how modern AI models differ from traditional analytical approaches.
2 Aug 2026
Read appraisal →Ultrasound in medicine & biology
A Systematic Review: The Application of Attention Mechanisms in Medical Ultrasound Image Processing
Among numerous medical imaging modalities, ultrasound imaging is one of the most commonly used diagnostic methods in clinical practice. However, ultrasound diagnosis heavily relies on physician experience, and diagnostic results often lack reproducibility. In recent years, with the rapid development of artificial intelligence technology, which provides new impetus for the automated medical ultrasound image processing and analysis. Among the numerous deep learning approaches proposed, attention mechanisms have become a key component for improving network robustness to cope with challenges in ultrasound imaging, such as low contrast, blurred boundaries, and variable object morphologies. This paper systematically reviews the attention mechanisms employed in medical ultrasound image analysis, which can be roughly divided these mechanisms into three categories based on the differences in feature focus dimensions: channel attention, spatial attention, and hybrid attention. Most importantly, we not only summarized the application scenarios and effectiveness of various attention mechanisms but also analyzed the potential challenges faced in the future.
2 Aug 2026
Read appraisal →Clinical & translational oncology : official publication of the Federation of Spanish Oncology Societies and of the National Cancer Institute of Mexico
Research progress of machine learning applications in gastric cancer diagnosis and therapy
Gastric cancer (GC), a malignant neoplasm originating from the gastric mucosal epithelium, represents one of the most prevalent cancers worldwide. Early detection is critical for improving treatment outcomes and patient prognosis. Recent advances in artificial intelligence (AI), particularly in machine learning, have introduced powerful computational and analytical capabilities that are increasingly being applied in GC research. Machine learning algorithms have shown considerable promise in enhancing the accuracy of GC diagnosis and optimizing therapeutic strategies. This review provides a concise overview of progress in machine learning applications within oncology, examines their current role and clinical utility in GC diagnosis and treatment, and highlights the transformative potential of machine learning in advancing GC management and patient care.
2 Aug 2026
Read appraisal →Cancer medicine
Applications of Artificial Intelligence in Cancer Diagnosis and Treatment
Driven by changes in lifestyle and environmental factors, the global incidence of cancer is steadily increasing, which has established it as a leading cause of mortality worldwide. The current paradigm for cancer diagnosis and treatment relies on conventional methods, such as imaging, endoscopy, and tissue biopsy, which present significant limitations regarding sensitivity in early screening, diagnostic specificity, and personalized treatment. Consequently, the development of more efficient and accurate technologies remains a major objective in modern oncology research, and artificial intelligence (AI) has emerged as a particularly promising solution. Through machine learning and deep learning algorithms, AI is reshaping cancer care by enabling automated detection of minute lesions during screening and quantitative analysis of pathological features for diagnosis. It may also advance tumor theranostics through multimodal data integration for treatment stratification, response prediction, and image-guided or targeted therapeutic decision-making, whereas providing data-driven recommendations for personalized treatment. Despite these prospects, medical AI development faces several key issues, including data bias, model explainability, clinical reliability and generalizability, emerging limitations of foundation models and generative AI, and regulatory and ethical issues that need to be addressed. By reviewing recent advances in AI across screening, diagnosis, theranostics, and treatment, we aim to clarify where these methods are already useful, where evidence remains limited, and why closer collaboration among clinicians, engineers, and data scientists is needed for clinical translation. We hope this review serves as a practical reference for researchers and clinicians evaluating how AI may be integrated into oncology in a more standardized, clinically responsible way.
1 Aug 2026
Read appraisal →Balkan medical journal
Bias and Fairness Across the Healthcare AI Lifecycle: A Clinician-Oriented Review
Artificial intelligence (AI) is increasingly being investigated and, in selected clinical settings, implemented to support diagnosis, triage, and workflow optimization. Although these systems have the potential to improve access, consistency, and efficiency, they may also reproduce or amplify health inequities when bias is introduced during development, evaluation, implementation, or postdeployment use. This clinician-oriented narrative review adopts a practical lifecycle approach to explain how algorithmic unfairness becomes clinically relevant, how clinicians can recognize it, and how institutions can mitigate its impact. We first outline the ethical, clinical, and mathematical dimensions of fairness. We then examine fairness risks and sources of bias across six stages of the healthcare AI lifecycle: problem formulation, data generation, model development, evaluation, implementation, and postdeployment monitoring and governance. Key mechanisms include biased proxy outcomes, unrepresentative or error-prone data and labels, model shortcut learning, hidden stratification, distribution shift, and human-AI interaction effects (e.g., automation bias and alert fatigue), all of which can create feedback loops and contribute to fairness drift over time. For each stage, we identify clinician-facing red flags and practical mitigation strategies, including defining clinically meaningful outcomes, using representative and well-documented datasets, conducting subgroup-stratified evaluations, performing external and prospective validation, justifying decision thresholds, implementing safeguards for human-AI interactions, and maintaining continuous postdeployment monitoring, including postmarket surveillance for regulated medical devices. Fairness cannot be ensured through a single metric, publication, regulatory clearance, or one-time validation. Instead, equitable healthcare AI requires transparent design, rigorous evaluation, local governance, and ongoing monitoring across diverse populations, clinical sites, devices, workflows, and time. Fairness should therefore be regarded as a continuous clinical and institutional responsibility rather than a downstream technical consideration.
1 Aug 2026
Read appraisal →Journal of medical ethics
AI interventions in cancer screening: balancing equity and cost-effectiveness
This paper examines the integration of artificial intelligence (AI) into cancer screening programmes, focusing on the associated equity challenges and resource allocation implications. While AI technologies promise significant benefits-such as improved diagnostic accuracy, shorter waiting times, reduced reliance on radiographers, and overall productivity gains and cost-effectiveness-current interventions disproportionately favour those already engaged in screening. This neglect of non-attenders, who face the worst cancer outcomes, exacerbates existing health disparities and undermines the core objectives of screening programmes.Using breast cancer screening as a case study, we argue that AI interventions must not only improve health outcomes and demonstrate cost-effectiveness but also address inequities by prioritising non-attenders. To this end, we advocate for the design and implementation of cost-saving AI interventions. Such interventions could enable reinvestment into strategies specifically aimed at increasing engagement among non-attenders, thereby reducing disparities in cancer outcomes. Decision modelling is presented as a practical method to identify and evaluate these cost-saving interventions. Furthermore, the paper calls for greater transparency in decision-making, urging policymakers to explicitly account for the equity implications and opportunity costs associated with AI investments. Only then will they be able to balance the promise of technological innovation with the ethical imperative to improve health outcomes for all, particularly underserved populations. Methods such as distributional cost-effectiveness analysis are recommended to quantify and address disparities, ensuring more equitable healthcare delivery.
25 July 2026
Read appraisal →Journal of medical Internet research
Model and Task-Aware Test-Time Scaling Strategies for Large Language and Vision-Language Models in Medicine: Evaluation Study
BACKGROUND: Test-time scaling has emerged as a promising method to enhance the reasoning capabilities of large language models (LLMs) and vision-language models (VLMs) during inference without additional training. While foundational studies established scaling paradigms in general domains, their applicability to the unique complexities of medical AI remains underexplored. OBJECTIVE: This study aims to conduct a comprehensive investigation of test-time scaling in the medical domain. We evaluate the impact of scaling across different model sizes and task complexities. Furthermore, we seek to identify domain-specific bottlenecks and assess model robustness against user-driven perturbations, such as misleading clinical authority. METHODS: This study evaluated a diverse set of general and medical-specific LLMs and VLMs. Experiments used five textual medical benchmarks comprising over 5500 questions and two multimodal benchmarks comprising 7000 samples. Performance was measured under three scaling conditions: increasing token budgets, iterative sequential scaling, and parallel scaling. Robustness was tested by embedding misleading hints with varying tones and levels of simulated clinical expertise into prompts. RESULTS: For nonreasoning LLMs, accuracy saturated quickly, with token usage often remaining under 500 tokens regardless of budget increases. Reasoning models demonstrated significant performance gains on complex tasks as token budgets increased. Notably, we identified distinct domain-specific behaviors. First, current VLMs showed a structural bottleneck in integrating visual clues and experienced limited benefit from token expansion. Second, medically fine-tuned LLMs excelled in clinical question answering but exhibited degraded scaling efficiency on calculation tasks compared to general-domain models. This reflects a disparity between qualitative clinical alignment and procedural logic. Third, while optimal scaling improved robustness, models exhibited a cognitive vulnerability by readily abandoning correct reasoning when confronted with misleading expert physician hints. Regarding scaling strategies, parallel scaling outperformed sequential scaling on easier tasks. Conversely, extended sequential scaling or increased budgets proved essential for complex problem-solving. CONCLUSIONS: Test-time scaling rules from general domains do not perfectly translate to medical AI. Longer reasoning is not universally beneficial. Concise reasoning with parallel scaling is optimal for simpler tasks. An extended chain of thought via sequential scaling or increased budgets is required for complex problems. Furthermore, safe clinical deployment requires addressing fundamental vision-language alignment, balancing clinical and procedural reasoning, and mitigating vulnerabilities to perceived clinical authority.
24 July 2026
Read appraisal →Journal of medical systems
Security Analysis of a Federated Learning Framework for Medical Image-to-Image Translation
Federated Learning (FL) emerged as a privacy-preserving paradigm for collaborative training of deep learning models across institutions without sharing patient data. This approach has been applied to complex tasks such as medical image-to-image (I2I) translation, including MRI-to-synthetic CT (sCT) generation. However, existing federated I2I frameworks often assume privacy preservation as an inherent property of FL rather than a requirement to be explicitly validated, leaving their robustness to representative adversarial threat scenarios largely unexplored. In this study, we evaluated the vulnerability of a federated MRI-to-sCT translation framework (FedSynthCT-Brain) to three representative attack classes: Deep Leakage from Gradients (DLG), Federated Membership Inference Attack (FedMIA), and data poisoning. The efficacy of corresponding defense mechanisms, such as Secure Aggregation (SecAgg) and Byzantine-robust median aggregation (FedMedian), were assessed. DLG enabled only the recovery of coarse anatomical structures, with no clinically identifiable details (SSIM ≤ 0.16, PSNR ≤ 11 dB) across clients, suggesting limited vulnerability under the evaluated DLG setting. In contrast, FedMIA achieved high membership discrimination, with AUC scores between 0.92 and 0.99, revealing a critical privacy vulnerability. The introduction of SecAgg reduced AUC values to near-random levels (0.23-0.56) across all centers without impacting synthesis quality. Under high-noise poisoning, the standard federated averaging (FedAvg) aggregation rendered the federation inoperative, while FedMedian restored performance close to the no-poisoning baseline in most scenarios, with significant residual degradation in specific center configurations. At low noise levels, the advantage of FedMedian was less consistent, as low-level noise injection may be indistinguishable from natural heterogeneity across centers, potentially enabling stealthy degradation. These findings demonstrate that federated I2I translation frameworks are not inherently secure and require explicit, multi-layered evaluation. As FL is increasingly adopted in clinical workflows, our results underscore the necessity of integrating cryptographic, algorithmic, and infrastructural safeguards for secure deployment.
6 July 2026
Read appraisal →IEEE transactions on bio-medical engineering
Deep Learning-Based Surrogate Model of Subject-Specific Finite-Element Analysis for Vertebrae
Subject-specific finite-element analysis (FEA) models enable accurate simulation of vertebral biomechanics but are often time-consuming to construct and solve under varying conditions. This study presents a novel deep learning (DL)/machine learning (ML)-based surrogate model that predicts stress distributions in vertebral bodies with high efficiency. The model integrates vertebral shape encoding and employs separate decoding branches for surface and internal nodes. It was trained on 3,960 synthetic L1 vertebrae generated via data augmentation from 42 real computed tomography (CT) scans. Evaluation on independent test samples yielded a mean absolute error (MAE) of 0.0596 MPa and an $\mathrm{R}^{2}$ of 0.864 for von Mises stress. Visualization results confirm strong agreement between predicted and FEA-computed stress patterns, with localized discrepancies observed at the anteroinferior margin and pedicles. Moreover, an end-to-end automated pipeline was established based on the developed model, reducing the total processing time from 90-120min to approximately 134-154s per subject. These findings highlight the potential of the proposed surrogate model to facilitate rapid, subject-specific biomechanical assessments in clinical workflows.
3 July 2026
Read appraisal →Magnetic resonance in medical sciences : MRMS : an official journal of Japan Society of Magnetic Resonance in Medicine
Integrating Artificial Intelligence into Prostate MR Imaging: Technical Foundations, Clinical Applications, and Workflow Implications
Prostate MRI has become a cornerstone of contemporary prostate cancer diagnosis, enabling improved detection of clinically significant disease while reducing unnecessary biopsies and overtreatment. However, prostate MRI remains technically demanding, time-consuming, and subject to inter-reader variability, particularly as healthcare systems move toward abbreviated protocols such as non-contrast MRI (biparametric MRI). In this context, artificial intelligence (AI) has emerged as a promising tool to enhance image quality, diagnostic consistency, and workflow efficiency across the prostate MRI pathway. This non-systematic narrative review provides a comprehensive overview of the technical foundations, clinical applications, and workflow implications of AI integration into prostate MRI. It summarizes key concepts in machine learning and deep learning relevant to prostate imaging and reviews current evidence supporting AI-based solutions for image quality assessment and reconstruction, automated prostate segmentation, lesion detection, and risk stratification. Particular attention is given to human-AI collaboration models, the role of AI in supporting equivocal lesions, and the integration of imaging with clinical variables for personalized risk estimation. In addition, it discusses the impact of AI on reporting efficiency, training, and standardization, as well as the current landscape of commercially available AI tools. Despite encouraging results from large multicenter studies, important challenges remain, including heterogeneity in study design, limited prospective validation, generalizability across institutions, and ethical and regulatory considerations. Overall, AI should be regarded as a complementary decision-support technology rather than a replacement for radiologists. Thoughtful implementation, robust validation, and appropriate user training are essential to ensure that AI meaningfully enhances the quality, efficiency, and reliability of prostate MRI-based care.
2 July 2026
Read appraisal →IEEE transactions on bio-medical engineering
EMI Cancellation for Shielding-Free Ultra-Low-Field MRI
OBJECTIVE: Ultra-Low-Field Magnetic Resonance Imaging (ULF MRI) offers low cost and portability but suffers from electromagnetic interference (EMI) in unshielded environments. This study developed a deep learning-based active EMI suppression method to overcome these limitations. METHODS: Using a 68mT ULF MRI system, human body coupling was identified as a primary EMI pathway. We proposed EMIC-Net, a U-Net architecture incorporating Transformer and hybrid attention mechanisms, to learn the data-driven nonlinear mapping from sensing coil signals to radio-frequency (RF) receiver coil interference. Acquired data underwent phase and gain compensation prior to model training. The model's efficacy was validated through in vivo human brain imaging, comparing its performance with EDITER and standard CNN methods, and by assessing the impact of varying EMI coil numbers and training data volumes. RESULTS: EMIC-Net effectively suppressed complex dynamic EMI. It restored image SNR from 2.35 dB to 17.63 dB, with significant PSNR and SSIM improvements. Image quality neared shielded acquisitions and surpassed comparative methods. Three EMI coils provided optimal balance, and the model showed data efficiency, requiring a small dataset for effective training. CONCLUSION: The EMIC-Net method accurately predicts and efficiently removes EMI in unshielded ULF MRI, offering superior performance and practicality. SIGNIFICANCE: This research promotes portable, low-cost ULF MRI for primary healthcare and bedside diagnosis. It also offers insights for mitigating complex EMI issues in other RF sensing domains.
2 July 2026
Read appraisal →IEEE transactions on pattern analysis and machine intelligence
A Survey on Interpretability in Visual Recognition
Visual recognition models have achieved unprecedented success in various tasks. While researchers aim to understand the underlying mechanisms of these models, the growing demand for deployment in safety-critical areas like autonomous driving and medical diagnostics has accelerated the development of eXplainable AI (XAI). Distinct from generic XAI, visual recognition XAI is positioned at the intersection of vision and language, which represent the two most fundamental human modalities and form the cornerstones of multimodal intelligence. This paper provides a systematic survey of XAI in visual recognition by establishing a multi-dimensional taxonomy from a human-centered perspective based on intent, object, presentation, and methodology. Beyond categorization, we summarize critical evaluation desiderata and metrics, conducting an extensive qualitative assessment across different categories and demonstrating quantitative benchmarks within specific dimensions. Furthermore, we explore the interpretability of Multimodal Large Language Models and practical applications, identifying emerging trends and opportunities. By synthesizing these diverse perspectives, this survey provides an insightful roadmap to inspire future research on the interpretability of visual recognition models.
1 July 2026
Read appraisal →Medical image analysis
Autodidactic dense anatomical models
Humans effortlessly interpret images by parsing them into part-whole hierarchies. Yet, deep learning models, despite excelling at capturing multi-level features, often fail to explicitly encode these part-whole hierarchies-an essential aspect of medical imaging, which boasts anatomical hierarchies in nature. To address this limitation, we introduce Adam-v2, a self-supervised learning framework that explicitly learns to encode inherent part-whole hierarchies within medical images through three key branches: (1) "localizability", which acquires discriminative representations to distinguish different anatomical structures; (2) "composability", which learns each anatomical structure in a parts-to-whole manner; and (3) "decomposability", which comprehends each anatomical structure in a whole-to-parts manner. Our extensive experiments showcase Adam-v2's advanced capability in anatomy understanding, unveiled through its embeddings (Eve-v2). Particularly, Eve-v2 demonstrates a zero-shot understanding of anatomy as revealed through its learned properties: ➀ preserving localizability of anatomical structures and ➁ encoding part-whole relations of anatomical structures as well as its emergent properties: ➂ understanding anatomical layouts via interpolation and extrapolation, ➃ associating each image pixel with (dense) semantic embeddings, ➄ recognizing anatomical symmetries, ➅ generating embeddings consistent across scales, and ➆ matching anatomical structures across images of the same patient with different diseases, images of different patients, and augmented views of the same image. Adam-v2 also offers robust and generalizable representations and stands out in ➇ few-shot learning, ➈ full-transfer learning, and ➉ novelty and anomaly detection. These capabilities and performance directly stem from our crafted anatomy learning strategy, which explicitly constructs hierarchies for distinct anatomical structures from unlabeled medical images. Project page: GitHub.com/JLiangLab/Eden.
1 July 2026
Read appraisal →The British journal of radiology
Artificial intelligence in nuclear medicine
Artificial intelligence (AI) holds great promise for advancing diagnostics and treatment in nuclear medicine. The rapid growth of AI over the past decade has been largely driven by advances in hardware components such as graphics processing units (GPUs) and the introduction of deep learning (DL) and convolutional neural networks (CNN). The integration of AI and medical imaging has the potential to revolutionize nuclear medicine by, for example, accelerating image acquisition, enhancing image quality, enabling advanced image generation, assisting image interpretation, and aiding treatment planning. Clinical applications have been demonstrated for most medical specialties, including oncology, neurology, and radionuclide therapy. The utilization of AI to provide automated, standardized procedures can help bring advanced imaging from major university centres to smaller local clinics, thus benefiting a broader range of patients. Additionally, AI has vast potential for predicting optimal treatment strategies, assessing risk, optimizing patient flow and outcomes, and even improving productivity, but these capabilities have yet to be fully utilized. The fraction of clinical AI applications in general healthcare reaching beyond the prototyping phase is reported to be as low as 2%. Indeed, in nuclear medicine, very few AI developments have reached commercial maturity. Currently, most AI applications in nuclear medicine follow the imaging flow from image acquisition and reconstruction, post-processing and image preparation, image analysis, and decision support for clinical interpretation. Below, we will briefly review selected areas and comment on challenges and opportunities for AI in nuclear medicine, with a special focus on the transition from development to clinical implementation.
1 July 2026
Read appraisal →Nephrology, dialysis, transplantation : official publication of the European Dialysis and Transplant Association - European Renal Association
Clinical applications of artificial intelligence in autosomal dominant polycystic kidney disease
Autosomal dominant polycystic kidney disease (ADPKD) is the most common genetic kidney disorder leading to kidney failure. Recent advancements in artificial intelligence (AI) are transforming the diagnosis, risk stratification, management and prognostication in ADPKD by enabling more accurate assessments and individualized care. AI-powered imaging tools enhance the measurement of total kidney volume (TKV), improving the precision and efficiency of monitoring disease progression and more reliable assessment of therapeutic response. Machine learning algorithms can integrate genetic, imaging and clinical data to predict kidney function decline, facilitating personalized treatment strategies. In addition, AI is being used to identify genetic variants and to refine genotype-phenotype relationships, offering deeper insights into disease variability. AI-enabled monitoring technologies can support longitudinal tracking of TKV measurements obtained through magnetic resonance imaging or computed tomography, thereby improving clinical decision-making and management. Furthermore, AI can optimize clinical trial design by improving patient selection, prediction of treatment responses and safety monitoring. Despite these promising developments, integrating AI into the clinical practice remains controversial and may pose several challenges. This review highlights the emerging clinical applications of AI in ADPKD, emphasizing its potential to advance precision medicine and improve patient outcomes.
1 July 2026
Read appraisal →PloS one
Few-shot skin lesion classification with Adaptive Multi-Scale Convolutional Attention Network
Computer-aided diagnosis of skin lesions faces core challenges, including scale diversity, blurred boundaries, intra-class morphological variations, and sparse data. Existing methods often rely on fixed receptive fields or generic attention mechanisms, struggling to fully adapt to the unique characteristics of skin lesions. To address this, we propose an Adaptive Multi-scale Convolutional Attention Network (AMCANet), which aims to achieve accurate and robust classification of skin lesions with limited data. AMCANet comprises three core modules: the adaptive multi-scale convolution module dynamically adjusts the receptive field to accommodate lesions of varying sizes; the hierarchical channel attention module integrates multi-level semantic information across different resolutions; and the skin spatial attention module leverages image gradient information to enhance lesion boundaries and local texture features. Extensive few-shot experiments on the HAM10000 and PAD-UFES-20 public datasets demonstrate that AMCANet significantly outperforms existing baseline models across multiple metrics, exhibiting promising generalization capabilities on the evaluated datasets. Qualitative and visual analyses further validate the model's ability to extract discriminative features and effectively focus on lesion regions. This study proposes a deep learning model, which demonstrates certain effectiveness in classifying skin lesions even with a small number of samples, providing a potential direction for future research.
27 June 2026
Read appraisal →Placenta
Deep learning-based early prediction of gestational diabetes mellitus through first-trimester placental texture analysis
INTRODUCTION: This study aimed to develop a multi-parameter fusion model for early GDM risk prediction and validate its performance through external multicenter testing. METHODS: A total of 628 pregnant women at 11+0-13+6 weeks were enrolled from two medical centers. The Center I cohort was divided into training (n = 356) and testing sets (n = 153). Radiomic features (1,289) and deep learning features (2,048) were extracted from placental ultrasound images. Feature-level fusion resulted in 3337 features, which were selected using Spearman correlation, mRMR, and LASSO. Five models were built: Rad Model, DTL Model, DLR Model, Clinic Model, and Combined Model. Performance was assessed using ROC analysis, DCA, and calibration curves. RESULTS: The Combined Model achieved the best overall performance, with an area under the ROC curve (AUC) of 0.879 in the internal validation, significantly outperforming any single-modality model (P < 0.05). DCA demonstrated that the fusion-based model provided higher net clinical benefit across a wide range of threshold probabilities compared with both "treat-all" and "treat-none" strategies. The calibration curve showed excellent agreement between predicted and observed probabilities (Hosmer-Lemeshow test, P > 0.05). DISCUSSION: The multimodal fusion model enhanced early GDM prediction by detecting subtle placental changes in first-trimester, enabling timely intervention and personalized decision-making.
27 June 2026
Read appraisal →Journal of medical systems
Beyond Validation: Operationalising Post-Deployment Surveillance of AI Medical Devices in Clinical Practice
Artificial intelligence medical devices are increasingly deployed in clinical practice, yet practical approaches to post-deployment monitoring remain poorly defined. We present a structured, decision-oriented approach to monitoring within healthcare institutions, grounded in our own deployment experience. By framing surveillance as a set of interdependent decisions, this model supports effective performance assessment and governance-linked corrective action, enabling safer and more accountable integration of AI into routine clinical care.
21 June 2026
Read appraisal →Urolithiasis
A validated custom pipeline for three-dimensional kidney stone renderings to create an open access repository
Three-dimensional (3D) rendering of urologic pathology plays an important role in simulation-based education, surgical training, and computer vision research; however, a standardized, open-access repository of high-fidelity kidney stone models stratified by chemical composition is lacking. We developed and validated a reproducible photogrammetry-based pipeline to generate realistic 3D kidney stone renderings. Chemically characterized human stones composed of calcium oxalate monohydrate (COM) (n = 11), uric acid (UA) (n = 5), cystine (n = 4), magnesium ammonium phosphate hexahydrate/carbonate apatite (MAPH/CA) (n = 2), and calcium hydrogen phosphate dihydrate (CHPD) (n = 3) were photographed using a custom-built rotating stage and dual fixed 4 K cameras. Rendered models were sent to 25 endourologists using a 5-point Likert-scale survey assessing geometric and surface texture fidelity. Successful 3D renderings were obtained for 8/11 COM stones, 5/5 UA stones, 2/2 MAPH/CA fragments, and 3/3 CHPD fragments, while all cystine stones failed to render. Across stone types, mean fidelity scores were highest for UA and COM stones (mean 3.8-3.9), intermediate for calcium phosphate stones (mean 3.6-3.8), and lowest for struvite stones (mean 3.0-3.3). Geometry scores were higher than texture scores overall, though this difference was not significant. Significant differences in geometric fidelity were observed across stone compositions (χ² = 9.30, p = 0.026). Inter-rater reliability was poor for individual evaluators (ICC = 0.10) but moderate for aggregated mean ratings (ICC = 0.67). This validated workflow enables the creation of generally realistic, open-access 3D kidney stone models (github.com/uro-glidar/3d-rendering-diverse-stones) for simulation, education, and future machine learning applications in endourology.
21 June 2026
Read appraisal →Cell systems
Fusing imaging and metabolic modeling via multimodal deep learning in ovarian cancer
Integrating genotype (e.g., transcriptomics), phenotype (e.g., imaging), and tumor microenvironment (e.g., metabolomics) is crucial to elucidating the molecular basis of ovarian cancer. However, there is a lack of robust multimodal integration methods when only a limited number of common samples is available. Here, we generate patient-specific metabolic models starting from transcriptomics data and integrate them with imaging data. We show that this multimodal integration-never attempted before-improves survival estimation and enables a mechanistic interpretation of the predictions. We assess the robustness of our approach with different combinations of transcriptomics, fluxomics, and 3D computerized tomography (CT) imaging data, correctly stratifying patients based on risk. Fusing metabolic modeling with imaging and transcriptomics significantly improves model accuracy compared with widely used transcriptomics-imaging approaches and elucidates critical metabolic reactions. Our approach is general and can be applied to other cancer types where coupled imaging-transcriptomics data are available. A record of this paper's transparent peer review process is included in the supplemental information.
19 June 2026
Read appraisal →