Research Appraisals
Evidence-based critical appraisals of the latest medical research, systematically evaluated using Oxford CEBM methodology.
Showing 10 appraisals
Journal of chemical information and modeling
Functionally Guided Graph Learning for Robust Cross-Patient Cell-Type Annotation in Single-Cell RNA Sequencing
Cross-patient cell-type annotation in single-cell RNA sequencing (scRNA-seq) remains challenging due to pronounced interpatient heterogeneity and distribution shifts across patient-specific cellular contexts. Conventional annotation approaches often rely on proximity-driven graph construction or expression similarity, which may introduce spurious cell-cell connections and lead to unstable knowledge transfer across patients. To address this limitation, we propose PathoGraph, a functionally guided graph learning framework for robust cross-patient cell-type annotation. The proposed method integrates KEGG-7-based biosemantic graph structure learning with cross-patient representation adaptation. Specifically, pathway-derived functional semantic profiles are incorporated to refine patient-specific cell graphs, encouraging biologically coherent neighborhoods and suppressing noise introduced by purely expression-based similarity. Based on the refined graphs, a cross-patient representation adaptation mechanism further aligns embeddings between labeled reference patients and unlabeled query patients to facilitate reliable annotation transfer. Experiments on three cross-patient scRNA-seq data sets, including leukemia, breast invasive carcinoma, and colorectal cancer data sets, demonstrate that PathoGraph achieves stable annotation performance across 32 directed reference-to-query transfer tasks. Across all tasks, PathoGraph obtained an average ACC of 84.28% and an F1-score of 84.08%, showing competitive and stable performance compared with representative marker-based, correlation-based, and model-based annotation methods. Ablation studies further show that removing the biosemantic graph learning module reduces the average accuracy to 83.48%, highlighting the importance of functional-guided graph refinement. In addition, post hoc functional relevance analyses in immune-cell and cancer-associated contexts suggest that the learned cell-cell graphs capture biologically relevant neighborhood structures beyond expression-driven proximity. The source code and processed data are publicly available at: https://github.com/LiYuechao1998/PathoGraph.
29 July 2026
Read appraisal →Journal of chemical information and modeling
Gradient-Guided Graph Contrastive Learning for Mass Spectrometry-Based Proteomics Clustering
Single-cell proteomic data generated by mass spectrometry-based technologies provide direct insights into cellular functional states and have become increasingly important for revealing cellular heterogeneity in complex biological systems. Accurate clustering of such data is essential for identifying functionally distinct cell subpopulations and understanding biological processes such as immune responses, tumor heterogeneity, and cell fate regulation. However, mass spectrometry-based single-cell proteomic data are often characterized by high dimensionality, measurement noise, technical bias, and complex nonlinear structures, which pose major challenges to conventional clustering methods. To address these issues, this study proposes a gradient-information-guided graph contrastive learning framework for single-cell proteomic clustering. The proposed method adaptively reconstructs intercellular relationship graphs through gradient-guided structure learning and introduces a gradient-weighted contrastive loss to alleviate the influence of false-negative samples. By better preserving similarity among biologically related cells, the framework learns more robust and biologically meaningful representations. Experimental results on multiple data sets demonstrate that the proposed method outperforms conventional clustering approaches and existing graph contrastive learning methods in terms of clustering accuracy, stability, and biological consistency. Overall, this work provides an effective framework for clustering mass spectrometry-based single-cell proteomic data and offers new insights into the application of graph contrastive learning in bioinformatics.
28 July 2026
Read appraisal →mSystems
MeLSI: Metric Learning for Statistical Inference in microbiome community composition analysis
Microbiome beta diversity analysis relies on distance-based methods, including permutational multivariate analysis of variance (PERMANOVA) combined with fixed ecological distance metrics (Bray-Curtis, Euclidean, Jaccard, and UniFrac), which treat all microbial taxa uniformly, regardless of their biological relevance to community differences. This "one-size-fits-all" approach may miss subtle but biologically meaningful patterns in complex microbiome data. We present Metric Learning for Statistical Inference (MeLSI), a novel machine learning framework that learns data-adaptive distance metrics optimized for detecting community composition differences in multivariate microbiome analyses. MeLSI employs an ensemble of weak learners using bootstrap sampling, feature subsampling, and gradient-based optimization to learn optimal feature weights, combined with rigorous permutation testing for statistical inference. The learned metrics can be used with PERMANOVA for hypothesis testing and with principal coordinates analysis for ordination visualization. Comprehensive validation on synthetic benchmarks and real data sets shows that MeLSI maintains proper type I error control while delivering competitive or superior statistical power for detecting subtle community shifts and, crucially, supplies interpretable feature-weight profiles that clarify which taxa drive group separation. On the DietSwap data set, MeLSI was the only method to achieve significance at α = 0.05, demonstrating that adaptive weighting can detect diet-induced community shifts that fixed metrics miss. Across all data sets, the learned feature weights identified biologically relevant taxa while providing actionable insight that no fixed distance metric can supply. MeLSI therefore offers a statistically rigorous tool that augments beta diversity analysis with transparent, data-driven interpretability.IMPORTANCEUnderstanding which microbes differ between groups of interest could reveal therapeutic targets and diagnostic biomarkers. However, current analysis methods treat all microbes equally (similar to using the same ruler to measure everything, regardless of what matters most). This means subtle but biologically important differences may go undetected, especially when only a few key species drive disease states while hundreds of "bystander" species add noise. Metric Learning for Statistical Inference (MeLSI) solves this by learning which microbes matter most for each specific comparison. In comparing male and female gut microbiomes, MeLSI identified specific bacterial families driving the differences, providing actionable biological insights that standard methods miss. This capability is particularly crucial for detecting early disease biomarkers, where differences are subtle and masked by biological variability. By telling researchers not just whether groups differ, but which specific microbes drive those differences, MeLSI accelerates the path from microbiome data to testable biological hypotheses and clinical applications.
22 July 2026
Read appraisal →Cancer cell
OpenIO: An open framework for AI-native immunotherapy
We propose Open Immune Oncology (OpenIO), a framework integrating generative AI and omics to advance precision oncology. By leveraging biological scaling laws and foundation models, we aim to transition immunotherapy from empirical screening to rational, AI-native engineering of therapeutic interventions.
15 July 2026
Read appraisal →Nucleic acids research
ChIP-Atlas 2025 update: 10-year anniversary of a data-mining platform for exploring epigenomic landscape
ChIP-Atlas (https://chip-atlas.org/) is a data-mining platform that systematically processes and integrates public epigenomic profiling data, including ChIP-seq, ATAC-seq, DNase-seq, and Bisulfite-seq, together with curated sample metadata, currently encompassing >460 000 experiments across multiple model organisms. In this 2025 update, an experiment-level quality control framework was embedded in every experiment page, enabling intuitive assessment of data reliability and representativeness. In addition, a parametric analysis of gene set enrichment-based module was implemented, allowing enrichment analysis directly from continuous RNA-seq count tables without relying on threshold-defined, discrete gene lists, thereby helping elucidate underlying gene regulatory mechanisms. Combined with sustained data expansion, these updates advance ChIP-Atlas into a more quantitative and flexible platform for exploring the epigenomic landscape in multicellular organisms, marking its 10th anniversary.
13 July 2026
Read appraisal →Archives of microbiology
SigMine and OPathDb: a literature-mining pipeline and database of potential opportunistic pathogens
Conversion of unstructured biomedical literature into structured knowledge for identifying cross-domain associations between biological entities remains a challenging task. SigMine is an automated pipeline constructed to mine biomedical literature to identify significantly associated biological entities. SigMine performs biomedical entity recognition from PMC articles using the EuropePMC Annotation API. Advanced entity recognition was performed using Python scripting, NCBI E-Utilities, and an n-gram algorithm followed by extensive data cleaning and mapping against standard databases. Statistical evaluation identified significantly co-occurring entities. The entire workflow was automated through a modular framework developed in Python v3.13 with a Tkinter-based Graphical User Interface. SigMine enhances usability while retaining the flexibility to use new dictionaries for annotation. SigMine was used to construct a literature-derived potential human Opportunistic Pathogens Database (OPathDb), housing 5,626 potential opportunistic pathogens significantly co-occurring with 1440 diseases and 7121 genes mined from 25,000 PMC articles. Additional annotation of 598 significantly co-occurring metabolites and 30 affected tissues is available for 3204 and 227 pathogens, respectively. OpathDb has a user-friendly query interface searchable by organism, disease, tissue, gene, protein and metabolite available at https://www.opathdb.cbsblab-nsut.in . Organism-entity associations can be visualized as weighted networks, with color-coded nodes and significance-scaled edges. Significant associations of opportunistic pathogens like Akkermansia mucinifila with colorectal cancer and Segatella copri with glucose intolerance can be identified through OpathDb. Through this database, the SigMine framework demonstrates conversion of unstructured text in vast and heterogenous corpora into standardized and well-organized information. Statistically inferred associations in OPathDb are potential candidates for clinical and experimental validation.
13 July 2026
Read appraisal →Briefings in bioinformatics
EssTFNet: integration of adaptive time-frequency and DNA language models for interpretable human essential gene prediction
Essential genes are defined as indispensable for an organism's survival. The loss of function of these genes results in cell death or an inability to complete the normal life cycle. Research on essential genes is pivotal in elucidating the origin and evolution of life, as well as in identifying potential therapeutic targets. Therefore, predicting essential genes is of great scientific importance and has many applications in basic research and the biomedical field. In this study, we propose EssTFNet, a novel, interpretable deep learning framework that combines adaptive time-frequency analysis with a DNA language model to achieve accurate prediction of human essential genes while enabling mechanistic biological interpretation. EssTFNet leverages the architecture of ATFNet, which maps DNA and protein sequences into equivalent time-series signals to extract periodic and nonstationary features, enhancing the model's capacity to capture complex sequence patterns. Through feature selection and architectural optimization, EssTFNet achieves a favorable balance among prediction accuracy, model interpretability, and cross-tissue generalization. On the S1 benchmark task, EssTFNet outperformed mainstream sequence-based deep learning methods, achieving an area under the curve of 0.9679 and an area under the precision-recall curve of 0.8491. Additionally, the DeepLIFT attribution method was employed to identify functional motifs associated with gene essentiality, offering valuable insights for experimental validation. For the convenience of researchers, we have developed an easy-to-use web server and made it along with the source code in a GitHub repository: https://github.com/QIANJINYDX/EssTFNet. Overall, this study presents a potentially useful methodological framework for human essential genes prediction, which could provide valuable insights for future research and applications in this field.
5 July 2026
Read appraisal →BioEssays : news and reviews in molecular, cellular and developmental biology
AI in Genomics: From Variant Calling to Multi-Omics Integration
Artificial intelligence (AI) strategies are revolutionizing genomics by extracting complex patterns that traditional statistical pipelines are likely to miss. This mini-review aims to provide a concise overview of how AI is transforming major genomic technologies including variant calling, gene expression analysis, single-cell transcriptomics, CRISPR-Cas9 optimization, and multi-omics integration. In genome sequencing, machine learning variant callers greatly improve the accuracy and the rate at which single nucleotide and structural variants are called. In bulk RNA-Seq, AI augmented quantification, denoising, and differential expression modules complement the highly established STAR-featureCounts-DESeq2 pipeline, revealing subtle signals in big data sets. In single cell transcriptomics, deep learning approaches enhance batch correction, automate cell type annotation, and track developmental trajectories, hence clarifying cellular heterogeneity. AI-assisted guide RNA design, outcome prediction, and nuclease engineering enable more efficient CRISPR-Cas9 editing, reducing experimental cycles, and off-target effects. Finally, integrated platforms that combine genomic, transcriptomic, epigenomic, proteomic, and metabolomic layers provide an integrative view of cellular regulation and disease mechanisms. The review also covers current limitations, sparsity of data, model bias, privacy, and the need for standardized benchmarks and offers future directions in the form of interpretable models, collaborative learning, and open science practices. Together, these developments render AI an indispensable partner to unravel genomic complexity and accelerate precision medicine applications.
3 July 2026
Read appraisal →Open biology
Deep learning in tumour genomics: from multi-omics integration to precision oncology
Cancer remains a leading cause of death globally, with nearly 10 million deaths in 2020. Advances in genomic technologies have revolutionized cancer research, shifting focus towards precision medicine based on comprehensive tumour genomic profiling. Concurrently, deep learning (DL) has emerged as a powerful paradigm for complex biological data. This review critically assesses recent advances in DL applications for tumour genomics, emphasizing four key domains: DNA sequencing analysis for mutation detection, gene expression profiling for cancer subtype classification, methylation function prediction for epigenetic characterization and integrative multi-omics approaches for comprehensive tumour profiling. We systematically analyse how different DL architectures-including convolutional neural networks, recurrent neural networks, graph neural networks, autoencoders and transformers-address specific challenges in cancer genomics. Our review highlights how these approaches significantly enhance detection sensitivity for genomic alterations, improve cancer subtype stratification, identify novel biomarkers and optimize therapeutic target selection. We examine technical challenges in DL implementation, including model interpretability, data scarcity, computational requirements and integration issues, alongside emerging solutions such as explainable AI, federated learning, and multi-modal frameworks. By synthesizing methodological innovations and identifying research directions, this review provides bioinformaticians and cancer researchers with a roadmap for leveraging DL to advance precision oncology.
2 July 2026
Read appraisal →Nucleic acids research
SpliceSelectNet: a hierarchical Transformer-based deep learning model for splice site prediction
Accurate RNA splicing is essential for gene expression and protein function, yet the mechanisms governing splice site recognition remain incompletely understood. Aberrant splicing caused by mutations can lead to severe diseases, including cancer and genetic disorders, underscoring the need for accurate computational tools to predict splice sites and detect disruptions. Existing methods have made significant advances in splice site prediction but are often limited in handling long-range dependencies due to high computational costs, a factor critical to splicing regulation. Moreover, many models lack interpretability, hindering efforts to elucidate the underlying biological mechanisms. Here, we present SpliceSelectNet (SSNet), a hierarchical Transformer-based deep learning model that predicts splice sites from DNA sequences spanning up to 100 kb. By integrating local and global attention mechanisms, SSNet efficiently captures both proximal and distal regulatory signals while maintaining single-nucleotide resolution. Across multiple benchmark datasets, SSNet achieves state-of-the-art performance in splice site prediction and aberrant splicing detection. Systematic in silico mutagenesis demonstrates that attention scores reflect functional sequence importance, supporting their biological relevance. Long-range sequence perturbation experiments further show that SSNet captures distal regulatory effects beyond conventional receptive fields. Together, these results establish SSNet as a biologically interpretable framework for modeling long-range splicing regulation from genomic sequence.
23 June 2026
Read appraisal →