Research Appraisalother

Predictions from deep learning propose substantial protein-carbohydrate interplay

Proceedings of the National Academy of Sciences of the United States of AmericaCanner, Samuel W, Schnaar, Ronald L, Gray, Jeffrey J26 May 2026DOI

Clinical Snapshot

85CEBM
Evidence: Moderateother

PICO Framework

P — PopulationProtein datasets from six organisms (E. coli, M. musculus, H. sapiens, S. cerevisiae, C. elegans, D. melanogaster) and manually curated carbohydrate-binding proteins
I — InterventionDeep learning neural network models (PiCAP and CAPSIF2) for predicting protein-carbohydrate interactions
C — ComparatorPrevious computational models and current estimates of <5% protein-carbohydrate binding prevalence
O — OutcomesPrediction accuracy for protein-carbohydrate binding (90% balanced accuracy), residue-level prediction performance (Dice coefficient 0.57), and estimated prevalence of carbohydrate-binding proteins across proteomes

Bottom Line

This computational study presents PiCAP and CAPSIF2, neural network models that predict protein-carbohydrate interactions with 90% accuracy. The research challenges current estimates suggesting <5% of proteins bind carbohydrates, instead predicting 35-40% prevalence across six species' proteomes. The models show particular strength in identifying extracellular and cell surface protein interactions (75% predicted to bind carbohydrates). While the computational approach offers significant advantages over experimental screening, the findings require experimental validation. The substantial increase in predicted carbohydrate-binding prevalence has important implications for understanding cellular processes, drug development, and therapeutic targeting. The tools could accelerate glycobiology research and inform pharmaceutical development, though clinicians should await experimental confirmation of specific predictions before clinical application.

Evidence: Moderate

Key Findings

  • P Value: Not applicable for this computational study design

  • Effect Size: 7-8 fold increase in estimated carbohydrate-binding protein prevalence (from <5% to 35-40%)

  • Primary Outcome: 90% balanced accuracy for protein-level carbohydrate binding predictions

  • Nnt Or Sensitivity: Dice coefficient of 0.57 for residue-level predictions, outperforming previous models

  • Confidence Interval: Not clearly reported

Clinical Application

High feasibility as computational tool requiring only protein sequence data Relevant to Australian biotechnology research, pharmaceutical development, and academic institutions. Could inform TGA drug evaluation processes for carbohydrate-targeting therapeutics and support research funded by NHMRC grants in structural biology Researchers in structural biology, drug discovery, and glycobiology; applicable to human protein analysis

Abstract

Noncovalent interaction between proteins and carbohydrates (sugars, glycans) is the basis for biological functions from metabolic regulation to intercellular recognition. It is a grand challenge to identify the protein-carbohydrate interactomes in organisms. Direct experiments would require extensive libraries of glycans to distinguish binding from nonbinding proteins. Computational screening of proteins for carbohydrate binding potential provides an attractive alternative. Current estimates propose that <5% of proteins bind carbohydrates, a number that is not well established. We therefore developed a neural network, "Protein interaction of Carbohydrates Predictor" (PiCAP), to predict whether a protein noncovalently binds to a carbohydrate. We trained PiCAP on a manually curated dataset of known carbohydrate binders and proteins that we identified as likely not to bind carbohydrates (transcription factors, cytoskeletal components, and small-molecule-binding proteins). PiCAP achieves 90% balanced accuracy on protein-level predictions of carbohydrate binding/nonbinding. Using the same datasets, we developed Carbohydrate Protein Site Identifier 2 (CAPSIF2) to predict protein residues that interact noncovalently with carbohydrates. CAPSIF2 achieves a Dice coefficient of 0.57 on residue-level predictions on our independent test dataset, outperforming previous models. To demonstrate the models' biological applicability, we investigated human cell surface proteins and further predicted the likelihood of carbohydrate binding in six proteomes (Escherichia coli, Mus musculus, Homo sapiens, Saccharomyces cerevisiae, Caenorhabditis elegans, Drosophila melanogaster). PiCAP predicts that ~35 to 40% of proteins in these proteomes bind carbohydrates, with 75% of extracellular and cell surface proteins predicted to bind. The PiCAP predicted binders are enriched for functions including growth factor receptor binding, inflammation, and cell-cell adhesion.

References

  1. 1.Canner, S. W., Schnaar, R. L., & Gray, J. J. (2026). Predictions from deep learning propose substantial protein-carbohydrate interplay. Proceedings of the National Academy of Sciences of the United States of America. https://doi.org/10.1073/pnas.2523342123
Share:XLinkedIn

This content is for educational purposes for healthcare professionals only and does not constitute clinical advice. Clinical decisions should be based on individual patient assessment, current guidelines, and appropriate specialist consultation. Editorial Standards · Privacy Policy · Terms of Service