Predictions from deep learning propose substantial protein-carbohydrate interplay
Clinical Snapshot
PICO Framework
| P — Population | Protein datasets from six organisms (E. coli, M. musculus, H. sapiens, S. cerevisiae, C. elegans, D. melanogaster) and manually curated carbohydrate-binding proteins |
| I — Intervention | Deep learning neural network models (PiCAP and CAPSIF2) for predicting protein-carbohydrate interactions |
| C — Comparator | Previous computational models and current estimates of <5% protein-carbohydrate binding prevalence |
| O — Outcomes | Prediction accuracy for protein-carbohydrate binding (90% balanced accuracy), residue-level prediction performance (Dice coefficient 0.57), and estimated prevalence of carbohydrate-binding proteins across proteomes |
Bottom Line
This computational study presents PiCAP and CAPSIF2, neural network models that predict protein-carbohydrate interactions with 90% accuracy. The research challenges current estimates suggesting <5% of proteins bind carbohydrates, instead predicting 35-40% prevalence across six species' proteomes. The models show particular strength in identifying extracellular and cell surface protein interactions (75% predicted to bind carbohydrates). While the computational approach offers significant advantages over experimental screening, the findings require experimental validation. The substantial increase in predicted carbohydrate-binding prevalence has important implications for understanding cellular processes, drug development, and therapeutic targeting. The tools could accelerate glycobiology research and inform pharmaceutical development, though clinicians should await experimental confirmation of specific predictions before clinical application.
Key Findings
P Value: Not applicable for this computational study design
Effect Size: 7-8 fold increase in estimated carbohydrate-binding protein prevalence (from <5% to 35-40%)
Primary Outcome: 90% balanced accuracy for protein-level carbohydrate binding predictions
Nnt Or Sensitivity: Dice coefficient of 0.57 for residue-level predictions, outperforming previous models
Confidence Interval: Not clearly reported
Clinical Application
High feasibility as computational tool requiring only protein sequence data Relevant to Australian biotechnology research, pharmaceutical development, and academic institutions. Could inform TGA drug evaluation processes for carbohydrate-targeting therapeutics and support research funded by NHMRC grants in structural biology Researchers in structural biology, drug discovery, and glycobiology; applicable to human protein analysis
Abstract
Noncovalent interaction between proteins and carbohydrates (sugars, glycans) is the basis for biological functions from metabolic regulation to intercellular recognition. It is a grand challenge to identify the protein-carbohydrate interactomes in organisms. Direct experiments would require extensive libraries of glycans to distinguish binding from nonbinding proteins. Computational screening of proteins for carbohydrate binding potential provides an attractive alternative. Current estimates propose that <5% of proteins bind carbohydrates, a number that is not well established. We therefore developed a neural network, "Protein interaction of Carbohydrates Predictor" (PiCAP), to predict whether a protein noncovalently binds to a carbohydrate. We trained PiCAP on a manually curated dataset of known carbohydrate binders and proteins that we identified as likely not to bind carbohydrates (transcription factors, cytoskeletal components, and small-molecule-binding proteins). PiCAP achieves 90% balanced accuracy on protein-level predictions of carbohydrate binding/nonbinding. Using the same datasets, we developed Carbohydrate Protein Site Identifier 2 (CAPSIF2) to predict protein residues that interact noncovalently with carbohydrates. CAPSIF2 achieves a Dice coefficient of 0.57 on residue-level predictions on our independent test dataset, outperforming previous models. To demonstrate the models' biological applicability, we investigated human cell surface proteins and further predicted the likelihood of carbohydrate binding in six proteomes (Escherichia coli, Mus musculus, Homo sapiens, Saccharomyces cerevisiae, Caenorhabditis elegans, Drosophila melanogaster). PiCAP predicts that ~35 to 40% of proteins in these proteomes bind carbohydrates, with 75% of extracellular and cell surface proteins predicted to bind. The PiCAP predicted binders are enriched for functions including growth factor receptor binding, inflammation, and cell-cell adhesion.
References
- 1.Canner, S. W., Schnaar, R. L., & Gray, J. J. (2026). Predictions from deep learning propose substantial protein-carbohydrate interplay. Proceedings of the National Academy of Sciences of the United States of America. https://doi.org/10.1073/pnas.2523342123
Related Research
Current opinion in chemical biology
Prediction of protein-protein interactions and co-complex models with deep learning
3 Aug 2026
Bioconjugate chemistry
Artificial Intelligence for Discovery in Life Sciences
16 July 2026
Journal of chemical theory and computation
Predicting and Decoding Allosteric Binding Sites Using Protein Language Models and Structure-Based Machine Learning: An Energy Landscape-Guided Explainable AI Framework
27 May 2026
This content is for educational purposes for healthcare professionals only and does not constitute clinical advice. Clinical decisions should be based on individual patient assessment, current guidelines, and appropriate specialist consultation. Editorial Standards · Privacy Policy · Terms of Service