Artificial intelligence for early detection of diabetic retinopathy: A vision transformer-based approach
Clinical Snapshot
PICO Framework
| P — Population | Patients with fundus retinal images from public datasets (EyePACS and APTOS 2019) representing a spectrum of diabetic retinopathy severity levels (No DR to Proliferative DR) |
| I — Intervention | Compact Convolutional Transformer (CCT)-based Vision Transformer (ViT) model for automated classification of diabetic retinopathy severity |
| C — Comparator | State-of-the-art deep learning architectures (CNN-based models, other transformer variants) evaluated on the same datasets |
| O — Outcomes | Classification accuracy, F1-score across DR severity levels, sensitivity, specificity, and overall diagnostic performance for early DR detection |
Bottom Line
This paper presents a Compact Convolutional Transformer-based model for automated diabetic retinopathy classification, reporting 97% accuracy and F1-scores above 0.95 on two public benchmark datasets. While the transformer architecture represents a technically interesting approach to retinal image analysis, the study has critical methodological deficiencies that preclude any clinical recommendation. Most fundamentally, sensitivity and specificity — the essential metrics for a diagnostic accuracy study — are not reported, and no confidence intervals are provided for any outcome. The reference standard relies on pre-existing dataset labels without independent clinical verification or blinding. There is no external validation, no reporting of ungradable images, and no clinical utility analysis. The study was conducted entirely by computer scientists without apparent ophthalmology collaboration, and no prospective clinical evaluation has been performed. Benchmark dataset performance, even when impressive, is a poor surrogate for real-world diagnostic accuracy. Australian clinicians and health services should not adopt or recommend this tool based on the current evidence. Prospective clinical validation in representative screening populations, with full STARD-compliant reporting, independent ophthalmologist adjudication, and TGA regulatory assessment, would be required before any clinical translation could be considered.
Key Findings
P Value: Not reported
Effect Size: Overall classification accuracy of 97% on benchmark datasets
Primary Outcome: Multi-class classification of diabetic retinopathy severity (No DR, Mild, Moderate, Severe, Proliferative DR) using CCT-based Vision Transformer on EyePACS and APTOS 2019 datasets
Nnt Or Sensitivity: Sensitivity and specificity not reported. F1-scores reported as >0.95 across all DR severity classes. No 2×2 contingency table data provided. Positive and negative likelihood ratios cannot be calculated from reported data.
Confidence Interval: Not reported
Clinical Application
Technical feasibility of the CCT model is demonstrated on benchmark data. Real-world clinical feasibility is undemonstrated. Integration into existing retinal camera systems, electronic health records, or telehealth platforms is not described. Computational requirements, inference time, and deployment infrastructure are not reported. Regulatory approval pathways (e.g., FDA, CE marking, TGA) are not discussed. In Australia, diabetic retinopathy screening is recommended annually for all people with diabetes, as per RACGP and Diabetes Australia guidelines. The National Diabetes Services Scheme (NDSS) supports access to screening. Telehealth-based retinal screening programmes exist in remote and rural Australia (e.g., the KeepSight programme). AI-assisted DR screening tools require TGA approval as a Class IIa or IIb medical device software (SaMD) under the TGA's Software as a Medical Device framework. This study does not meet the evidentiary standards required for TGA regulatory submission or RACGP guideline endorsement. No Australian dataset validation has been performed. The model's performance on images from Australian community screening settings — which may differ in camera type, image quality, and patient demographics — is entirely unknown. Theoretically applicable to adults with diabetes undergoing fundus photography-based DR screening. However, the study population is defined only by benchmark dataset membership, not by clinical characteristics. Applicability to specific populations (type 1 vs. type 2 diabetes, duration of disease, comorbidities, image acquisition device type) cannot be determined from this paper.
Abstract
BACKGROUND: Early identification of diabetic retinopathy (DR), which is a primary cause of vision impairment globally, is a crucial phasis for effective intervention and treatment. Traditional screening workflows rely on manual diagnosis by ophthalmologists, which remains the gold standard but can be time-consuming and subject to variability due to human factors. To support and enhance the screening process, artificial intelligence (AI)-based tools have shown promise in automating DR detection, particularly with recent advances in deep learning. However, medical images with long-range dependencies and spatial linkages can be challenging for CNN-based algorithms to handle. METHODS: This paper proposes a Vision Transformer (ViT)-based model, specifically using a Compact Convolutional Transformer (CCT), for early automated detection of DR. The model uses self-attention techniques to improve feature extraction and classification performance; combining three main stages: the CCT tokenizer, transformer encoder, and sequence pooling. The proposed approach was trained on public datasets (EyePACS and APTOS 2019) and evaluated against state-of-the-art deep learning architectures. RESULTS: Our experimental findings demonstrate that ViT performs among the best in the current state of the art with an overall accuracy of 97% and F1-scores above 0.95 across all DR severity levels. Our system is primarily designed for the pre-screening stage of diabetic retinopathy workflows, enabling rapid and reliable identification of potential DR cases for further clinical evaluation. CONCLUSION: These results highlight the potential of transformer-based designs in medical picture analysis, as well as the implications for telemedicine and e-health solutions in real-time, especially in cases of low-resource settings.
References
- 1.ElAdel, A., Filali, I., & Zaied, M. (2026). Artificial intelligence for early detection of diabetic retinopathy: A vision transformer-based approach. PLOS ONE. https://doi.org/10.1371/journal.pone.0350854
Related Research
Indian journal of ophthalmology
Internet of things in assistive technology for people with visual impairment: A scoping review
1 Aug 2026
BMJ open ophthalmology
Screening for diabetic retinopathy with artificial intelligence in a primary care setting: a comparative cost analysis
1 Aug 2026
Science advances
Multispectral infrared-to-full-color upconversion expanding human vision
1 Aug 2026
This content is for educational purposes for healthcare professionals only and does not constitute clinical advice. Clinical decisions should be based on individual patient assessment, current guidelines, and appropriate specialist consultation. Editorial Standards · Privacy Policy · Terms of Service