Mapping the evolution of deep learning and computer vision in robotic surgery: a bibliometric analysis of surgical video intelligence, instrument perception, and clinical translation.
Clinical Snapshot
PICO Framework
| P — Population | Published literature (2010–2026) on deep learning and computer vision applications in robotic surgery, retrieved from the Web of Science Core Collection (1,186 documents across 356 sources) |
| I — Intervention | Bibliometric and visualisation analysis using Bibliometrix/Biblioshiny, VOSviewer, and CiteSpace to map research trends, collaboration networks, keyword co-occurrence, and citation bursts |
| C — Comparator | No direct comparator; temporal comparison of early-phase (image guidance, registration, navigation) versus contemporary research themes (surgical video intelligence, instrument perception, autonomous assistance) |
| O — Outcomes | Research output growth, international collaboration patterns, leading contributing nations and institutions, most influential journals and authors, dominant methodological drivers (U-Net, residual learning, transformers, foundation models), and conceptual evolution of the field toward clinical translation |
Bottom Line
This bibliometric analysis maps 16 years of deep learning and computer vision research in robotic surgery, identifying rapid growth (16.09% annually), China–USA dominance, and a conceptual shift from image guidance toward surgical video intelligence and autonomous assistance. The study is methodologically sound as a scientometric exercise but carries important limitations: restriction to Web of Science excludes major engineering conference literature where much foundational work resides; no inferential statistics are reported; and a misattributed DOI raises editorial quality concerns. The findings are best interpreted as a research landscape orientation rather than clinical evidence. For senior clinicians and programme leaders, the key actionable insight is that the field remains pre-translational — instrument perception algorithms and workflow recognition systems are algorithmically mature but lack the prospective, multi-institutional clinical validation required for confident adoption. Australian robotic surgery programmes evaluating AI-embedded platforms should demand TGA-compliant SaMD evidence packages and prospective outcome data before clinical integration, rather than relying on publication volume or citation metrics as proxies for clinical readiness.
Key Findings
P Value: Not reported — descriptive bibliometric study only
Effect Size: Annual growth rate of 16.09% in publication output; peak of 216 publications in 2025; 1,186 total documents across 356 sources
Primary Outcome: Characterisation of the intellectual structure, growth trajectory, and thematic evolution of deep learning and computer vision research in robotic surgery from 2010 to 2026
Nnt Or Sensitivity: Not applicable — no clinical outcome data. Key scientometric indicators: 5,296 contributing authors; 30.69% international co-authorship rate; IEEE Robotics and Automation Letters identified as most productive and locally influential source; China leading in output volume, United States leading in citation impact
Confidence Interval: Not reported — no inferential statistics provided
Clinical Application
The findings are directly applicable to research strategy and curriculum development rather than bedside clinical practice. Institutions planning investment in surgical robotics platforms with embedded computer vision capabilities can use this landscape analysis to identify validated versus experimental technology domains. The identified gaps — prospective validation, multi-institutional datasets, integrated scene understanding — provide a practical checklist for evaluating vendor claims. Australian robotic surgery is expanding rapidly, with da Vinci systems operating across major metropolitan hospitals and newer platforms (Hugo RAS, Versius) entering the market under TGA regulatory oversight. The Therapeutic Goods Administration (TGA) requires clinical evidence of safety and efficacy for software-as-a-medical-device (SaMD) components, including computer vision modules — a regulatory gap this bibliometric analysis implicitly highlights by noting the absence of prospective clinical validation. The RACGP and relevant surgical colleges (RACS, USANZ, GESA) have not yet issued specific guidance on AI-assisted robotic surgery. Australian participation in multi-institutional dataset consortia (e.g., CholecT50, MISAW) would address the data diversity gaps identified. PBS implications are indirect but relevant: if deep learning-enabled robotic surgery demonstrates improved outcomes, MBS item reclassification and cost-effectiveness evaluation by MSAC would be required. Australian researchers should note that the study's WoS-only corpus may undercount contributions from Australian institutions publishing in conference venues. Robotic surgery researchers, surgical informaticians, health technology assessment bodies, and clinical programme leaders evaluating the maturity and translational readiness of deep learning-enabled robotic surgical systems
Abstract
Deep learning and computer vision are increasingly embedded in robotic surgery, yet the development and translational direction of this research domain remain incompletely characterized. We conducted a bibliometric and visualization analysis of publications retrieved from the Web of Science Core Collection using Bibliometrix/Biblioshiny, VOSviewer, and CiteSpace. A total of 1,186 documents published between 2010 and 2026 across 356 sources were included. Scientific output increased rapidly, with an annual growth rate of 16.09% and a peak of 216 publications in 2025. The field involved 5,296 authors, and international collaboration accounted for 30.69% of publications. IEEE Robotics and Automation Letters was the most productive and locally influential source. China and the United States were the leading contributors, with China showing the most rapid recent expansion and the United States retaining the highest citation impact. Citation-burst and keyword analyses identified U-Net, residual learning, transformer architectures, and foundation-model-enabled segmentation as major methodological drivers. The conceptual structure evolved from image guidance, registration, and navigation toward surgical video intelligence, instrument perception, workflow understanding, autonomous assistance, and clinical translation. Instrument perception emerged as a central link between algorithmic development and operative application. Future progress will require diverse multi-institutional datasets, external and prospective validation, integrated scene understanding, and rigorous evaluation of intelligent assistance within real robotic surgical workflows.
References
- 1.Yang, D., Shang, F., Xu, Y., Liu, J., Wang, J., Li, Y., Ma, X., & Lv, D. (2026). Mapping the evolution of deep learning and computer vision in robotic surgery: a bibliometric analysis of surgical video intelligence, instrument perception, and clinical translation. Journal of Robotic Surgery. https://doi.org/10.1109/TMI.2021.3069471 [Note: DOI as supplied by publisher metadata — independent verification recommended due to apparent DOI–article mismatch]
Related Research
Journal of robotic surgery
From robot assistance to surgical intelligence: global research trends and emerging frontiers of artificial intelligence-enhanced robotic surgery in urology.
27 July 2026
Journal of robotic surgery
Robot-assisted versus open kidney transplantation: an umbrella review of systematic reviews and meta-analyses
27 July 2026
Journal of robotic surgery
Reliability and readability of AI chatbot responses to patient questions about robot-assisted radical cystectomy
23 July 2026
This content is for educational purposes for healthcare professionals only and does not constitute clinical advice. Clinical decisions should be based on individual patient assessment, current guidelines, and appropriate specialist consultation. Editorial Standards · Privacy Policy · Terms of Service