AI meets sleep surgery: assessing drug-induced sleep endoscopy interpretation with a large language model.
Clinical Snapshot
PICO Framework
| P — Population | Adults (n=16) undergoing drug-induced sleep endoscopy (DISE) for obstructive sleep apnea at a single tertiary academic centre in Israel |
| I — Intervention | Large language model (LLM) interpretation of anonymised DISE procedural videos, clinical vignettes, and examination findings, with generation of treatment recommendations |
| C — Comparator | Two sleep surgery experts, a senior otolaryngology resident, and the contemporaneous clinical report (reference standard) |
| O — Outcomes | Concordance with the clinical reference standard for video quality assessment, airway manoeuvre identification, anatomical site collapse grading (velum, oropharynx, tongue base, epiglottis), jaw thrust response, and treatment recommendations; assessed via Cohen's kappa and intraclass correlation coefficients; safety and reproducibility of LLM-generated recommendations |
Bottom Line
This small, single-centre proof-of-concept study (n=16) demonstrates that a large language model can interpret drug-induced sleep endoscopy videos with agreement statistics approximating those of expert sleep surgeons across most anatomical domains, and can generate treatment recommendations that are safe and consistent with accepted clinical practice. The blinded, prospective design and reproducibility testing are methodological strengths. However, the critically small sample size, absence of confidence intervals, undisclosed LLM identity and prompt methodology, and single tertiary-centre setting severely limit the conclusions that can be drawn. Agreement with a reference standard in 16 cases does not establish diagnostic accuracy, clinical utility, or generalisability. The epiglottis — the most surgically consequential and technically challenging domain — showed only moderate agreement for both human raters and the LLM, which warrants careful attention. For Australian sleep surgeons, this study is hypothesis-generating rather than practice-changing. Larger, multicentre, prospective validation studies with full methodological transparency, confidence intervals, and patient outcome data are required before LLM-assisted DISE interpretation could be considered for clinical integration. Regulatory pathways through the TGA's SaMD framework would also need to be navigated prior to any deployment.
Key Findings
P Value: Not reported
Effect Size: LLM achieved perfect or near-perfect agreement with the reference standard for velum, oropharynx, tongue base, and jaw thrust response (matching expert human rater performance); moderate agreement for epiglottic collapse (also consistent with human expert performance at this domain)
Primary Outcome: Concordance between LLM interpretation and the clinical reference standard across five DISE domains: velum, oropharynx, tongue base, epiglottis, and jaw thrust response
Nnt Or Sensitivity: Agreement statistics (Cohen's kappa and ICC) reported but specific numerical values not provided in the abstract; qualitative descriptors used (perfect, near-perfect, moderate). No sensitivity/specificity data reported. No unsafe LLM treatment recommendations identified across all 16 cases.
Confidence Interval: Not reported in available data — a critical methodological omission given n=16
Clinical Application
Feasibility in routine clinical practice remains unproven. Key barriers include: the specific LLM and prompt structure are not disclosed; integration into existing endoscopy reporting workflows is not described; data privacy and patient consent frameworks for uploading procedural videos to LLM platforms are not addressed; and the study does not report processing time, cost, or infrastructure requirements. The technology is promising but not yet operationally defined for clinical deployment. In Australia, DISE is performed at a limited number of tertiary ENT and sleep surgery centres, with access concentrated in major metropolitan hospitals. The RACGP and Australasian Sleep Association (ASA) do not currently include LLM-assisted DISE interpretation in clinical guidelines. The TGA would need to classify and regulate any LLM deployed as a medical device for diagnostic purposes under the Software as a Medical Device (SaMD) framework. PBS listing for DISE-guided surgical procedures (e.g., hypoglossal nerve stimulation, palatal surgery) requires robust pre-operative assessment, and any adjunct tool would need to demonstrate clinical non-inferiority or superiority in an Australian population before adoption. The findings are of academic interest but are not yet practice-changing for Australian sleep surgeons. Adults with suspected or confirmed obstructive sleep apnea being evaluated for upper airway surgery via drug-induced sleep endoscopy, particularly in settings where sleep surgery expertise is limited or where standardisation of DISE reporting is a priority
Abstract
STUDY OBJECTIVES: To evaluate the performance and safety of a large language model in interpreting drug-induced sleep endoscopy (DISE) videos and providing treatment recommendations for obstructive sleep apnea, compared to expert human raters and the contemporaneous clinical report. METHODS: This prospective, blinded study included 16 adults undergoing drug-induced sleep endoscopy at a tertiary academic center. For each case, an anonymized procedural video, clinical vignette, and examination findings were independently reviewed by two sleep surgery experts, a senior otolaryngology resident, and a large language model. All the raters assessed the video quality, airway maneuvers, airway collapse at each anatomical site, and recommended therapy. Concordance with the clinical reference standard was evaluated using Cohen's kappa and intraclass correlation coefficients. Safety and reproducibility were assessed through subgroup and error-type analyses. RESULTS: Human raters demonstrated perfect or near-perfect agreement with the reference standard for the velum, oropharynx, tongue base, and jaw thrust response, and moderate agreement for epiglottic collapse. The large language model matched expert performance for all domains except the epiglottis, where moderate agreement was observed. Model-generated treatment recommendations were safe, consistent with accepted clinical practice, and highly reproducible between independent runs. No unsafe or discordant recommendations were identified. CONCLUSIONS: In this single-center study of 16 patients, a large language model accurately interpreted DISE videos and generated safe recommendations consistent with accepted clinical practice, approximating expert performance in most domains. Larger multicenter cohorts are needed to validate these findings and confirm generalizability. Statement of Significance This study demonstrates that artificial intelligence can interpret complex airway videos in sleep surgery and provide safe, expert-level treatment advice. By comparing the performance of a large language model with that of experienced clinicians, our findings suggest that advanced technology can help standardize decision-making in a highly subjective area of care. This work highlights the promise of artificial intelligence as an adjunct to clinical judgment, especially in settings where expert access is limited. However, important questions remain about the use of artificial intelligence in challenging cases and its integration into real-world practice. Future research should focus on larger, more diverse groups of patients and on ensuring ongoing oversight and safety.
References
- 1.Hack, S., Alsleibi, S., Shemesh, S., & Nakache, G. (2026). AI meets sleep surgery: assessing drug-induced sleep endoscopy interpretation with a large language model. Sleep. https://doi.org/10.1093/sleep/zsaf338
This content is for educational purposes for healthcare professionals only and does not constitute clinical advice. Clinical decisions should be based on individual patient assessment, current guidelines, and appropriate specialist consultation. Editorial Standards · Privacy Policy · Terms of Service