Initial-Visit Specialty Triage in Rare Diseases Using Large Language Models: Retrospective Benchmarking Study
Clinical Snapshot
PICO Framework
| P — Population | Patients with rare diseases presenting for initial-visit specialty triage, represented across five datasets including publication-derived case sets, RareBench-derived datasets, and a Facial phenotype-Gene-Disease Dataset-derived set |
| I — Intervention | Fourteen large language models (LLMs) evaluated for specialty triage accuracy, including Claude-opus-4-5, GPT-5.1, and twelve others, assessed over five independent runs per case |
| C — Comparator | Registered nurses and non-medical participants (human comparators on the publication-derived case set only); inter-model comparisons across proprietary vs open-weight, thinking vs non-thinking, and varying parameter scales |
| O — Outcomes | Primary: specialty triage accuracy (proportion of correct specialty assignments). Secondary: response time per case (seconds), consistency across five runs, subgroup performance by model type, reasoning mode, parameter scale, and phenotype count |
Bottom Line
This retrospective benchmarking study evaluates 14 large language models for specialty triage in rare diseases across five curated datasets, finding accuracy ranging from 44% to 71%, with the best-performing model (Claude-opus-4-5) outperforming registered nurses and non-medical participants on a single dataset. While the concept is clinically compelling — rare disease patients suffer prolonged diagnostic odysseys partly due to misdirected initial referrals — the study has critical methodological limitations that preclude clinical translation. Accuracy is reported without confidence intervals; standard diagnostic metrics (sensitivity, specificity, likelihood ratios) are absent; the reference standard derivation is opaque; and all cases are drawn from curated, confirmed datasets that do not reflect real-world triage complexity. The risk of data contamination — where LLMs may have been trained on the very case reports used for evaluation — is not addressed and could substantially inflate apparent accuracy. Human comparators were restricted to nurses and non-medical participants rather than experienced clinicians. Until prospective validation on unstructured real-world presentations is conducted, with appropriate regulatory oversight and clinician-supervised workflows, these findings should be interpreted as hypothesis-generating only. Senior clinicians should not adopt LLM-based triage tools on the basis of this evidence alone.
Key Findings
P Value: Not reported in abstract
Effect Size: Best-performing LLM (Claude-opus-4-5) achieved accuracy of 0.7141; LLMs as a group achieved mean accuracy of 0.5978 vs registered nurses 0.4914 and non-medical participants 0.4573 on the publication-derived case set
Primary Outcome: Specialty triage accuracy across five rare disease datasets, ranging from 0.4378 to 0.7141 across 14 LLMs
Nnt Or Sensitivity: Sensitivity and specificity not reported; accuracy (proportion correct) is the sole diagnostic performance metric. Consistency for top model (Claude-opus-4-5): 0.9653 across five runs. Response time range: 3.39 s/case (GPT-5.1) to 10.79 s/case (Claude-opus-4-5)
Confidence Interval: Not reported for any accuracy estimate
Clinical Application
Technical feasibility is demonstrated in a benchmarking context. Real-world feasibility requires integration into electronic health record or triage workflows, clinician oversight mechanisms, regulatory approval, and validation on prospective unstructured clinical data. None of these are addressed in the current study. In Australia, rare disease patients face well-documented diagnostic delays, often navigating multiple GP and specialist encounters before diagnosis. The RACGP and Rare Voices Australia have highlighted the need for improved triage pathways. However, no LLM-based triage tool is currently TGA-approved as a medical device software (SaMD) in Australia. Any clinical deployment would require TGA Software as a Medical Device (SaMD) regulatory assessment under the Therapeutic Goods Act 1989. PBS implications are indirect — earlier correct specialty referral could reduce unnecessary specialist consultations and investigations. The study's Chinese institutional context (West China Hospital, Sichuan University) and use of Chinese-context datasets limits direct generalisability to Australian rare disease epidemiology, healthcare system structure, and specialty taxonomy. Patients with suspected rare diseases at first point of clinical contact requiring specialty triage — particularly relevant in settings with limited access to specialist expertise or rare disease coordinators
Abstract
BACKGROUND: Specialty triage at first contact is an overlooked step in early diagnostic pathways for rare diseases. Patients often present with overlapping, multisystem, and atypical manifestations, making first-visit specialty selection challenging and potentially prolonging diagnostic pathways. OBJECTIVE: The aim of this study is to evaluate the accuracy, response time, and consistency of large language models (LLMs) for initial-visit specialty triage in rare diseases across multiple datasets, and to compare their performance with registered nurses and nonmedical participants. METHODS: In this retrospective benchmarking study, we used 5 rare disease datasets: a publication-derived case set, 3 RareBench-derived datasets, and a Facial phenotype-Gene-Disease Dataset-derived set. Fourteen LLMs were evaluated over 5 independent runs per case. Performance was assessed using accuracy, response time, and consistency, with subgroup analyses by model accessibility, reasoning mode, parameter scale, and phenotype count. Human comparison was conducted on the publication-derived case set using registered nurses and nonmedical participants. RESULTS: Across datasets, model accuracy ranged from 0.4378 to 0.7141. Claude-opus-4-5 achieved the highest accuracy (0.7141) and consistency (0.9653), averaging 10.79 seconds per case. GPT-5.1 had the shortest response time (3.39 s/case) and high accuracy (0.6948). Proprietary models had numerically higher average accuracy than open-weight models (0.6973 vs 0.6365). Nonthinking models achieved higher average accuracy than thinking models (0.6789 vs 0.5826) and had shorter response times, although this exploratory comparison was based on a small number of thinking models. Accuracy varied by phenotype count, with higher performance in cases with 1 to 2 or more than 14 phenotypes. On the publication-derived case set, LLMs achieved higher average accuracy than registered nurses and nonmedical participants (0.5978 vs 0.4914 and 0.4573). CONCLUSIONS: LLMs showed potential as assistive tools for initial-visit specialty triage in rare diseases. Model choice, reasoning mode, and phenotype information density influenced performance, but subgroup findings should be interpreted cautiously. Future work should evaluate LLM-based specialty triage in prospective clinical settings and develop clinician-supervised workflows with traceable evidence support.
References
- 1.Song, J., Xu, Z., Xiao, M., Bi, C., Zhang, Y., Zheng, X., Li, X., Cao, Q., Lu, Z., Yang, H., & Shen, B. (2026). Initial-visit specialty triage in rare diseases using large language models: Retrospective benchmarking study. Journal of Medical Internet Research. https://doi.org/10.2196/101711
Related Research
Orphanet journal of rare diseases
Clinical outcomes in alpha-mannosidosis: a systematic review of therapeutic approaches
23 July 2026
Nucleic acids research
FABIAN-variant 2026: improved prediction of the effects of DNA variants on transcription factor binding
13 July 2026
Nucleic acids research
SNPnexus: an enhanced web platform for large-scale and multi-sample variant analysis (2025 update)
13 July 2026
This content is for educational purposes for healthcare professionals only and does not constitute clinical advice. Clinical decisions should be based on individual patient assessment, current guidelines, and appropriate specialist consultation. Editorial Standards · Privacy Policy · Terms of Service