Research Appraisaldiagnostic

Advancing Alzheimer Disease Prediction With Large Language Model-Based Linguistic Feature Analysis: Development and Validation Study

JMIR medical informaticsHsu, Ming-Hsia, Hwang, San-Yih, Tsai, Yi-Hang et al.28 May 2026DOI

Clinical Snapshot

60CEBM
Evidence: Moderatediagnostic

PICO Framework

P — PopulationPatients with suspected Alzheimer disease undergoing speech-based assessment using the ADReSSo 2021 dataset
I — InterventionLarge language model-based linguistic feature analysis of transcribed speech (readability, fluency, richness of detail, keyword relevance)
C — ComparatorExisting LLM-driven methodologies and conventional diagnostic approaches
O — OutcomesDiagnostic accuracy metrics (sensitivity, specificity, precision, F1-score) for Alzheimer disease prediction

Bottom Line

This validation study demonstrates promising diagnostic accuracy for a speech-based Alzheimer disease screening tool using large language model analysis of linguistic features. The framework achieved high sensitivity (91%) and specificity (96%) with good explainability. However, critical methodological details are missing, including confidence intervals, clear reference standards, and comprehensive bias assessment. While the technology shows potential for non-invasive cognitive screening, clinical validation in real-world settings is needed before implementation. The privacy-preserving local deployment option (F1-score 82%) may address data security concerns in healthcare settings. Clinicians should view this as early-stage research requiring further validation rather than ready-for-practice technology. The approach could eventually complement, not replace, standard cognitive assessment tools in appropriate clinical contexts.

Evidence: Moderate

Key Findings

  • P Value: Not reported

  • Effect Size: F1-score of 91.05% (mean across 3 independent runs)

  • Primary Outcome: Diagnostic accuracy for Alzheimer disease prediction using LLM-based linguistic analysis

  • Nnt Or Sensitivity: Sensitivity 91.08%, Specificity 96.29%, Precision 91.52%

  • Confidence Interval: Not reported

Clinical Application

Moderate - requires speech transcription capabilities and technical infrastructure, though local deployment option available Could complement existing cognitive assessment pathways under RACGP guidelines for dementia screening, potentially reducing referral burden to memory clinics. Would require TGA evaluation for clinical use and integration with existing PBS-funded cognitive assessment protocols. Patients with suspected cognitive decline requiring non-invasive screening assessment

Abstract

BACKGROUND: Alzheimer disease (AD) is a progressive neurodegenerative disorder with rapidly growing global prevalence. Early detection is critical for timely intervention; yet, conventional diagnostic methods remain costly and invasive. Speech-based assessment has emerged as a noninvasive alternative, as AD characteristically impairs linguistic abilities including fluency, coherence, and informational content. Recent advances in large language models (LLMs) offer new opportunities to extract structured linguistic features from transcribed speech for automated AD classification. However, existing LLM-based approaches often lack transparency and clinical interpretability, limiting their adoption in clinical workflows. OBJECTIVE: This study aims to investigate the influence of linguistic features extracted from transcribed speech, as analyzed by LLMs, on the accuracy and interpretability of AD prediction. METHODS: We propose a framework that leverages LLMs to analyze linguistic features extracted from transcribed speech for AD classification. Our approach focuses on 4 key aspects, including readability, fluency, richness of detail, and keyword relevance. To enhance classification accuracy, the framework integrates transcript embeddings with feature explanation embeddings, forming a comprehensive linguistic representation. We conducted extensive ablation studies to evaluate the contributions of individual features and benchmarked our framework against existing LLM-driven methodologies through pairwise explainability evaluations. Output stability was assessed across 3 independent pipeline runs. A fully local configuration (Llama 3 8B + nomic-embed-text) was tested to evaluate privacy-preserving deployment feasibility. Explainability was assessed via LLM-based pairwise comparison (Gemini-3.1-flash-lite) against the method of Bang et al across 54 correctly classified cases and by blinded evaluation from 2 neurologists. RESULTS: The proposed framework achieved a mean precision of 91.52%, a sensitivity of 91.08%, a specificity of 96.29%, and F1-score of 91.05% across 3 independent runs on the ADReSSo 2021 dataset, outperforming existing LLM-based approaches. A fully-local configuration (Llama 3 8B+nomic-embed-text, requiring no cloud application programming interface access) achieved an F1-score of 81.58%, demonstrating framework transferability to privacy-preserving deployment environments. Keyword relevance was the most influential feature (F1-score drop of 13.22 pp when removed). Explainability evaluations showed our method was preferred in 49 out of 54 cases via Gemini-3.1-flash-lite, with human experts preferring our method in 89 of 108 blinded assessments. CONCLUSIONS: These findings highlight that a structured linguistic feature analysis using LLMs provides a robust and interpretable framework for preliminary AD detection. Our approach offers a scalable and accessible solution that bridges artificial intelligence-driven text analysis with clinical applications, supporting early detection of cognitive decline through noninvasive assessment methods.

References

  1. 1.Hsu, M. H., Hwang, S. Y., Tsai, Y. H., Chang, Y. C., Liang, C. K., & Chang, C. Y. (2026). Advancing Alzheimer Disease Prediction With Large Language Model-Based Linguistic Feature Analysis: Development and Validation Study. JMIR Medical Informatics. https://doi.org/10.2196/86965
Share:XLinkedIn

This content is for educational purposes for healthcare professionals only and does not constitute clinical advice. Clinical decisions should be based on individual patient assessment, current guidelines, and appropriate specialist consultation. Editorial Standards · Privacy Policy · Terms of Service