Evidence-Based Medicine

Research Appraisals

Evidence-based critical appraisals of the latest medical research, systematically evaluated using Oxford CEBM methodology.

Showing 4 appraisals

qualitativeEvidence: Moderate
80CEBM

Journal of medical Internet research

Acceptability of Technologies to Support Early Dementia Detection: Qualitative Study With the Boston University Alzheimer's Disease Center Cohort

BACKGROUND: Dementia is on the rise globally due to increasing life expectancies and population growth. Digital technologies may help detect early signs, enabling timely interventions to slow or reverse cognitive decline. However, to support the successful implementation of these digital technologies into health care settings, they must be acceptable to target users. Older adults and those with mild cognitive impairment (MCI) are at risk of developing dementia in later life and need to be able to use these technologies in order for this intervention to be approved and implemented in clinical practice. OBJECTIVE: This study explored the perspectives of older adults and those living with a clinical diagnosis of MCI on the acceptability of using various digital technologies that have the potential to support early dementia detection. METHODS: Participants were recruited from Boston University's Alzheimer's Disease Research Center. Participants selected at least 2 technologies from 9 different wearables and software to use for 2 weeks, at 3-month intervals, over a total duration of 2 years. A subgroup of self-selecting participants was interviewed after the first 2 weeks of use to gather initial perspectives regarding the acceptability of using the digital technologies. An inductive framework thematic analysis approach was used, assisted by NVivo (version 14.23.2; QSR International). RESULTS: In total, 13 individuals living with a clinical diagnosis of MCI and 11 adults aged 65 years and older were interviewed. Our analysis identified five key themes: (1) gamification, (2) wearability, (3) user guidance, (4) burden of use, and (5) usefulness. Gamified apps were generally liked, although users with little experience of digital games needed time to adjust. Wearables resembling everyday accessories (eg, watches) were preferred, but complaints about tight or uncomfortable straps were frequently reported. Clear instructions were critical to support correct use, but many participants would have liked more troubleshooting support when technical issues arose. The use of 5 or more devices led to a high burden, especially when devices had practicality issues such as not being waterproof. Devices offering personal feedback were perceived as useful to satisfy personal interests, though some questioned their usefulness within health care. Participants raised concerns about losing valued personal interactions with health care professionals and questioned how their existing health conditions and treatment for such conditions may affect the validity of the data collected by the devices. CONCLUSIONS: These findings can guide researchers in choosing appropriate devices and minimizing burden. Future work should explore the views of those experiencing digital exclusion to ensure equitable access to dementia-detection technologies.

31 May 2026

Read appraisal →
diagnosticEvidence: Moderate
60CEBM

JMIR medical informatics

Advancing Alzheimer Disease Prediction With Large Language Model-Based Linguistic Feature Analysis: Development and Validation Study

BACKGROUND: Alzheimer disease (AD) is a progressive neurodegenerative disorder with rapidly growing global prevalence. Early detection is critical for timely intervention; yet, conventional diagnostic methods remain costly and invasive. Speech-based assessment has emerged as a noninvasive alternative, as AD characteristically impairs linguistic abilities including fluency, coherence, and informational content. Recent advances in large language models (LLMs) offer new opportunities to extract structured linguistic features from transcribed speech for automated AD classification. However, existing LLM-based approaches often lack transparency and clinical interpretability, limiting their adoption in clinical workflows. OBJECTIVE: This study aims to investigate the influence of linguistic features extracted from transcribed speech, as analyzed by LLMs, on the accuracy and interpretability of AD prediction. METHODS: We propose a framework that leverages LLMs to analyze linguistic features extracted from transcribed speech for AD classification. Our approach focuses on 4 key aspects, including readability, fluency, richness of detail, and keyword relevance. To enhance classification accuracy, the framework integrates transcript embeddings with feature explanation embeddings, forming a comprehensive linguistic representation. We conducted extensive ablation studies to evaluate the contributions of individual features and benchmarked our framework against existing LLM-driven methodologies through pairwise explainability evaluations. Output stability was assessed across 3 independent pipeline runs. A fully local configuration (Llama 3 8B + nomic-embed-text) was tested to evaluate privacy-preserving deployment feasibility. Explainability was assessed via LLM-based pairwise comparison (Gemini-3.1-flash-lite) against the method of Bang et al across 54 correctly classified cases and by blinded evaluation from 2 neurologists. RESULTS: The proposed framework achieved a mean precision of 91.52%, a sensitivity of 91.08%, a specificity of 96.29%, and F1-score of 91.05% across 3 independent runs on the ADReSSo 2021 dataset, outperforming existing LLM-based approaches. A fully-local configuration (Llama 3 8B+nomic-embed-text, requiring no cloud application programming interface access) achieved an F1-score of 81.58%, demonstrating framework transferability to privacy-preserving deployment environments. Keyword relevance was the most influential feature (F1-score drop of 13.22 pp when removed). Explainability evaluations showed our method was preferred in 49 out of 54 cases via Gemini-3.1-flash-lite, with human experts preferring our method in 89 of 108 blinded assessments. CONCLUSIONS: These findings highlight that a structured linguistic feature analysis using LLMs provides a robust and interpretable framework for preliminary AD detection. Our approach offers a scalable and accessible solution that bridges artificial intelligence-driven text analysis with clinical applications, supporting early detection of cognitive decline through noninvasive assessment methods.

31 May 2026

Read appraisal →
observationalEvidence: Moderate
70CEBM

JMIR formative research

At-Home Sleep Electroencephalography Assessment in Young and Older Adults Using a Novel Wireless Soft Electronics Sleep Monitoring System: Experimental Study

BACKGROUND: Sleep quality declines with age and is a known contributor to multiple chronic health conditions, including Alzheimer disease. Emerging evidence suggests that certain electroencephalography (EEG) neural signatures measured during sleep may be predictive of cognitive decline in older adults. Sleep EEG signals are traditionally measured using bulky, rigid, and uncomfortable equipment in an unfamiliar laboratory setting, which can negatively impact sleep signals. Due to these limitations, sleep EEG data acquisition is typically limited to a single night. OBJECTIVE: This study aimed to validate our recently developed portable, skin-like EEG monitoring patch for 7 nights in the home environment in a pilot sample of young and older adults by evaluating usability and acceptance, and replicating age-related differences in sleep architecture observed in the polysomnography literature. METHODS: Eighteen young adults and 18 cognitively unimpaired older adults without sleep disorders were enrolled (data from 11 young adults and 12 older adults were included in the analyses) in a 7-night study during which they wore novel, gel-free, wireless, ultrathin, skin-conforming, sleep monitoring, fabric-based patches. These patches were self-applied to the forehead and face for optimal usability and comfort. The patches incorporate laser-cut mesh electrodes with low-profile electronics (including a rechargeable battery and amplifier) and transmit EEG signals to a participant-controlled, Bluetooth-enabled, tablet-based data acquisition app. An automated algorithm was used to stage sleep and assess microarchitecture features from the EEG commonly impacted for each participant. Averages across nights were computed for these sleep features for each participant. RESULTS: Young and older adults reported that the sleep patch was easy to use and comfortable to wear. There was no loss of signal power over 7 nights of wear across participants (retained-data signal-to-noise ratio over the 7-d period: young adult, mean 20.69, SD 12.78, maximum 52.13, minimum 5.19; older adult, mean 22.10, SD 9.39, maximum 49.96, minimum 13.79). Most datasets not retained were lost due to poor reference electrode adhesion on the nose (75/101, 74% of lost datasets in young adults and 57/88, 65% in older adults). Trained sleep technologists verified that the retained datasets were of sufficient quality to be scored without difficulty. Expected age-group differences in sleep features were observed, including age-related reductions in stage N3 sleep (young adult, mean 18.55, SD 6.70; older adult, mean 10.40, SD 6.43; Mann-Whitney U=42.0; P=.01) and reduced sleep spindle density (young adult, mean 2.92, SD 2.24; older adult, mean 0.94, SD 1.33; Mann-Whitney U=45.0; P=.006). CONCLUSIONS: This study demonstrates that our novel, comfortable, wearable patch can reliably measure physiological sleep data over multiple nights at home in adults across the lifespan, thereby making multinight sleep assessment in cognitive aging studies and clinical research more accessible than traditional polysomnography. In future studies, the small, lightweight system, which is highly scalable, can be shipped inexpensively to participants' homes, making this technology and research accessible to individuals who may have difficulty traveling or who are hesitant to travel to a laboratory or clinic.

11 May 2026

Read appraisal →
otherEvidence: Moderate
80CEBM

Behavior research methods

Neural cognitive diagnosis modeling incorporating response times

Cognitive diagnosis is a fundamental issue in the field of intelligent education, aiming to identify students' mastery of specific knowledge concepts. With computerized testing, response time (RT) is a process data that can be collected. Incorporating RT in cognitive diagnosis assessment can enhance diagnostic accuracy. However, RT is only considered in traditional statistical cognitive diagnosis models. Compared with traditional statistical diagnostic models, cognitive diagnosis models based on neural networks have advantages such as high precision and strong generalization ability. Therefore, this paper proposes a JRT-NCD (joint response times neural cognitive diagnosis) model that uses neural networks to model the complex nonlinear interactions between exercises and students and incorporates RT as a new feature to refine diagnostic results on student abilities. Research findings on three datasets of PISA2012, 2MFC, and ASSIST09 indicate that: (1) In comparison with traditional statistical models, neural networks have better fitting capabilities for real complex nonlinear data; (2) compared to the NCD model that disregards RT, JRT-NCD achieves higher diagnostic accuracy while maintaining its interpretability, and reduces the misleading effects of "overspeed behavior" on the diagnostic results.

27 Apr 2026

Read appraisal →