Research Appraisals
Evidence-based critical appraisals of the latest medical research, systematically evaluated using Oxford CEBM methodology.
Showing 28 appraisals
Cognitive, affective & behavioral neuroscience
Emotion detection unveiled: A cognitive-computational synthesis of physiological models, machine learning, and datasets.
This comprehensive survey synthesizes state-of-the-art advancements in emotion recognition based on physiological signals, specifically focusing on the paradigm shift occurring between 2021 and 2025. Crucially, we move beyond a technical review by establishing a novel Cognitive-Computational Synthesis Framework (CCSF). This framework explicitly maps multimodal physiological manifestations (e.g., electroencephalogram (EEG), electrocardiogram (ECG), and galvanic skin response (GSR)) to underlying cognitive processes, such as attentional allocation, arousal regulation, and perceptual bias, providing a theoretical foundation for explainable AI (XAI) in affective computing. We meticulously examine the transition from traditional machine learning to advanced deep learning architectures, highlighting how recent innovations in Transformers, self-supervised learning, and diffusion models have shattered previous performance plateaus. While earlier dimensional models were often limited to 70-75% accuracy, this survey details how modern architectures now achieve benchmarks exceeding 95% on seminal datasets like SEED and DREAMER. Furthermore, the survey provides a rigorous analysis of 40 key studies (identified via PRISMA protocols), evaluating them based on their validation strategies, cross-subject generalizability, and adversarial robustness. By bridging the gap between raw physiological data and cognitive theory, this work offers a strategic roadmap for the next generation of robust, interpretable, and real-time emotion recognition systems.
3 Aug 2026
Read appraisal →Cognitive, affective & behavioral neuroscience
A review of current capabilities and future directions in machine-based emotion recognition
The ability of machines to recognize emotions automatically is becoming increasingly significant across many domains where emotional understanding is essential. Such technology is applied in customer interaction, marketing, healthcare, education, the automotive industry, entertainment, and security. Providing real-time insights into human affective states improves user engagement and enables systems to respond more intelligently. Nevertheless, progress in this field is hindered by the inherent complexity of emotions, cultural differences in expression, and technical limitations that make accurate detection challenging. This paper delivers a broad review of contemporary approaches to emotion recognition. It highlights techniques based on facial expression analysis (FER), oculometrics (OM), microexpressions identification (MER), and speech analysis (SER). Further attention is given to methods involving body posture, gesture, and gait, as well as tactile interaction, text-based emotion recognition, and methods based on self-reporting. In addition, physiological signal-driven methods are discussed in depth, including respiration signals (RS), galvanic skin response (GSR), electroencephalography (EEG), electromyography (EMG), skin temperature (SKT), cardiac signals (ECG, PPG, HRV), and touch dynamics (TD) analysis. This comprehensive overview lays the foundation for advancing research on machine-based emotion recognition.
2 Aug 2026
Read appraisal →JMIR mHealth and uHealth
Machine Learning Frameworks for Wearable-Based Stress Modeling in Naturalistic Settings: Scoping Review
BACKGROUND: Stress, as commonly recognized, is an integral part of modern life and can significantly affect both mental and physical health. While substantial advancements have been made in measuring physical fitness through wearable devices, the detection and assessment of mental stress remain in their early stages. OBJECTIVE: The objective of this paper is to review recent studies of wearable-based stress detection in naturalistic settings, with a specific focus on characterizing machine learning frameworks inspired by the model card approach. METHODS: This review was conducted using the PRISMA-ScR (Preferred Reporting Items for Systematic Reviews and Meta-Analyses extension for Scoping Reviews) checklist. A total of 353 articles were identified through searches in databases such as PubMed, MEDLINE, ScienceDirect, IEEE, ACM Digital Library, Web of Science, and Embase. Studies were considered eligible if they collected data from healthy adults in naturalistic settings using wearable devices and used machine learning models for stress detection. RESULTS: A total of 34 articles met the eligibility criteria, including 11 conference papers, 22 journal articles, and 1 preprint published between 2017 and 2024. From these studies, we analyzed key machine learning modeling decisions such as problem formulation, ground truth determination, and machine learning algorithms. Additionally, we examined the major contributions of each study, focusing on the challenges they addressed and the solutions they proposed. Based on these findings, we proposed a model card framework for reporting machine learning-based, wearable-based stress detection. CONCLUSIONS: This scoping review highlights recent trends in machine learning models for stress detection and measurement using wearable signals. It underscores the need for improved standardization in reporting practices for datasets and key machine learning decisions, as well as the importance of addressing critical challenges associated with data collection in real-world settings. We hope this review will support and strengthen ongoing research efforts, promote knowledge sharing, and promote collaboration among researchers-ultimately advancing the field as a community.
1 Aug 2026
Read appraisal →BMJ global health
A systematic review of MHPSS interventions targeting non-clinical Arabic-speaking refugees and/or displaced populations in the MENA region
BACKGROUND: Arabic-speaking refugees and displaced populations in the Middle East and North Africa (MENA) face ongoing adversity that negatively impacts their mental health. Despite increasing mental health needs, there is no published systematic review evaluating non-clinical mental health and psychosocial support (MHPSS) interventions in the MENA region for this population, particularly through a culturally grounded framework. METHODS: We conducted a systematic review of MHPSS interventions published from 2011 to 2022 targeting non-clinical Arabic-speaking adult refugees and/or displaced persons in MENA countries. Following Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) 2020 guidelines and PROSPERO registration (CRD42023421057), we searched ten databases and grey literature sources. Objectives were to identify, appraise and synthesise eligible MHPSS interventions and assess their alignment with the Integrative Complexity (IC)-Adaptation and Development After Persecution and Trauma (ADAPT) framework. Studies were appraised using four validated quality assessment tools and a novel tool assessing links to the IC-ADAPT framework, which integrates an eco-psycho-social model (ADAPT) and a cognitive-interactionist model (IC). A subset of ten top-rated interventions was synthesised for IC-ADAPT alignment. RESULTS: Thirty-eight studies met the inclusion criteria. Most interventions were conducted in Lebanon or Jordan, group-based and delivered by non-specialist facilitators in community settings. Parenting, resilience-building and trauma-focused therapies were the most common modalities. Most studies prioritised feasibility and effectiveness over efficacy. While none explicitly applied IC-ADAPT, several implicitly aligned with its eco-psychosocial systems (eg, bonds/networks, safety/security and roles/identities) and IC (eg, increased 'differentiation'), reporting reduced harsh parenting and improved reflective capacity. Culturally adapted interventions that promoted complex thinking, relational healing and community participation were linked with more favourable outcomes. CONCLUSION: This review highlights the value of culturally sensitive, non-clinical MHPSS interventions in enhancing mental health outcomes for Arabic-speaking refugee populations in MENA. Programmes grounded in eco-psychosocial and cognitive complexity frameworks-such as IC-ADAPT-offer effective, scalable strategies for addressing refugee mental health. Future work should embed these models explicitly to optimise intervention design, relevance and impact. PROSPERO REGISTRATION NUMBER: CRD42023421057.
31 July 2026
Read appraisal →PloS one
Development of a scale for measuring the perception of artificial intelligence among mental health consumers
BACKGROUND: Artificial Intelligence (AI) has emerged as a transformative force revolutionizing various sectors, including healthcare, particularly the mental health field. However, the acceptance and integration of AI technologies in different healthcare systems can be influenced by various factors, including cultural, social, and individual aspects. Nevertheless, there is a need for a valid and reliable tool for assessing AI's perception among healthcare consumers. AIM: To develop and validate a tool for the perception of AI among healthcare consumers and apply the tool to assess AI's perception among mental health consumers in the Jordanian healthcare system. METHOD: A cross-sectional descriptive correlational design was utilized in the study. Data was collected from a convenience sample of 431 mental health consumers visiting mental health clinics of the International Medical Corps and university hospitals in Jordan. Structured interviews were conducted using an AI Perception (AIP) questionnaire developed by the authors. The questionnaire's content validity was assessed by an expert panel. Using Principal Component Analysis (PCA), the construct validity of the tool was evaluated, and its internal consistency was examined using Cronbach's alpha. Descriptive statistics were used to assess the levels of AI perception among participants. RESULTS: The final AIP tool consisted of 20 items across 4 domains and has demonstrated strong internal consistency across its four domains: AI acceptance and readiness (α = 0.92), AI perceived importance (α = 0.92), AI perceived risk (α = 0.9), and AI perceived challenges (α = 0.85). The construct validity of the four-domain structure of the tool was supported by PCA. Additionally, the mean scores for each domain indicated the average level of agreement with AI perception items among participants. Specifically, the mean score for AI acceptance and readiness was [Formula: see text]). AI perceived importance was (2.18 [Formula: see text]0.83), AI perceived risk was (2.58[Formula: see text]0.92), and AI perceived challenge was (2.78 [Formula: see text] 0.87). CONCLUSION: The findings of this study resulted in developing a valid and reliable 20-item tool to assess AI's perception among mental health consumers. The tool can be used to assess the predictors of AI's readiness among mental health consumers. Therefore, aiding policymakers and other stakeholders in understanding the AI adoption barriers from the perspective of end-users. In addition, this study developed the AIP tool that can be validated and used among other populations in future research.
31 July 2026
Read appraisal →Depression and anxiety
Reducing Public Stigma Toward Suicide-Loss Survivors Through Brief Video Interventions: A Randomized Controlled Trial
BACKGROUND: Suicide-loss survivors (SLSs) experience substantial and often enduring psychological burden. These difficulties are compounded by public stigma, underscoring the need for scalable approaches to shift attitudes at the population level. In this study, we examined whether brief survivor-narrative videos can reduce public stigma of SLSs. METHODS: In a randomized controlled trial, 1351 adults (18-50) completed baseline measures and were allocated to view either a brief SLS narrative or a psychoeducational control. Public stigma toward SLSs and trait impressions were assessed at baseline and postexposure. RESULTS: Relative to control, the SLS video arm showed clear improvements immediately and at 30 days: stigma scores were lower and trait impressions more favorable, with attenuation over time. Item-level analyses indicated sizable immediate reductions for "Disconnected" (-24%) and "Cowardly" (-16%), and smaller but significant decreases for "Immoral" and "Irresponsible" (-12% each). CONCLUSIONS: Brief survivor-narrative videos can shift public attitudes toward SLSs and maintain part of that change over 1 month. As a low-cost, scalable complement to postvention, brief video contact offers a practical lever to improve the social climate surrounding SLSs. Deployed widely and reinforced over time, it can move communities from blame to empathy, strengthen everyday support, and advance survivors' recovery.
23 July 2026
Read appraisal →JMIR human factors
Video-Algorithmic Patient Monitoring in Mental Health Inpatient Settings: Qualitative Study of Patient or Consumer, Clinician, and Vendor Perspectives
BACKGROUND: Video-algorithmic patient monitoring (VAPM) combines remote, noncontact sensors and algorithmic analysis and is increasingly trialed in acute psychiatric and other care settings. While promoted for improving safety and reducing risk, it raises ethical concerns regarding safety, privacy and surveillance. Little is known about how those encountering VAPM in mental health care contexts anticipate its use and potential impacts, including where it has not yet been implemented. OBJECTIVE: This study aimed to explore the views of patients or mental health consumers, specialized mental health nurses and nurse academics, hospital managers, and technology vendors regarding the appropriateness and anticipated implications of VAPM in mental health inpatient care. METHODS: This qualitative study identified key stakeholders in Australia via networking techniques for participation in a deliberative workshop. A deliberative workshop was held, and the workshop discussion was audio-recorded, transcribed, and thematically analyzed, consistent with methods in health technology research, which enable exploration of different viewpoints, including convergences and divergences across stakeholder groups. RESULTS: In total, 16 stakeholders participated, exploring themes concerning (1) contestation over the rationale for VAPM in mental health settings, (2) VAPM reshaping care and relationships, (3) perceived harms of VAPM, (4) perceived observational support for safety and reduced disruption, (5) serious privacy implications of VAPM, (6) the need for appropriate governance, and (7) the potential for VAPM to transform, not augment, service delivery. General views differed across groups. Patients or service users expressed concerns about privacy, coercion, and the potential to intensify stigma. Mental health nurses were cautious but interested in possible benefits for safety and suicide prevention. Hospital managers and technology vendors largely emphasized safety gains. CONCLUSIONS: The findings suggest that the anticipated risks of VAPM are primarily experienced subjectively, as infringements on privacy, dignity, and trust, while purported benefits remain largely untested and unquantified. From a utilitarian perspective, direct comparison is therefore difficult-the risks are set out in the anticipated experiences of those with lived experience, and the benefits remain hypothetical. From this view, robust, independent evidence of real-world outcomes is required. Yet, for some participants, the very premise of such calculation was rejected, with privacy, dignity, and trust regarded as nonnegotiable, rather than items for trade-off. If VAPM is to be pursued at all, it should proceed only with extreme caution, with transparent evidence of outcomes, and with meaningful participation from those whose lives and care are most directly impacted.
22 July 2026
Read appraisal →Journal of global health
Animal-assisted therapy on psychological and physical outcomes: a meta-analysis of randomised controlled trials
BACKGROUND: Adults are at heightened risk of anxiety, stress, and depression; animal-assisted therapy (AAT) may serve as an effective approach to promote psychological well-being. In this study, we compared the effectiveness of AAT in improving depression, anxiety, stress, pain, and gait among adults with or without illness. METHODS: We systematically searched six electronic databases (PubMed, CINAHL, Embase, Web of Science, the Cochrane Library, and Scopus) and included all studies published up to August 2024. We used comprehensive meta-analysis software to complete the quantitative synthesis. We reported pooled effect sizes as Hedges' g with corresponding 95% confidence intervals (CIs), after applying a random-effects model. Furthermore, we assessed heterogeneity using Cochran's Q test and the I2 statistic. We applied the Cochrane Risk of Bias 2.0 tool to appraise the methodological quality of the included studies. The synthesis process followed PRISMA guidelines. RESULTS: From the 13,345 studies identified, 35 randomised controlled trials involving 2391 adults were included. Across diverse populations, AAT was associated with reductions in depression (Hedges' g = -0.403; 95% CI = -0.536, -0.271), anxiety (Hedges' g = -0.661; 95% CI = -1.069, -0.253), and stress (Hedges' g = -1.062; 95% CI = -1.849, -0.275) at post-intervention, although substantial between-study variability was observed. CONCLUSIONS: We demonstrated that AAT significantly improves anxiety, depression and stress in adults, but has no meaningful effect on pain or gait. Subgroup and meta-regression analyses indicate that psychological benefits depend on population and intervention characteristics, and are not moderated by age or gender. REGISTRATION: PROSPERO: CRD42024570108.
19 July 2026
Read appraisal →Journal of behavioral addictions
Efficacy of a mobile-based approach-avoidance task training (PROTECTapp) for problematic usage of the internet in young adults: A randomized controlled trial
BACKGROUND AND AIMS: Problematic usage of the internet (PUI) has been linked to impaired mental health and academic functioning in young adults. This randomized controlled trial evaluated the efficacy of a 3-week mobile-based approach-avoidance task (AAT) training (PROTECTapp) for reducing PUI in university students. METHODS: Ninety-two participants (Mage = 22.00 years, 69.6% women) with elevated levels of PUI were randomized to the PROTECTapp intervention (n = 45) or a waitlist control group (n = 47). Primary outcomes were PUI severity and internet-related craving. Secondary outcomes included motivation to change, psychopathological symptoms and academic functioning. Participants were assessed at baseline and postintervention; the intervention group completed additional 3- and 12-week follow-ups. RESULTS: Intention-to-treat analyses indicated greater reductions in PUI following the PROTECTapp intervention compared to the waitlist (p = .003; d = -0.80, 95% CI [-1.23, -0.38]). No significant effects emerged for craving or broader psychological outcomes (ps > .05), though favorable effects were observed on motivation to change (ambivalence: p = .020; d = -0.24, 95% CI [-0.65, 0.17]; taking steps: p = .002; d = 0.45 [0.04, 0.87]). Satisfaction with the intervention was moderate (M = 18.32 of 32), and participants completed on average 35.52 training sessions. Adverse events were reported infrequent (7.1%). DISCUSSION AND CONCLUSIONS: PROTECTapp is a promising mobile-based intervention to reduce PUI and enhance motivation to change in young adults. Its brevity, scalability, and safety profile highlight its potential as a low-threshold preventive or adjunctive intervention for young individuals at-risk.
18 July 2026
Read appraisal →European journal of cardiovascular nursing
The psychological effects of interventions targeting informal caregivers of patients with cardiovascular disease-a systematic review.
AIMS: To explore the psychological effects of interventions aimed at supporting informal caregivers involved in the care and treatment of patients with cardiovascular disease. METHODS AND RESULTS: Databases (PubMed, CINAHL, Embase, Cochrane Library, and PsycInfo) were searched for studies in accordance with the Cochrane Handbook guidelines. Inclusion criteria were: caregivers of patients with one or more cardiovascular diseases, patient and caregiver >18 years, caregivers included in the intervention, and, reporting of psychological outcomes specific to caregivers. Study designs were randomized controlled trials with a follow-up periods of >2 months. The RoB 2.0 bias assessment tool was used to assess risk of bias. Fifteen studies from nine countries were identified. Most interventions consisted of multiple components including educational face-to-face sessions, telephone support, and/or written resources. The analysis showed inconsistent results on caregiver outcomes, but significant improvements were reported in eight studies in at least one caregiver outcome. Analysis indicated that the studies providing more frequent contact with caregivers and patients were more likely to report significant improvements in caregiver outcomes. Risk of bias was judged as low in three studies, some concerns in nine studies, and as high in three studies. CONCLUSION: Due to inconsistency in results, this review yield uncertainty about whether interventions targeting caregivers of patients with cardiovascular disease improve caregivers' psychological outcomes. Therefore, further research is needed to develop effective interventions for caregivers, as they play a vital role in the daily care and treatment of patients with cardiovascular disease. REGISTRATION: PROSPERO: CRD420250654618.
17 July 2026
Read appraisal →Depression and anxiety
A Systematic Review of Community Pharmacy-Led Depression Services: Service Components, Outcomes, and Implementation Barriers and Facilitators
BACKGROUND: Depression is the most common mental ill health condition, and its prevalence is increasing. Despite this, its treatment is variable often due to a lack of capacity within healthcare systems to support this vulnerable population. Community pharmacy staff could offer additional support. This systematic review identifies depression services led by community pharmacy staff, their service components, outcomes, and barriers/facilitators to their implementation. METHODS: Four bibliographic databases were searched (Medline, EMBASE, PsycINFO, and CINAHL) from 2000 onwards. Title/abstract and full-text screening was conducted. Data on the service components were mapped to the Template for Intervention Description and Replication (TIDieR). Clinical, humanistic, economic, and service outcomes were charted. Barriers and facilitators were mapped to the Consolidated Framework for Implementation Research (CFIR). Quality assessment was performed using the Quality Assessment with Diverse Studies (QuADS) tool. RESULTS: Fifty studies were included. Seventeen studies identified general attitudes regarding community pharmacy services for depression, which were generally supportive. The majority (n = 33) explored an implemented depression service focusing on depression advice/education (n = 15), screening (n = 12), medication adherence (n = 4), medication review (n = 1), and disease therapy management (DTM) (n = 1). Clinical outcomes were the most commonly reported types of outcomes, with varied results. Key facilitators were linked to the pharmacy 'inner setting', including accessibility of community pharmacies, the use of private consultation rooms, and skills/training of staff. Barriers to service delivery related often to the external 'outer setting', especially societal stigma, low public awareness of pharmacy roles, funding constraints, and limited collaboration with other healthcare professionals. CONCLUSION: This international review identified a range of different services that community pharmacy staff can deliver to support people with depression, ranging from supporting diagnosis, health literacy, and management plans. The accessibility of community pharmacies for depression service delivery warrants further investigation. However, limited empirical evidence of clinical and economic outcomes and reported implementation barriers may complicate broader implementation.
13 July 2026
Read appraisal →JMIR mHealth and uHealth
Perceived Sensitivity of Sensor-Based Digital Health Data: Qualitative Interview Study
BACKGROUND: Digital health tools are increasingly used in mental health care to passively collect patient data and analyze health status outside of clinical settings. While technologies such as digital phenotyping, affective computing, and computational behavioral analysis offer new insights into symptom manifestation in daily life, they generate large volumes of potentially sensitive data that raise significant data privacy concerns, requiring high levels of patient awareness and consent. Empirical research is lacking on stakeholder understandings toward the sensitivity of these data and expectations for data stewardship, perspectives that are critical for developing robust informed consent and data protection policies for digital health data use. OBJECTIVE: This study aimed to explore key stakeholder perspectives on the sensitivity of computer perception (CP) data, trust in existing data protections, willingness to share CP data externally, and desire for transparency of CP data transactions outside of the clinical space. METHODS: As part of a larger, multisite study, we conducted qualitative interviews (n=40) via Zoom (Zoom Communications, Inc) with 20 adolescents (aged 12-17 years) familiar with CP tools and their caregivers (n=20). Interviews consisted of a series of open-ended questions regarding stakeholders' perspectives on privacy, data security, and the use and exchange of CP data. We developed a qualitative codebook to identify and label thematic patterns in responses to questions addressing the topics above, using thematic content analysis to identify themes inductively. Each interview was coded by merging work from at least two separate coders, and several team members contributed to qualitative analysis. RESULTS: Most adolescents and caregivers viewed CP data as highly sensitive and expressed a reluctance to share these data beyond their clinical teams. While many participants expressed trust in existing data protections to protect CP data, they often misunderstood or overestimated the extent of protections to safeguard CP data. CONCLUSIONS: Our findings underscore the critical need for clear and effective patient communication and education about the risks, benefits, and protections associated with CP data through informed consent protocols. To promote greater transparency, understanding, and trust, we recommend 5 strategies: educating patients about data protection; studying secondary data exchange and reidentification risks; strengthening transparency regulations; improving data traceability mechanisms, such as distributed ledger technologies, to enhance data traceability and auditability; and adopting dynamic consent models.
11 July 2026
Read appraisal →BMJ open
The Cyber Paranoia and Fear Scale-Updated (CPFS-U): development and implications for digital health engagement
OBJECTIVES: To update and revalidate the Cyber Paranoia and Fear Scale to reflect current technological contexts and examine its relevance to digital health readiness and engagement. METHODS: Using an online community sample (n=433), exploratory factor analysis was conducted to examine the factor structure of the revised item pool. Items were refined through consultation with Patient and Public Involvement and Engagement groups to ensure contemporary relevance and clarity. RESULTS: Analysis supported a four-factor structure representing artificial intelligence (AI) and digital dependence, technological risk awareness, perceived data vulnerability and surveillance-related mistrust. The updated Cyber Paranoia and Fear Scale-Updated (CPFS-U) demonstrated good internal consistency and supported construct validity. Cyber-paranoia and fear were conceptually and empirically distinct from general paranoia and anxiety, highlighting the specific cognitive and emotional responses elicited by digital technologies. CONCLUSIONS: The CPFS-U offers a psychometrically robust, modernised measure for understanding individuals' responses to digital and AI-based technologies. Its application in digital health research and practice can inform risk communication, user engagement strategies and the design of trustworthy digital interventions. By identifying individuals who may disengage due to online mistrust, the CPFS-U has the potential to inform more inclusive and psychologically informed digital health systems.
8 July 2026
Read appraisal →Philosophy, ethics, and humanities in medicine : PEHM
Moral diversity and the challenge of responsibility in AI-CDSS
The increasing integration of artificial intelligence in clinical decision support systems (AI-CDSS) has fueled expectations of more personalized and effective diagnostics and therapies. By incorporating machine learning methods, AI-CDSS promise enhanced predictive accuracy, improved stratification, and innovative individualized care. However, this technological optimism is accompanied by complex ethical challenges, including issues of explainability, trust, autonomy, and data security. At the core of these debates lies the question of responsibility, which involves both its attribution and diffusion, as well as the underlying normative standards guiding moral action. In the context of healthcare practice, responsibility is further complicated by moral diversity-the coexistence of varying moral values, cultural beliefs, and ethical frameworks among healthcare professionals, patients, and institutional stakeholders. This plurality challenges the establishment of a unified normative standard necessary for ethically sound responsibility attribution. This paper offers an analysis of moral diversity and AI-CDSS as a challenge for responsibility in healthcare environments. Using a relational concept of responsibility the study examines key areas in which moral diversity affects responsibility in AI-mediated decision-making. This includes algorithmic bias, healthcare professional and patient interaction and the role of patients. Through these examples, the paper explains how different normative standards intensify ethical complexity in AI-supported clinical contexts. It argues that greater ethical sensitivity to moral diversity is essential-both in the development of AI-CDSS and in their application within morally value-laden healthcare situations.
4 July 2026
Read appraisal →Behavior research methods
A validity-guided workflow for robust large language model research in psychology.
Large language models (LLMs) are rapidly being integrated into psychological and behavioral research as research tools, evaluation targets, human simulators, and cognitive models. Yet recent evidence reveals severe measurement unreliability: personality assessments degenerate under factor analysis, moral preferences reverse with punctuation changes, and theory-of-mind accuracy varies widely with trivial rephrasing. These "measurement phantoms"-statistical artifacts masquerading as psychological phenomena-threaten the validity of a growing body of research. Guided by the dual-validity framework that integrates psychometrics with causal inference, this article presents a six-stage workflow that scales validity requirements to research ambition-using LLMs to code text requires basic reliability and accuracy, whereas claims about psychological properties demand comprehensive construct validation. Researchers must (1) explicitly define their research goal and corresponding validity requirements, (2) develop and validate computational instruments through psychometric testing, (3) design experiments that control for computational confounds, (4) execute protocols transparently, (5) analyze data with methods appropriate for nonindependent observations, and (6) report findings within boundaries and use results to refine theory. The workflow is illustrated through an example of model evaluation-"LLM selfhood"-showing how systematic validation can distinguish genuine computational phenomena from measurement artifacts. By establishing validated computational instruments and transparent practices, this workflow provides a path toward building a robust empirical foundation for AI psychology research.
3 July 2026
Read appraisal →Journal of medical Internet research
How Does That Large Language Model Make You Feel?
People are increasingly turning to commercially available large language models (LLMs) for emotional support. In this News and Perspectives article, JMIR Correspondent Simon Spichak reports on the role of LLMs in mental health, speaking with experts about safety concerns, research gaps, and next steps.
2 July 2026
Read appraisal →The Cochrane database of systematic reviews
Views and experiences of weight management for people living with mobility-limiting conditions, intellectual disabilities or severe mental illness: a qualitative evidence synthesis
This is a protocol for a Cochrane Review (qualitative). The objectives are as follows: This qualitative evidence synthesis (QES) aims to address the question: What is known from qualitative evidence about behavioural weight management for people living with mobility-limiting conditions, intellectual disabilities, or severe mental illness? The five review objectives are as follows. To explore individuals' experiences with weight and behavioural weight-management programmes. This includes understanding how individuals living with mobility-limiting conditions, intellectual disabilities, severe mental illness and their carers experience managing their weight and how they perceive and engage with behavioural weight-management programmes, including how accessible, supportive, and relevant they find the programme elements. To identify barriers to and facilitators of managing a healthy weight and accessing behavioural weight-management programmes. Focusing on the unique obstacles encountered by each population group, we will examine issues such as physical accessibility, communication challenges, financial constraints, and gaps in tailored support that may prevent equitable participation. We will also include people's perspectives on structural and system-level barriers, as well as the influence of intersectional factors (e.g. gender and ethnicity), on the challenges of maintaining weight and accessing and benefiting from support. To understand the impact of stigma and discrimination on programme engagement and outcomes. This objective seeks to uncover how stigma related to weight, disability, or mental health influences individuals' willingness to participate, their inclusion and sense of belonging with others within the programmes, and overall programme satisfaction. It also examines the role of stigma as a potential barrier to success. To capture the perspectives of programme providers and healthcare practitioners. Gathering insights from healthcare providers and programme facilitators, this objective assesses which programme elements are considered successful and which present ongoing challenges. This includes evaluating aspects like accessibility, engagement strategies, and support mechanisms that may require improvement. To produce an inclusive piece of research through incorporating the insight of people with lived experience and professional expertise in interpreting and considering the implications of the evidence. By synthesising this qualitative evidence, the review aims to provide a comprehensive understanding of how to design and implement behavioural weight-management programmes that are more inclusive, effective, and responsive to the needs of these populations.
29 June 2026
Read appraisal →eLife
Disinformation elicits learning biases.
In open societies, disinformation is often considered a threat to the very fabric of democracy. However, we know little about how disinformation exerts its impact, especially its influence on individual learning processes. Guided by the notion that disinformation exerts its pernicious effects by capitalizing on learning biases, we ask which aspects of learning from potential disinformation align with ideal 'Bayesian' principles, and which exhibit biases deviating from these standards. To this end, we harnessed a reinforcement learning framework, offering computationally tractable models capable of estimating latent aspects of a learning process as well as identifying biases in learning. In two experiments, participants completed a two-armed bandit task, where they repeatedly chose between two lotteries and received outcome-feedback from sources of varying credibility, who occasionally disseminated disinformation by lying about true choice outcome (e.g., reporting non-reward when a reward was truly earned or vice versa). Computational modelling indicated that learning increased in tandem with source credibility, consistent with ideal-Bayesian principles. However, we also observed striking biases reflecting divergence from idealized Bayesian learning patterns. Notably, in one experiment individuals learned from sources that should have been ignored, as these were known to be fully unreliable. Additionally, the presence of disinformation elicited exaggerated learning from trustworthy information (akin to jumping to conclusions) and exacerbated a normalized measure of 'positivity bias' whereby individuals self-servingly boost their learning from positive, relative to negative, choice feedback. Thus, in the face of disinformation we identify specific cognitive mechanisms underlying learning biases, with potential implications for societal strategies aimed at mitigating its harmful impacts.
28 June 2026
Read appraisal →BMC palliative care
Factors associated with psychological distress among end-of-life care volunteers: a systematic review of quantitative and qualitative evidence
BACKGROUND: Volunteers are integral to end-of-life care, providing emotional, spiritual, and practical support. However, they often face emotionally demanding situations with limited training and supervision compared to professionals. Given the limited and fragmented literature on psychological distress experienced by end-of-life volunteers, this systematic review aimed to synthesise existing quantitative and qualitative evidence to identify factors associated with psychological distress. METHODS: We conducted a systematic literature review including qualitative and quantitative evidence. Five databases (MEDLINE, EMBASE, PsycINFO, Cochrane Database and Web of Science) were searched for original studies, complemented by citation and reference searches. Study quality was assessed using the Qualsyst tool. Quantitative findings were synthesised using an algorithm to evaluate evidence strength, and qualitative data were integrated through thematic meta-synthesis. RESULTS: Twenty-six studies (20 quantitative and 6 qualitative studies) met inclusion criteria. Quantitative research examined 49 volunteer-related, 18 service-related, and one volunteer-patient-interaction-related factor associated with anxiety, death anxiety, depression, burnout, and/or perceived stress. Moderate-strength evidence indicated that death anxiety was negatively associated with better health and well-being but unrelated to age, volunteer experience, or training. Furthermore, depression was negatively associated with volunteer training. Qualitative evidence was scarce, but highlighted additional patient-, interaction-, and service-level mechanisms. CONCLUSION: This review identifies a small, methodologically diverse evidence base on factors associated with psychological distress in end-of-life care volunteers. Quantitative evidence suggests a potential protective association between training and depression, though substantial heterogeneity limits firm conclusions. Limited qualitative evidence revealed patient-, interaction- and service-level factors that are rarely captured quantitatively. Robust theory-guided longitudinal studies are needed to better understand distress and resilience in this under-researched group.
28 June 2026
Read appraisal →The Keio journal of medicine
Comparison of Clinical Development of Digital Therapeutics among Japan, the United States, and Germany
Digital therapeutics (DTx) have demonstrated promising potential as a novel therapeutic approach, with their clinical development gaining momentum. Comparison of DTx development between Japan, a relative newcomer, and Germany and the U.S., known for favorable circumstances, is a topic of interest. DTx developed in Japan as of November 2023 were identified on the Pharmaceuticals and Medical Devices Agency (PMDA) website and in the literature. Similar DTx available in the U.S. and Germany were identified in the databases maintained by the U.S. Food and Drug Administration (USFDA) and Federal Institute for Drugs and Medical Devices (BfArM). Data on their clinical trials and regulatory status were obtained from these databases and their respective national clinical trial registries. The data obtained were compared and analyzed. By November 2023, Japan had developed DTx, encompassing 11 therapeutic areas. Seven DTx had reached or completed the confirmatory study stage, whereas five DTx were in the exploratory phase. Twenty DTx in the therapeutic areas were selected from the U.S. and Germany. A total of 27 DTx were reviewed regarding the designs of confirmatory studies and regulatory actions taken. Trial designs demonstrated more similarities than differences across the countries and therapeutic areas. Placebo trial design-associated difficulties were conspicuous. Although regulatory actions to place DTx in the market differed across countries, they effectively ensured the efficacy and safety of DTx that were suitable for marketing considering the most current science. The regulatory measures in the three countries seem to have positively impacted DTx development.
27 June 2026
Read appraisal →Scientific reports
Multimodal emotion recognition using hybrid deep feature fusion under speaker-independent evaluation
Emotion recognition is one of the most important and complex challenges for machines to understand, as most robots and AI agents struggle with human-centric perception and interpretation. Therefore, this paper introduces a novel multimodal emotion recognition system that analyzes emotions through two complementary channels: voice and facial expressions. The proposed approach is evaluated on the RAVDESS and CREMA-D datasets, which consist of acted emotional expressions across multiple discrete emotion categories. Utilizing an advanced multimodal deep feature fusion technique, the system combines handcrafted audio features (e.g., Mel-Frequency Cepstral Coefficients (MFCCs)) with deep visual features extracted from an attention-based VGGFace model. These features are integrated into a unified representation through a hybrid fusion strategy that jointly employs concatenation, cross-attention, gated fusion, and multiplicative fusion mechanisms to capture complementary cross-modal interactions. To ensure a comprehensive and realistic assessment, the model is evaluated under both random-split and strict speaker-independent protocols. On the RAVDESS dataset, the proposed system achieves an accuracy of 95.83% under random-split evaluation and 48.06% ± 9.76% accuracy under speaker-independent Leave-One-Speaker-Out (LOSO) testing, while on the CREMA-D dataset it attains 73.54% accuracy using random splits and 53.12% ± 2.65% accuracy under subject-exclusive speaker-independent 5-fold cross-validation.
27 June 2026
Read appraisal →JMIR formative research
Supporting Student Mental Health With the Safespace Generative AI Chatbot: Mixed Methods Feasibility Study
BACKGROUND: Generative artificial intelligence (GenAI) chatbots have the potential to provide personalized mental health support to individuals at scale. OBJECTIVE: This study evaluates the feasibility and usage patterns of the Safespace GenAI chatbot, an artificial intelligence (AI)-driven smartphone app that offers a large language model-powered interactive chatbot to support mental health. METHODS: Using a mixed methods approach, we explored baseline attitudes toward GenAI chatbots and chatbot usage patterns, conducted a qualitative content analysis of participants' experiences, and descriptively assessed patterns related to preintervention depressive symptoms. The study included an initial sample of 42 university students, 20 of whom actively used the chatbot over 2 to 4 weeks, generating 286 user-chatbot interactions. RESULTS: Preintervention surveys indicated that the majority of participants anticipated that the chatbot would be helpful (27/42, 64%) and that they trusted its privacy safeguards (39/42, 93%). Usage patterns suggested that the highest levels of interaction occurred early in the morning and late at night, when peer and professional support may be inaccessible. The qualitative analysis indicated that participants appreciated using the chatbot for reflection as a blended-care tool between their counseling sessions, while also naming technical barriers and specific design needs required to sustain engagement. In addition, our exploratory analyses descriptively showed that participants with elevated depression scores engaged in emotional disclosure during 99% (38 sessions with 8 participants) of their sessions, compared to 84% (26 sessions of 12 participants) of those with low symptoms. Due to the small sample size, future adequately powered studies are needed to inferentially examine these observed patterns. CONCLUSIONS: These findings provide initial insights into the usage and engagement dynamics of the Safespace GenAI chatbot and highlight directions for future research to optimize GenAI-driven mental health interventions.
27 June 2026
Read appraisal →BMJ open
Caregiver experiences of social isolation and loneliness in chronic kidney disease: systematic review of qualitative studies.
OBJECTIVES: This qualitative systematic review aimed to describe the experiences and perspectives of loneliness and social isolation among informal caregivers of people with chronic kidney disease (CKD). METHODS: Terms for caregivers, qualitative research, loneliness and social isolation and CKD were entered into MEDLINE, Embase, CINAHL and PsycINFO and searched from inception to May 2025. Qualitative studies that described social isolation and loneliness among caregivers of people with CKD were included in this review. Study characteristics were extracted into Microsoft Excel. Qualitative data from each study were imported into HyperRESEARCH and analysed using thematic synthesis. RESULTS: We included 19 articles involving 598 caregivers of people with CKD from 28 countries. Four major themes with subthemes were identified: confined by the patient's needs (social withdrawal due to unrelenting demands, separated from communities due to treatment, torn between work and caregiving responsibilities, foregoing social outings with family and friends, restricted by patient's diet, feeling protective against infection risk); limited care assistance exacerbating isolation (inadequate familial support, absence of respite care, lacking accessible guidance from health professionals); disrupting relationships and social roles (sacrificing social needs and identity, family conflicts worsening isolation, withdrawing to avoid stigma and ridicule, grappling with hopelessness, reluctance to share struggles); and coping with support resources (connecting with other families, seeking assistance from support services, finding support through faith). CONCLUSIONS: Caregivers of people with CKD experience restricted social participation and loss of social roles and identity, which can exacerbate feelings of loneliness and social isolation. Support services are needed to prevent and address social isolation and loneliness in CKD caregivers. PROSPERO REGISTRATION NUMBER: CRD420250637194.
26 June 2026
Read appraisal →Journal of medical Internet research
The Emerging Roles of AI in Self-Directed Stress Management: Systematic Review
BACKGROUND: Stress is widespread and carries substantial mental health, social, and economic burdens. Yet, access to clinician-led stress management remains constrained by service capacity, cost, and stigma. In response, artificial intelligence (AI)-enabled tools have rapidly proliferated as scalable, self-directed options. However, evidence on how these systems support stress management outside formal clinical settings remains fragmented. OBJECTIVE: This systematic review aimed to synthesize empirical evidence on how AI-enabled technologies are used for self-directed stress management. We mapped the emerging functions of these tools, the psychological frameworks informing their design, the populations and settings studied, and the outcomes reported. METHODS: We conducted a PRISMA (Preferred Reporting Items for Systematic Reviews and Meta-Analyses)-compliant systematic review of English-language studies published between 2000 and 2025. Six databases were searched (APA PsycINFO, PubMed, MEDLINE, Scopus, Web of Science Core Collection, ProQuest, and Google Scholar). RESULTS: Of 3008 records identified, 35 studies met the inclusion criteria. The methodological quality of included studies was critically appraised using the Mixed Methods Appraisal Tool (version 2018). Findings illustrated that AI-supported stress management can operate through 5 core functions, including psychological intervention, behavioral support, psychoeducation, companionship, and emotional support, and stress monitoring, detection, and triage. Across the reviewed studies, these functions supported self-directed stress management by helping users identify stress, regulate responses, and engage in coping outside formal clinical care. CONCLUSIONS: AI-enabled systems show preliminary promise for supporting self-directed stress management through multiple user-facing functions grounded in established psychological frameworks.
26 June 2026
Read appraisal →Nicotine & tobacco research : official journal of the Society for Research on Nicotine and Tobacco
Efficacy of Yoga in Tobacco Cessation: A Systematic Review and Meta-Analysis
INTRODUCTION: Tobacco consumption is a global epidemic with a high relapse rate. A complementary method like yoga, recognized for addressing psychological and behavioral aspects of addiction, was evaluated in this systematic review and metaanlaysis. METHODS: The objective was to comprehensively summarize the current evidence and evaluate the efficacy of yoga in tobacco cessation. This review was registered with PROSPERO (CRD42025643806). A systematic search across four databases identified randomized controlled trials (RCTs) in English up to September 2024. Eligible studies included adults (≥18 years) using any form of tobacco excluding vaping, randomized to yoga as monotherapy or adjunct therapy. The primary outcome was 7-day point prevalence abstinence (7PPA), while the secondary outcomes included quality of life, depression, anxiety, and mood states. Two reviewers independently extracted data and assessed quality using the Cochrane Risk of Bias (RoB2) tool. A meta-analysis was conducted for outcomes assessed in at least two studies. RESULTS: Seven RCTs were included, with five suitable for meta-analysis. The pooled odds ratio for 7PPA at the end of treatment was 1.50 (95% confidence interval = 0.60 to 3.73), suggesting a positive but inconsistent effect. Considerable heterogeneity was observed (I2 = 65%, p = .07; τ2 = 0.7127), indicating substantial variability across studies. Heterogeneity (I2 = 65%) markedly reduced to 16% when high-risk trials were excluded in sensitivity analysis. Studies reported improvements in depression, anxiety, and quality of life. CONCLUSION: This review consolidates the best available evidence on the effects of yoga interventions for tobacco cessation, showcasing the promising potential of yoga in smoking cessation. IMPLICATIONS: Active yoga styles like hatha, vinyasa, and Iyengar improved 7PPA by reducing stress and depression, while pranayama reduced cravings and negative affect. Future research should explore yoga-based cessation programs in diverse settings, especially low- and middle-income countries, to address unique challenges. Standardizing methodologies across populations will enable a more comprehensive evaluation of yoga's role in tobacco cessation and its potential to enhance global public health outcomes.
25 June 2026
Read appraisal →Journal of medical Internet research
You Can't Launch This: Trust as Infrastructure in Digital Behavioral Health
Digital behavioral health interventions frequently fail to scale, even when evidence-based and technically and operationally sound. In this News and Perspectives article, researcher, digital behavioral health platform founder, and JMIR Correspondent Trevor van Mierlo concludes a four-part series examining why this occurs, reporting on the foundational role of trust.
25 June 2026
Read appraisal →Scientific reports
Classifying mental stress from eye tracking data: deep learning approaches for out-of-the-lab conditions
Eye-tracking signals such as pupil diameter and gaze behavior have been widely used for stress detection, yet most approaches rely on task-specific features, controlled laboratory settings, or multimodal sensor combinations, limiting scalability in less controlled environments. This work investigates whether unimodal eye-tracking time-series data can support task-agnostic stress detection beyond static laboratory tasks. We analyze stress classification across two complementary datasets: a virtual reality goalkeeper task with moderate visuomotor activity and stable recording conditions, and a virtual job interview dataset reflecting less controlled settings with uncalibrated signals. The results show that these signals alone contain informative patterns related to stress-associated autonomic and oculomotor responses. Under favorable conditions, performance reaches up to [Formula: see text] macro-averaged F1-score. At the same time, performance varies substantially across datasets, indicating that effective learning depends strongly on data quality, calibration, signal characteristics, and task design. Overall, the findings demonstrate the potential of unimodal eye tracking as a lower-burden alternative to more complex multimodal systems, while highlighting that reliable stress detection is fundamentally conditioned by the interplay of data, signal representation, and modeling approach.
23 June 2026
Read appraisal →Progress in biomedical engineering (Bristol, England)
Beyond the lab: a review on neurophysiological mental states assessment in real-world settings.
Recent advancements in wearable, non-invasive neurophysiological sensors have increased interest in applying Human factors (HF) research beyond controlled laboratory settings. HF research aims to objectively quantify individuals' (alone or working in team) mental and emotional states to ensure safety, maintain good performances and prevent risks. These devices enable real time, unobtrusive monitoring of individuals in real-world environments, overcoming the limitations of traditional subjective assessments. Unlike self-reports, neurophysiological signals provide objective, real-time data, allowing for a more accurate and continuous understanding of mental states. This review examines the latest advancements over the past ten years in real-time monitoring of the most studied and operationally relevant mental states, including mental workload (MW), stress, attention, fatigue, drowsiness, and teamwork, by analyzing studies that contribute to research advancing toward real-world application direction. While wearable biosensors were used as one criterion for selecting studies, this review extends beyond device usage to explore crucial methodological aspects relevant to real-world applications, such as data quality, customized processing steps, and the robustness of results compared to traditional laboratory-grade devices. Thus, this review analyzes 123 articles exploring current advancements in real-time monitoring of MW, stress, attention, mental fatigue, drowsiness, and teamwork using wearable devices. The findings provide a comprehensive overview of methodology advancement toward real-time, objective individuals monitoring in real world settings.
20 June 2026
Read appraisal →