Research AppraisalSystematic Review

A scoping review on conversational AI in mental health: A human-centered perspective.

International journal of medical informaticsLi, Jingwei, Li, Huiran, Zhu, Hongwei et al.15 July 2026DOI

Clinical Snapshot

35CEBM
Evidence: WeakSystematic Review

PICO Framework

P — PopulationIndividuals across the mental health patient journey (prevention, detection, intervention, maintenance), including general population, patients with mood disorders, anxiety, stress-related disorders, and other mental health conditions
I — InterventionConversational artificial intelligence systems, including large language models (LLMs), chatbots, and dialogue agents applied to mental health support and care
C — ComparatorNo formal comparator specified; the review maps the research landscape rather than comparing interventions head-to-head
O — OutcomesResearch foci distribution across patient journey stages, mental disorder types studied, target populations, AI technologies employed, data sources used, evaluation metrics applied, and a synthesised human-centered design taxonomy

Bottom Line

This scoping review provides the most comprehensive landscape mapping to date of conversational AI in mental health, synthesising 677 studies from 10,293 records across fifteen databases. Its principal contribution is a human-centered taxonomy comprising four dimensions — emotional sensitivity, user-centric design, human-AI collaboration, and ethics and accountability — intended to guide future development. However, clinicians should interpret findings cautiously. The review maps what research exists, not what works: no efficacy data are synthesised, no quality appraisal of included studies was performed, and two-thirds of the evidence base originates from computer science rather than clinical research. Critical gaps identified — minimal research on prevention and maintenance stages, limited population diversity, heavy reliance on text-based data, and inconsistent evaluation metrics — confirm that conversational AI in mental health remains a nascent field with significant clinical translation challenges. The taxonomy offers a useful framework for procurement and design decisions, but should not be interpreted as evidence of clinical effectiveness. For Australian clinicians and health services considering AI-assisted mental health tools, this review underscores the need for rigorous clinical evaluation, TGA regulatory compliance, and explicit attention to equity for underserved populations before widespread adoption.

Evidence: Weak

Key Findings

  • Effect Size: Not applicable — scoping review with no pooled quantitative synthesis

  • Primary Outcome: Descriptive mapping of 677 studies across the mental health patient journey: prevention (8%), detection (23%), intervention (66%), maintenance (3%). LLMs are the predominant AI technology, particularly in intervention and maintenance stages. Mood, anxiety, and stress-related disorders are most investigated. A four-dimension human-centered taxonomy was developed: Emotional Sensitivity to Users, User-Centric Interaction Design, Human-AI Collaboration and Capability Enhancement, and Ethics and Accountability (13 sub-dimensions total).

  • Nnt Or Sensitivity: Not applicable — no clinical efficacy data synthesised; evaluation metrics reported to vary significantly by discipline, reflecting limited cross-disciplinary standardisation

  • Confidence Interval: Not applicable — no meta-analytic pooling performed

Clinical Application

The review does not evaluate clinical implementation feasibility directly. However, the identified gaps — particularly in multimodal data integration, population inclusivity (beyond text-literate, digitally engaged users), and maintenance-stage support — highlight substantial barriers to real-world deployment. The dominance of computer science research suggests many tools have not been evaluated in clinical settings with appropriate patient populations. In Australia, conversational AI mental health tools operate in a complex regulatory and funding environment. The TGA has begun developing guidance on AI-based Software as a Medical Device (SaMD), and tools making therapeutic claims require regulatory oversight. The PBS does not currently fund AI-based mental health interventions. RACGP guidelines emphasise shared decision-making and continuity of care — values aligned with the review's human-centered taxonomy but not yet operationalised for AI tools. The identified gaps in ethics, accountability, and equitable access are particularly salient in the Australian context given the mental health needs of rural and remote populations, Aboriginal and Torres Strait Islander communities, and culturally and linguistically diverse groups — populations largely absent from the current evidence base as described. Headspace, Beyond Blue, and state-based digital mental health platforms represent potential implementation contexts, but clinical governance frameworks for AI integration remain underdeveloped. The taxonomy and gap analysis are relevant to clinicians, health informaticians, and policymakers involved in designing, procuring, or evaluating conversational AI tools for mental health across all age groups and disorder types. Most directly applicable to those working in mood disorders, anxiety, and stress-related conditions, which dominate the current evidence base.

Abstract

BACKGROUND: Conversational AI offers scalable mental health support, with large language models (LLMs) enabling personalized interactions. Human-centered design is critical in this domain, yet a comprehensive synthesis from this perspective is lacking. This review maps conversational AI research in mental health across the patient journey and develops a human-centered taxonomy to guide future design. METHODS: Following PRISMA guidelines, we conducted a comprehensive search across fifteen multidisciplinary databases. We systematically analyzed the literature across six dimensions: research foci, mental disorder types, target populations, AI technologies, data sources, and evaluation metrics. A consensual taxonomy research method was employed to develop a human-centered design framework. RESULTS: Of 10,293 identified records, 677 studies met the inclusion criteria. Analysis reveals a marked increase in publications since 2020, predominantly from computer science (449 studies), followed by medicine (148) and social sciences (80). Research is skewed toward detection (23%) and intervention (66%) stages, with prevention (8%) and maintenance (3%) receiving less attention. Mood, anxiety, and stress-related disorders are the most investigated conditions. LLMs have emerged as the predominant AI technology, particularly within intervention and maintenance stages. Data sources continue to rely heavily on text-based inputs, with multimodal approaches still limited in adoption. Evaluation metrics vary significantly by discipline, reflecting limited cross-disciplinary integration. Through thematic synthesis, we developed a human-centered taxonomy comprising four primary dimensions: Emotional Sensitivity to Users, User-Centric Interaction Design, Human-AI Collaboration and Capability Enhancement, and Ethics and Accountability, with a total of thirteen sub-dimensions. CONCLUSIONS: This review provides a comprehensive, human-centered mapping of conversational AI research in mental health across the patient journey. Critical gaps remain in stage coverage, disorder diversity, population inclusivity, multimodal data integration, and interdisciplinary evaluation. The proposed taxonomy offers a structured framework to align AI development with human-centered principles, fostering empathetic, ethical, effective, and equitable mental health support.

References

  1. 1.Li, J., Li, H., Zhu, H., Walter, H., Banaschewski, T., Li, C., Li, X., & Zhang, Y. (2026). A scoping review on conversational AI in mental health: A human-centered perspective. International Journal of Medical Informatics. https://doi.org/10.1016/j.ijmedinf.2026.106460
Share:XLinkedIn

This content is for educational purposes for healthcare professionals only and does not constitute clinical advice. Clinical decisions should be based on individual patient assessment, current guidelines, and appropriate specialist consultation. Editorial Standards · Privacy Policy · Terms of Service