AI Agents Are Coming: 5-Stage Taxonomy of Language-Based AI Systems for Psychiatry, Psychotherapy, and Counseling
Clinical Snapshot
PICO Framework
| P — Population | Clinicians, researchers, and developers working in psychiatry, psychotherapy, and counselling; patients receiving or potentially receiving language-based digital mental health interventions |
| I — Intervention | A proposed 5-stage taxonomy for classifying language-based autonomous systems in mental health contexts, distinguishing technical functionality from clinical effectiveness |
| C — Comparator | Existing classification frameworks (notably autonomous driving analogues) applied to AI systems in healthcare and mental health |
| O — Outcomes | Conceptual clarity, clinical utility, and safety guidance for deploying language-based systems across levels of autonomy in mental health care; benchmarking standards for therapeutic capability evaluation |
Bottom Line
This theoretical paper from the University of Graz and Stockholm University proposes a 5-stage taxonomy for classifying language-based autonomous systems in psychiatry and psychotherapy, progressing from static knowledge tools (Level 1) to fully autonomous therapy capability (Level 5). Its most important conceptual contribution is the explicit separation of technical performance from clinical effectiveness — a distinction with direct implications for safe deployment and governance. The critique of autonomous driving analogues as inappropriate models for mental health AI is well-reasoned and timely. However, the taxonomy remains unvalidated empirically: level boundaries are conceptual rather than operationalised, the evidence synthesis methodology is opaque, and high-risk clinical populations are inadequately addressed. The novel terminology introduced requires consensus validation before adoption in regulatory or clinical governance contexts. For Australian clinicians and health administrators, the framework offers a useful starting vocabulary for TGA, RACGP, and RANZCP policy discussions, but should not be applied prescriptively until prospective validation studies are conducted. Senior clinicians should treat this as a well-argued position paper that identifies important questions rather than providing definitive answers. The call for dynamic therapeutic capability benchmarking over static knowledge testing is the paper's most actionable and clinically grounded recommendation.
Key Findings
P Value: Not reported — no statistical testing conducted
Effect Size: Not applicable — no quantitative effect size reported; this is a theoretical framework paper
Primary Outcome: A 5-stage taxonomy for language-based autonomous systems in mental health: Level 1 (Knowledge — static benchmark tasks), Level 2 (Elementary — dynamic therapeutic microskills), Level 3 (Integration — cross-module consistency, basic case conceptualisation, blended therapy under human oversight), Level 4 (Saturation — therapist-in-the-loop, minimal supervision), Level 5 (Mastery — technically capable of autonomous therapy)
Nnt Or Sensitivity: Not applicable — no clinical trial data; the key conceptual finding is that Level 4 or 5 technical performance does not automatically confer full treatment effectiveness, even with high treatment fidelity
Confidence Interval: Not reported — no empirical data presented
Clinical Application
The taxonomy is immediately applicable as a conceptual reference framework for clinical governance discussions, ethics committee reviews, and research protocol design. However, operationalisation for regulatory or procurement purposes would require further empirical validation and consensus development. Blended therapy models (Level 3) are the most immediately feasible clinical application given current technology maturity. In Australia, the TGA regulates software as a medical device (SaMD) under the Therapeutic Goods (Medical Devices) Regulations 2002, and language-based autonomous mental health systems at Levels 3–5 would likely require TGA classification and conformity assessment. The RACGP and RANZCP have not yet published specific guidelines on autonomous AI in psychotherapy, making this taxonomy potentially valuable for informing future position statements. The PBS does not currently fund AI-delivered psychotherapy, and Medicare Benefits Schedule items for psychological therapies (e.g., Better Access initiative) are restricted to registered practitioners — autonomous systems at any level would require legislative amendment for reimbursement. The Australian Digital Health Agency's National Digital Health Strategy 2023–2028 provides a relevant policy context. Headspace, Beyond Blue, and other digital mental health platforms operating in Australia would benefit from a validated version of this taxonomy for service governance. Equity considerations are particularly salient in Australia given the mental health workforce shortage in rural and remote areas, where autonomous systems might be considered for access expansion. Mental health clinicians, psychiatrists, psychologists, counsellors, and digital health developers evaluating or deploying language-based autonomous systems; health service administrators and regulators developing governance frameworks for AI-assisted mental health care
Abstract
The rapid evolution of large language models has accelerated the development of agentic artificial intelligence (AI) systems capable of pursuing autonomous goals, creating an urgent need for structural frameworks in psychiatry and psychotherapy. While existing classifications often draw parallels to autonomous driving, this paper argues that the mental health domain requires a distinct, domain-specific theoretical foundation, as the 2 domains differ fundamentally in their semantic, ideographic, and epistemological demands. Furthermore, they differ in their end goals, for which we introduce terms such as agentic guidance capability. To guide clinicians and researchers through these developments, we propose a 5-stage taxonomy for language-based AI systems that differentiates technical functionality from clinical effectiveness. The taxonomy progresses from level 1 (knowledge level), in which systems perform static benchmark tasks, to level 2 (elementary level), characterized by dynamic engagement in specific therapeutic microskills. At level 3 (integration level), systems achieve consistency across and within modules, as well as basic case-level conceptualization suitable for blended therapy under human oversight. Level 4 (saturation level) describes therapist-in-the-loop systems capable of autonomous functioning with minimal supervision, whereas level 5 (mastery level) represents AI systems that are technically capable of performing autonomous therapy. By distinguishing technical functionality from clinical effectiveness, we conclude that level 4 or level 5 performance does not automatically translate into full treatment effectiveness, even if high treatment fidelity can be achieved. We conclude by emphasizing the need to shift benchmarking from static knowledge tests to dynamic evaluations of therapeutic capabilities in order to safely navigate the transition toward autonomous care.
References
- 1.Schuster, R., Plessen, C. Y., Carlbring, P., & Walther, A. (2026). AI Agents Are Coming: 5-Stage Taxonomy of Language-Based AI Systems for Psychiatry, Psychotherapy, and Counseling. JMIR Mental Health. https://doi.org/10.2196/91746
Related Research
Journal of psychiatric research
Systematic review of machine learning and deep learning models for EEG-based detection of depression.
2 Aug 2026
Sleep medicine reviews
Artificial intelligence in sleep medicine I: Diagnosis, treatment, care, and research
2 Aug 2026
International journal of yoga therapy
Randomized Controlled Trial on the Effect of Yoga on Anxiety Severity, Quality of Life, and Biomarkers in Individuals with Anxiety Disorder
31 July 2026
This content is for educational purposes for healthcare professionals only and does not constitute clinical advice. Clinical decisions should be based on individual patient assessment, current guidelines, and appropriate specialist consultation. Editorial Standards · Privacy Policy · Terms of Service