Research Appraisals
Evidence-based critical appraisals of the latest medical research, systematically evaluated using Oxford CEBM methodology.
Showing 5 appraisals
Proceedings of the National Academy of Sciences of the United States of America
The backfiring effect of weak AI safety regulation
Recent policy proposals aim to improve the safety of general-purpose AI, but there is little understanding of the efficacy of different regulatory approaches. We present a strategic model that explores interactions between safety regulation, general-purpose AI technology creators, and domain specialists-those who adapt the technology for specific applications. Our analysis examines how regulatory measures targeting different parts of the AI development chain affect the outcome of this game. Our model assumes AI technology is characterized by two key attributes: safety and performance. The regulator first sets a minimum safety requirement that applies to one or both players. The general-purpose creator then invests in the technology, establishing its initial safety and performance levels. Next, domain specialists refine the AI for their use cases, updating the safety and performance levels and taking the product to market. Resulting revenue is shared between the specialist and generalist. Our analysis reveals two insights: first, weak safety regulation imposed predominantly on domain specialists can backfire. While it might seem logical to regulate AI use cases, our analysis shows that weak regulations targeting domain specialists alone can reduce safety in a large class of parameterizations. Second, in contrast to the previous finding, we observe that stronger, well-placed regulation can mutually benefit all players. When regulators impose appropriate safety standards on both general-purpose AI creators and domain specialists, the regulation can function as a commitment device, leading to safety and performance gains, surpassing what is achievable under no regulation or regulating only one player.
28 July 2026
Read appraisal →Journal of medical Internet research
Addressing the Key Challenges of Decentralized Clinical Trials in Europe: Multistakeholder Perspective Delphi Study
BACKGROUND: Decentralized clinical trials (DCTs) represent an emerging model in clinical research, accelerated by the restrictions imposed during the COVID-19 pandemic. By leveraging digital technologies and local health care resources, DCTs aim to increase accessibility and reduce participant burden compared to traditional site-based models, which often face recruitment failures and high attrition rates. While various regulatory initiatives in Europe, such as the Accelerating Clinical Trials in the European Union program and the European Medicines Regulatory Network recommendation paper (updated in October 2025), have sought to facilitate their implementation, the widespread adoption of DCTs remains limited due to significant operational, regulatory, and technological challenges, including platform fragmentation and gaps in digital literacy. OBJECTIVE: This study aimed to identify and prioritize actionable solutions to the main challenges of DCT implementation in Europe from a multistakeholder perspective, gathering insights to address specific ethical, legal, and operational barriers. METHODS: Building on a preceding strengths, weaknesses, opportunities, and threats analysis, a 2-round Delphi study was conducted, involving 26 experts in clinical trials, ethics, law, regulation, and patient engagement between March and May 2023. In the first round, 309 open-ended responses were collected via REDCap (Research Electronic Data Capture; Vanderbilt University) surveys and underwent systematic inductive content analysis using ATLAS.ti (ATLAS.ti Scientific Software Development GmbH) with independent double coding. This process resulted in 244 unique proposals that were categorized according to 6 key challenges. In the second round, 39 synthesized proposals were evaluated using a 4-point Likert scale. Consensus was defined as ≥80% agreement on the appropriateness of each proposal. RESULTS: High levels of consensus were achieved, with 32 out of the 39 proposals reaching the threshold and 14 achieving 100% unanimity. Overall, 82% of the proposals were rated as "appropriate" or "very appropriate." Key recommendations included providing support and training for health care professionals, enhancing investigational medicinal product and biological sample logistics through validated technologies, improving collaboration with local health care providers, fostering regulatory harmonization while respecting national specificities, strengthening capacity-building initiatives, and promoting accessible, user-friendly digital tools supported by hybrid trial models. Conversely, proposals such as peer-to-peer participant support and the centralization of ethics reviews at the European Union level failed to reach consensus. CONCLUSIONS: The study offers a prioritized compilation of expert-driven recommendations for overcoming current barriers to DCT implementation in Europe. The adoption of these recommendations could support the development of more inclusive, efficient, and sustainable decentralized research frameworks across diverse health care systems.
12 July 2026
Read appraisal →Journal of medical Internet research
Personal Health Large Language Models and the Negotiation of Medical Authority in Clinical Care: Opportunities, Risks, and Governance
Personal health large language models (PH-LLMs) are patient-facing conversational systems that synthesize user-entered information, patient-generated health data, wearable data, and selected personal health records-where users choose to connect them-into personalized, longitudinal, action-oriented health narratives. Unlike generic health chatbots that mainly provide one-off responses to isolated questions, PH-LLMs may generate continuing interpretations, priorities, and candidate next steps that patients bring into clinical encounters. In contrast to electronic health record-tethered clinical artificial intelligence (AI), they often originate outside institutional oversight and may be selected, used, or trusted by patients before professional review. This viewpoint examines how PH-LLMs may reshape the negotiation of medical authority by contributing to a shift from the traditional dyadic clinician-patient relationship toward a triadic model of negotiated authority, in which clinicians may increasingly need to mediate among clinical evidence, patient values, and algorithmic narratives. PH-LLMs may support patient participation by organizing symptoms, contextualizing home-monitoring and wearable data, improving health literacy, assisting chronic disease self-management, and preparing patients for more collaborative visits. Patient-facing AI narratives may also introduce distinct risks. At the individual level, these include inaccurate or incomplete responses arising from imprecise queries or missing context, misinterpretation of otherwise accurate information in the absence of clinical context, and contextually biased or poorly matched advice across demographic, cultural, linguistic, disability-related, or socioeconomic contexts. At the system level, they include authority conflict when AI recommendations diverge from clinical judgment, fragmentation of clinical truth, privacy and data-governance concerns, diffusion of accountability when harm results from advice produced outside clinical governance, and inequitable access to premium tools and continuous monitoring devices. To address these challenges, we propose a 3-layer clinical governance framework for patient-brought PH-LLM narratives. The first layer, evidence and provenance, makes AI-generated narratives epistemically legible by clarifying platform identity, data sources, temporal anchoring, uncertainty, and privacy-relevant data-use and retention conditions. The second layer, clinical arbitration and workflow integration, uses risk-stratified intake, proportionate documentation, escalation triggers, and equity-preserving workflows to embed PH-LLM outputs into routine care. The third layer, competence and accountability, defines the communication competencies, AI literacy supports, institutional responsibilities, vendor accountability, and risk-proportionate verification duties needed for triadic care. This framework is a conceptual and governance-oriented proposal rather than a validated clinical protocol. Future empirical work should evaluate its feasibility, documentation burden, equity effects, clinical safety impact, and acceptability among patients, clinicians, and health systems. Governed through these interdependent layers, PH-LLMs may serve as supporting infrastructure for safer, person-centered longitudinal care.
26 June 2026
Read appraisal →Current psychiatry reports
Artificial Intelligence in Child and Adolescent Psychiatry: A Narrative Review of Recent Clinical Applications and Ethical Considerations
PURPOSE OF REVIEW: This narrative review examines recent literature of artificial intelligence (AI) in child and adolescent psychiatry. With increasing mental health disorders in young people along with persistent workforce shortages, AI has emerged as a potential tool to improve efficiency and support clinical decision making. However, using AI raises important ethical concerns which are summarized in our review. RECENT FINDINGS: Most AI applications in child and adolescent psychiatry remain early in development and are not ready for routine clinical use. AI scribes are likely impractical in many child psychiatry settings because of multiparty visits, consent concerns, and sensitive clinical discussions. Many multimodal diagnostic tools using machine learning still require further testing and can be impractical when relying on costly diagnostics like neuroimaging. Similarly, AI-assisted therapeutics requiring physical hardware like robotics and virtual or augmented reality devices can also be prohibitively expensive. One of the few exceptions includes AI-enabled video and eye-tracking approaches for autism diagnosis. Chatbots and robot companions may provide modest improvements for depression, but evidence remains limited in the pediatric population with risk of serious harm. Concerns of AI include misinformation, algorithmic biases, privacy risks, crisis mismanagement, and excessive emotional attachments to chatbots. AI may eventually support child and adolescent psychiatry, but current evidence supports cautious, supervised use rather than broad clinical adoption. Clinicians should help families understand AI's limits, encourage digital literacy, and ensure that AI remains an adjunct to human care rather than a substitute.
22 June 2026
Read appraisal →Current psychiatry reports
The Digital Mirror: Clinical Potentials and Relational Risks of Generative AI in Mental Health Interventions
PURPOSE OF REVIEW: This review explores the rapidly evolving integration of Generative Artificial Intelligence (GenAI) in mental health care. It aims to evaluate current applications in assessment, treatment planning, and psychotherapeutic interventions, while critically examining the clinical risks, ethical dilemmas, and the future potential of GenAI as an adjunctive tool rather than a replacement for human-delivered therapy. RECENT FINDINGS: Recent studies indicate that AI models can effectively assist in diagnostic reasoning, biomarker identification via EEG, and the prediction of symptom trajectories from session transcripts. Randomized controlled trials (RCTs) suggest that GenAI chatbots significantly reduce anxiety and depressive symptoms in the short term, particularly in settings with limited access to clinicians. However, human-led therapy remains superior in fostering deep emotional engagement and clinical impact. Significant risks identified include the potential for GenAI to foster dependency, reinforce maladaptive schemas or delusional ideation through "sycophantic" mirroring, and raise complex ethical-legal challenges regarding the reporting of criminal disclosures. AI represents a transformative adjunctive layer in mental health, offering scalable support for assessment, training, and between-session monitoring. While technological advances in personalization, multimodality, and immersive virtual reality enhance its clinical utility, GenAI lacks the authentic relational depth and"calibrated mismatches" essential for autonomy and transformative change. Future integration must prioritize a human-centered, blended approach, where GenAI is strictly supervised by clinicians within a robust ethical and regulatory framework to preserve the essential heart of the therapeutic connection. Research priorities, interim clinical safeguards, and recommendations for navigating the gap between current evidence and real-world adoption need to be defined and implemented.
19 June 2026
Read appraisal →