Research Appraisals
Evidence-based critical appraisals of the latest medical research, systematically evaluated using Oxford CEBM methodology.
Showing 9 appraisals
Rheumatic diseases clinics of North America
Artificial Intelligence Regulation in the United States: Current Landscape and Implications for Rheumatology
Artificial intelligence (AI) is increasingly embedded in clinical tools used in rheumatology, including imaging interpretation, longitudinal disease monitoring, and electronic health record-based decision support. AI has moved from the periphery of biomedical research to an operational component of clinical care, increasingly embedded in electronic health records, imaging platforms, and decision support systems. In rheumatology, where care is longitudinal, AI systems offer substantial promise-but also poses distinct risks. Together, these commitments-advancing humanity, ensuring equity, engaging impacted individuals, improving workforce well-being, monitoring performance, innovating and learning, and promoting sustainability-operationalize trustworthy AI systems across a patient's care trajectory.
2 Aug 2026
Read appraisal →Proceedings of the National Academy of Sciences of the United States of America
The backfiring effect of weak AI safety regulation
Recent policy proposals aim to improve the safety of general-purpose AI, but there is little understanding of the efficacy of different regulatory approaches. We present a strategic model that explores interactions between safety regulation, general-purpose AI technology creators, and domain specialists-those who adapt the technology for specific applications. Our analysis examines how regulatory measures targeting different parts of the AI development chain affect the outcome of this game. Our model assumes AI technology is characterized by two key attributes: safety and performance. The regulator first sets a minimum safety requirement that applies to one or both players. The general-purpose creator then invests in the technology, establishing its initial safety and performance levels. Next, domain specialists refine the AI for their use cases, updating the safety and performance levels and taking the product to market. Resulting revenue is shared between the specialist and generalist. Our analysis reveals two insights: first, weak safety regulation imposed predominantly on domain specialists can backfire. While it might seem logical to regulate AI use cases, our analysis shows that weak regulations targeting domain specialists alone can reduce safety in a large class of parameterizations. Second, in contrast to the previous finding, we observe that stronger, well-placed regulation can mutually benefit all players. When regulators impose appropriate safety standards on both general-purpose AI creators and domain specialists, the regulation can function as a commitment device, leading to safety and performance gains, surpassing what is achievable under no regulation or regulating only one player.
28 July 2026
Read appraisal →JMIR mental health
AI Agents Are Coming: 5-Stage Taxonomy of Language-Based AI Systems for Psychiatry, Psychotherapy, and Counseling
The rapid evolution of large language models has accelerated the development of agentic artificial intelligence (AI) systems capable of pursuing autonomous goals, creating an urgent need for structural frameworks in psychiatry and psychotherapy. While existing classifications often draw parallels to autonomous driving, this paper argues that the mental health domain requires a distinct, domain-specific theoretical foundation, as the 2 domains differ fundamentally in their semantic, ideographic, and epistemological demands. Furthermore, they differ in their end goals, for which we introduce terms such as agentic guidance capability. To guide clinicians and researchers through these developments, we propose a 5-stage taxonomy for language-based AI systems that differentiates technical functionality from clinical effectiveness. The taxonomy progresses from level 1 (knowledge level), in which systems perform static benchmark tasks, to level 2 (elementary level), characterized by dynamic engagement in specific therapeutic microskills. At level 3 (integration level), systems achieve consistency across and within modules, as well as basic case-level conceptualization suitable for blended therapy under human oversight. Level 4 (saturation level) describes therapist-in-the-loop systems capable of autonomous functioning with minimal supervision, whereas level 5 (mastery level) represents AI systems that are technically capable of performing autonomous therapy. By distinguishing technical functionality from clinical effectiveness, we conclude that level 4 or level 5 performance does not automatically translate into full treatment effectiveness, even if high treatment fidelity can be achieved. We conclude by emphasizing the need to shift benchmarking from static knowledge tests to dynamic evaluations of therapeutic capabilities in order to safely navigate the transition toward autonomous care.
15 July 2026
Read appraisal →Journal of medical Internet research
Addressing the Key Challenges of Decentralized Clinical Trials in Europe: Multistakeholder Perspective Delphi Study
BACKGROUND: Decentralized clinical trials (DCTs) represent an emerging model in clinical research, accelerated by the restrictions imposed during the COVID-19 pandemic. By leveraging digital technologies and local health care resources, DCTs aim to increase accessibility and reduce participant burden compared to traditional site-based models, which often face recruitment failures and high attrition rates. While various regulatory initiatives in Europe, such as the Accelerating Clinical Trials in the European Union program and the European Medicines Regulatory Network recommendation paper (updated in October 2025), have sought to facilitate their implementation, the widespread adoption of DCTs remains limited due to significant operational, regulatory, and technological challenges, including platform fragmentation and gaps in digital literacy. OBJECTIVE: This study aimed to identify and prioritize actionable solutions to the main challenges of DCT implementation in Europe from a multistakeholder perspective, gathering insights to address specific ethical, legal, and operational barriers. METHODS: Building on a preceding strengths, weaknesses, opportunities, and threats analysis, a 2-round Delphi study was conducted, involving 26 experts in clinical trials, ethics, law, regulation, and patient engagement between March and May 2023. In the first round, 309 open-ended responses were collected via REDCap (Research Electronic Data Capture; Vanderbilt University) surveys and underwent systematic inductive content analysis using ATLAS.ti (ATLAS.ti Scientific Software Development GmbH) with independent double coding. This process resulted in 244 unique proposals that were categorized according to 6 key challenges. In the second round, 39 synthesized proposals were evaluated using a 4-point Likert scale. Consensus was defined as ≥80% agreement on the appropriateness of each proposal. RESULTS: High levels of consensus were achieved, with 32 out of the 39 proposals reaching the threshold and 14 achieving 100% unanimity. Overall, 82% of the proposals were rated as "appropriate" or "very appropriate." Key recommendations included providing support and training for health care professionals, enhancing investigational medicinal product and biological sample logistics through validated technologies, improving collaboration with local health care providers, fostering regulatory harmonization while respecting national specificities, strengthening capacity-building initiatives, and promoting accessible, user-friendly digital tools supported by hybrid trial models. Conversely, proposals such as peer-to-peer participant support and the centralization of ethics reviews at the European Union level failed to reach consensus. CONCLUSIONS: The study offers a prioritized compilation of expert-driven recommendations for overcoming current barriers to DCT implementation in Europe. The adoption of these recommendations could support the development of more inclusive, efficient, and sustainable decentralized research frameworks across diverse health care systems.
12 July 2026
Read appraisal →Journal of medical Internet research
How Does That Large Language Model Make You Feel?
People are increasingly turning to commercially available large language models (LLMs) for emotional support. In this News and Perspectives article, JMIR Correspondent Simon Spichak reports on the role of LLMs in mental health, speaking with experts about safety concerns, research gaps, and next steps.
2 July 2026
Read appraisal →Joint Commission journal on quality and patient safety
Large Language Models as Physician Recommenders: Limitations, Alternatives, and the Path Toward Hybrid AI Systems
Digital tools are increasingly used by patients to access health information and navigate care, including choosing clinicians, but the evidence supporting "best doctor" recommendations varies widely across available methods. This manuscript compares the emerging use of consumer-facing large language models (LLMs) for this purpose with four alternative physician-selection strategies: patient networks, expert referrals, consumer rating/search platforms, and outcome-informed artificial intelligence / machine learning (AI/ML) recommender systems. Consumer rating platforms mainly reflect patient experiences rather than technical skills and often show weak or inconsistent links with objective performance measures. General-purpose LLMs may produce convincing rationales that highlight visibility and reputation, but they remain limited by hallucinations, prompt sensitivity, and uncertainty about whether their explanations truly reflect the underlying reasoning. Outcome-informed AI/ML recommender systems could, in theory, incorporate risk-adjusted clinical performance, but they face significant challenges, including unreliable physician-level measurement, incomplete data, residual confounding, and limited transparency, which can undermine trust and informed consent. We suggest that a hybrid system should be seen as a guiding concept and a testable approach rather than a final solution. Such a system could integrate proven, auditable performance measures with transparent patient-preference filters and verifiable, patient-facing explanations. We present key design principles, including clinical relevance, explainability, autonomy, equity, and governance, in line with current US Food and Drug Administration (FDA) guidance related to clinical decision support and adaptive AI life cycle management.
1 July 2026
Read appraisal →Journal of medical Internet research
A Futures Framework for Clinical AI Governance: Anticipating Emerging Risks, Shifting Roles, and Regulatory Challenges
This viewpoint develops the futures framework for clinical artificial intelligence governance (FF-CAIG), a conceptual and anticipatory framework for organizing emerging governance challenges in clinical artificial intelligence (AI). Although life cycle-oriented oversight is increasingly reflected in clinical AI regulation and institutional governance, existing approaches remain more developed for near-term validation, current-state assurance, and retrospective risk detection than for longer-horizon sociotechnical change. This gap is increasingly relevant, as AI systems become more complex, adaptive, and autonomous, and as they become more deeply embedded in care relationships and accountability structures. FF-CAIG is grounded in 3 futures methodologies: the 3 horizons model, scenario planning, and causal layered analysis. It is operationalized through an emerging clinical AI risk taxonomy that links these methods to governance domains. Its practical outputs include horizon classification, risk-domain mapping, scenario stress-testing findings, accountability-chain mapping, and horizon-scaled minimum governance actions for deployment or continued use. Applied across near-term, transitional, and longer-term horizons, the framework proposes cross-horizon priorities, including stronger predeployment equity evaluation, clearer life cycle accountability, clinician AI oversight competencies, and safeguards for increasingly autonomous or AI-mediated care systems. We illustrate FF-CAIG through 3 representative clinical AI deployment patterns and discuss its limitations, including differential compliance burdens, risks of overdocumentation, variable feasibility across jurisdictions, and the need for empirical validation. FF-CAIG is intended not as a prescriptive policy instrument or validated assessment tool, but as a structured analytic approach for regulators, health system leaders, developers, and researchers seeking prospective and systems-oriented approaches to clinical AI governance.
30 June 2026
Read appraisal →The Behavioral and brain sciences
Fair contracts for artificial intelligence?
This comment explores resource-rational contractualism (RRC) as a framework for artificial intelligence (AI) alignment and reasoning. It assesses RRC's capacity to capture the insights of philosophical contractualism and canvasses its extension to AI alignment. It suggests that a modified version of RRC, centered on impartial deliberation, can be used to identify principles for alignment and to guide efficient AI moral cognition.
27 June 2026
Read appraisal →Journal of medical Internet research
Ethical Considerations in Personal Health Large Language Models
Personal health large language models (PH-LLMs) have rapidly evolved from research prototypes into consumer-facing, data-linked systems that support symptom triage, medication questions, mental health check-ins, and longitudinal self-management. Their direct-to-consumer use without clinical oversight creates a distinct ethical risk profile that general artificial intelligence governance frameworks do not fully address. This viewpoint focuses on text-based, platform-mediated PH-LLMs and synthesizes PH-LLM-specific challenges across 6 domains: privacy, accuracy, equity, transparency, human-artificial intelligence interaction, and regulatory governance. These risks may be amplified by health literacy gaps, longitudinal data aggregation, persuasive conversational design, and fragmented oversight across the consumer-clinical boundary. Grounded in the 4 principles of biomedical ethics, we propose a governance framework that operationalizes beneficence, nonmaleficence, autonomy, and justice through design and deployment controls, including health literacy-aligned communication, crisis and pharmacological safeguards, hallucination mitigation, role disclosure, granular consent, fairness auditing, and accessible design. We further outline implementation mechanisms, including risk-tiered certification, tiered accountability, and postdeployment oversight through adverse-event reporting, transparency reporting, and independent safety evaluation. This framework is intended as an evidence-informed but partly anticipatory approach to governing PH-LLMs in personal health management.
19 June 2026
Read appraisal →