AI agents are sensitive to nudges.
Clinical Snapshot
PICO Framework
| P — Population | Large language models (LLMs) deployed as autonomous decision-making agents, benchmarked against human decision-makers |
| I — Intervention | Exposure to four forms of choice architecture (nudges): defaults, suggestions, information highlighting, and 'optimal' nudges derived from a resource-rational model of human choice |
| C — Comparator | Human decision-making behaviour under identical choice architecture conditions (used as a normative baseline for calibrated sensitivity to nudges) |
| O — Outcomes | Degree of behavioural sensitivity to nudges (responsiveness to choice architecture manipulations), payoff outcomes, information acquisition costs, and stability of decision behaviour across prompting strategies |
Bottom Line
This experimental study from MIT and Dartmouth demonstrates that large language model agents are substantially more sensitive to subtle changes in choice architecture — including defaults, suggestions, and information highlighting — than human decision-makers under equivalent conditions. This 'behavioural brittleness' occurs even in non-adversarial settings, meaning that minor, unintentional variations in how options are presented to an LLM agent can produce disproportionately large shifts in its decisions, toward both better and worse outcomes. Chain-of-thought prompting and providing in-context human data do not reliably stabilise this behaviour. Reasoning-optimised models offer partial mitigation but are inconsistent and computationally expensive. For senior clinicians and health system leaders, the critical implication is this: LLM agents deployed in clinical workflows — whether for triage, prescribing support, referral management, or patient communication — may behave unpredictably when interface design, prompt wording, or information presentation varies, even subtly. This is a patient safety concern that precedes adversarial manipulation. Robust pre-deployment testing across varied choice architectures, ongoing monitoring, and governance frameworks that account for nudge sensitivity are urgently needed before broad clinical deployment of autonomous LLM agents.
Key Findings
P Value: Not reported in abstract; requires full-text review
Effect Size: Not quantitatively specified in abstract; described qualitatively as LLMs being 'far more responsive' than humans, with weak cues producing larger behavioural shifts in LLMs than in humans
Primary Outcome: LLMs are substantially more sensitive to choice architecture manipulations (nudges) than human decision-makers across all four nudge types tested (defaults, suggestions, information highlighting, and resource-rational optimal nudges)
Nnt Or Sensitivity: Not applicable in traditional sense; directional finding: reasoning-optimised LLMs partially restore human-level nudge sensitivity in some configurations but do so inconsistently and at high computational cost
Confidence Interval: Not reported in abstract; requires full-text review
Clinical Application
Findings are immediately relevant to system designers and clinical informaticists deploying LLM agents. Mitigation strategies (reasoning-optimised models, careful prompt engineering) are technically feasible but computationally costly and inconsistently effective. No pharmacological or procedural intervention is required; implications are for governance, interface design, and safety testing protocols. Directly relevant to the Australian Digital Health Agency's national digital health strategy and the increasing deployment of LLM-based tools in My Health Record integrations, telehealth platforms, and primary care decision support. The TGA's emerging regulatory framework for Software as a Medical Device (SaMD) does not yet comprehensively address behavioural brittleness of LLM agents under choice architecture variation. RACGP guidelines on clinical decision support tools do not currently specify nudge-sensitivity testing requirements. The Australian Commission on Safety and Quality in Health Care (ACSQHC) should consider this evidence when developing standards for AI-assisted clinical decision-making. PBS and MBS implications are indirect but significant if LLM agents influence prescribing defaults or referral pathways. Any clinical or healthcare administrative setting deploying LLM-based autonomous agents for decision support, triage, care coordination, patient communication, or clinical documentation — particularly where interface design, default settings, or information presentation may vary
Abstract
Large language models (LLMs) are increasingly deployed as autonomous agents that make choices and use tools on behalf of users. Yet, we have limited evidence about how their decisions are shaped by their environment. We adapt a human decision-making task to test leading LLMs under four forms of choice architecture: defaults, suggestions, information highlighting, and "optimal" nudges derived from a resource-rational model of human choice. We treat human behavior as a baseline for predictable sensitivity to such interventions. Across models and prompting strategies, LLMs often depart substantially from this baseline. They sometimes pay excessive costs to acquire information, sometimes ignore available information, and, most crucially, are far more responsive to nudges than humans, such that weak cues that slightly shift human behavior have larger effects on model choices, toward both better and worse payoff outcomes. Chain-of-thought prompting and in-context human data do not reliably stabilize behavior. Recent reasoning-optimized LLMs can, in some configurations, restore more human-level sensitivity to nudges, but do so inconsistently and at substantial computational cost. These results point to an important and largely neglected safety concern: LLM agents can be behaviorally brittle under subtle changes in choice architecture, even in the absence of adversarial settings.
References
- 1.Cherep, M., Maes, P., & Singh, N. (2026). AI agents are sensitive to nudges. Proceedings of the National Academy of Sciences of the United States of America. https://doi.org/10.1073/pnas.2537030123
Related Research
Clinical gastroenterology and hepatology : the official clinical practice journal of the American Gastroenterological Association
Artificial Intelligence Tools for Gastrointestinal Research: A Practical Guide
2 Aug 2026
Clinical spine surgery
An Introduction to Machine Learning for the Practicing Spine Surgeon
2 Aug 2026
Rheumatic diseases clinics of North America
Sources of Bias in Clinical Artificial Intelligence and Applications in Rheumatology
2 Aug 2026
This content is for educational purposes for healthcare professionals only and does not constitute clinical advice. Clinical decisions should be based on individual patient assessment, current guidelines, and appropriate specialist consultation. Editorial Standards · Privacy Policy · Terms of Service