Research Appraisalother

The backfiring effect of weak AI safety regulation

Proceedings of the National Academy of Sciences of the United States of AmericaLaufer, Benjamin, Kleinberg, Jon, Heidari, Hoda28 July 2026DOI

Clinical Snapshot

50CEBM
Evidence: Weakother

PICO Framework

P — PopulationGeneral-purpose AI technology creators and domain specialists (those who adapt AI for specific applications) operating under regulatory frameworks
I — InterventionSafety regulation imposed on general-purpose AI creators and/or domain specialists, including weak regulation targeting domain specialists alone versus stronger dual-layer regulation targeting both parties
C — ComparatorNo regulation; regulation targeting only domain specialists; regulation targeting only general-purpose AI creators
O — OutcomesSafety and performance levels of AI systems at market; equilibrium outcomes for all players (regulator, general-purpose creator, domain specialist); revenue sharing and welfare implications

Bottom Line

This PNAS theoretical paper by Laufer, Kleinberg, and Heidari presents a game-theoretic model of AI safety regulation, revealing two important insights for policymakers and health technology regulators. First, weak safety regulation applied only to domain specialists — those who adapt general-purpose AI for specific applications such as clinical decision support — can paradoxically reduce overall AI safety. This counterintuitive 'backfiring effect' arises because downstream-only regulation alters the strategic incentives of upstream general-purpose AI developers in ways that reduce their safety investment. Second, stronger, well-designed regulation applied to both general-purpose AI creators and domain specialists can function as a commitment device, producing simultaneous gains in safety and performance that benefit all parties. For Australian healthcare regulators, this finding challenges frameworks that focus regulatory burden exclusively on clinical AI deployers while leaving foundation model developers largely unregulated. The TGA's SaMD pathway and emerging AI governance proposals should consider whether dual-layer regulatory obligations — spanning both upstream model developers and downstream clinical deployers — are necessary to achieve genuine safety improvements. The study's limitations are significant: it is purely theoretical, uses a highly stylised two-player model, and lacks empirical validation. Findings should inform regulatory design thinking but not substitute for evidence from real-world regulatory experiments.

Evidence: Weak

Key Findings

  • P Value: Not applicable — theoretical model; no hypothesis testing performed

  • Effect Size: Not quantified empirically; results are mathematical propositions establishing directional effects within the model's parameter space. The backfiring effect is shown to hold across a 'large class of parameterizations' — the precise scope of this class is not specified in the abstract.

  • Primary Outcome: Weak safety regulation imposed predominantly on domain specialists can backfire, reducing overall AI safety across a large class of model parameterisations. Conversely, stronger regulation applied to both general-purpose AI creators and domain specialists can function as a commitment device, yielding simultaneous safety and performance gains that exceed outcomes achievable under no regulation or single-player regulation.

  • Nnt Or Sensitivity: Not applicable — no empirical effect size metrics. The key theoretical metric is the equilibrium safety level under each regulatory regime relative to the no-regulation baseline, demonstrating that dual-layer regulation dominates single-layer regulation in the model's parameter space.

  • Confidence Interval: Not applicable — theoretical model; no confidence intervals reported

Clinical Application

The theoretical insights are directly translatable into regulatory design principles. Practically, implementing dual-layer regulation (targeting both foundation model developers and clinical AI deployers) requires regulatory capacity, clear definitional boundaries between 'general-purpose' and 'domain-specific' AI, and coordination between technology regulators and health regulators. These are feasible but non-trivial implementation challenges. This paper has direct relevance to Australia's evolving AI regulatory landscape. The Australian Government's voluntary AI Safety Standard (2024) and the Department of Industry, Science and Resources' AI regulatory review are at a critical juncture. The paper's central warning — that weak, use-case-only regulation can reduce safety — is pertinent to proposals that focus regulatory obligations primarily on AI deployers (domain specialists) rather than foundation model developers. In the healthcare context, the TGA's Software as a Medical Device (SaMD) framework currently targets clinical AI deployers but has limited reach over upstream foundation model developers. The Therapeutic Goods Administration and the Australian Digital Health Agency should consider whether their current frameworks inadvertently replicate the 'weak domain-specialist-only' regulatory structure identified as problematic in this model. RACGP guidance on AI in general practice similarly focuses on point-of-care deployment rather than upstream model governance. The paper supports a case for coordinated regulation spanning both the AI development and deployment layers. Healthcare AI developers, clinical AI deployers (e.g., radiology AI vendors, clinical decision support system developers), hospital procurement teams, and health technology regulators considering AI safety governance frameworks

Abstract

Recent policy proposals aim to improve the safety of general-purpose AI, but there is little understanding of the efficacy of different regulatory approaches. We present a strategic model that explores interactions between safety regulation, general-purpose AI technology creators, and domain specialists-those who adapt the technology for specific applications. Our analysis examines how regulatory measures targeting different parts of the AI development chain affect the outcome of this game. Our model assumes AI technology is characterized by two key attributes: safety and performance. The regulator first sets a minimum safety requirement that applies to one or both players. The general-purpose creator then invests in the technology, establishing its initial safety and performance levels. Next, domain specialists refine the AI for their use cases, updating the safety and performance levels and taking the product to market. Resulting revenue is shared between the specialist and generalist. Our analysis reveals two insights: first, weak safety regulation imposed predominantly on domain specialists can backfire. While it might seem logical to regulate AI use cases, our analysis shows that weak regulations targeting domain specialists alone can reduce safety in a large class of parameterizations. Second, in contrast to the previous finding, we observe that stronger, well-placed regulation can mutually benefit all players. When regulators impose appropriate safety standards on both general-purpose AI creators and domain specialists, the regulation can function as a commitment device, leading to safety and performance gains, surpassing what is achievable under no regulation or regulating only one player.

References

  1. 1.Laufer, B., Kleinberg, J., & Heidari, H. (2026). The backfiring effect of weak AI safety regulation. Proceedings of the National Academy of Sciences of the United States of America. https://doi.org/10.1073/pnas.2509768123
Share:XLinkedIn

This content is for educational purposes for healthcare professionals only and does not constitute clinical advice. Clinical decisions should be based on individual patient assessment, current guidelines, and appropriate specialist consultation. Editorial Standards · Privacy Policy · Terms of Service