Sociodemographic bias in large language model clinical trial screening
Clinical Snapshot
PICO Framework
| P — Population | Phase II–III US adult randomised controlled trial (RCT) protocols (2023–2024), evaluated using physician-validated clinical vignettes representing adult patients with 33 sociodemographic identity variants |
| I — Intervention | Clinical vignettes with sociodemographic identity labels (race/ethnicity, sex, socioeconomic status, housing status, and other identity markers) applied to nine large language models (LLMs) performing clinical trial eligibility screening |
| C — Comparator | Control vignettes without sociodemographic identity labels, evaluated by the same nine LLMs under identical conditions |
| O — Outcomes | Primary: LLM-generated eligibility judgments (eligible vs ineligible) across identity variants. Secondary: LLM assessments of adherence likelihood, resource availability, and patient–researcher trust domains |
Bottom Line
This large cross-sectional vignette study evaluated nine LLMs across 58 US RCT protocols and 5.3 million eligibility assessments, finding that LLMs apply explicit inclusion/exclusion criteria consistently regardless of patient sociodemographic identity. However, when LLMs were required to make inferences about patient behaviour, resources, or trustworthiness — domains beyond explicit protocol criteria — significant disparities emerged, with homelessness producing the largest negative effect on eligibility judgments. Race and ethnicity effects were largely attenuated after adjustment for socioeconomic status, suggesting SES is the primary driver of apparent racial disparities in LLM screening. For senior clinicians and trial administrators, the key message is nuanced: LLMs may be reliable for rule-based eligibility checking but introduce meaningful bias when performing inferential assessments of social determinants. Before deploying LLM tools in trial recruitment workflows, institutions should mandate prospective bias auditing across sociodemographic variants — particularly for housing status and SES. In the Australian context, this is especially pertinent for trials enrolling Aboriginal and Torres Strait Islander peoples and rural populations, where socioeconomic disadvantage intersects with existing barriers to trial participation. Human clinical oversight of inferential screening domains remains essential.
Key Findings
P Value: Not reported in abstract.
Effect Size: Specific adjusted effect sizes not reported in abstract. Homelessness described as producing 'the largest negative eligibility shift' and 'pronounced effects' in adherence, resources, and trust domains. Quantitative estimates require full paper review.
Primary Outcome: LLM eligibility judgments were largely stable across most sociodemographic identity variants when applied to explicit inclusion/exclusion criteria. Race and ethnicity showed minimal independent effects on eligibility after adjustment for socioeconomic status. Homelessness produced the largest negative shift in eligibility judgments across all identity variants tested.
Nnt Or Sensitivity: Not applicable to this cross-sectional bias evaluation study. Relevant analogue would be the adjusted risk difference in eligibility probability for homeless vs. control identity — not available from abstract alone.
Confidence Interval: Not reported in abstract. Mixed-effects models used; interval estimates presumably available in full paper.
Clinical Application
LLM-based trial screening tools are increasingly commercially available and technically feasible. However, this study identifies a critical implementation risk: inferential bias against homeless and socioeconomically disadvantaged patients in domains beyond explicit eligibility criteria. Deployment without bias auditing is not currently supported by the evidence. Institutions should require prospective bias evaluation — including homelessness and SES-related identity variants — before clinical deployment. Human oversight for inferential screening domains (adherence prediction, resource assessment) remains essential. In Australia, clinical trial screening is governed by TGA regulatory frameworks and NHMRC ethical guidelines, with increasing interest in digital health tools for trial recruitment efficiency. The RACGP and ACCRM have emphasised equitable access to clinical trials for rural, remote, and Indigenous populations — groups with elevated rates of housing instability and socioeconomic disadvantage. This study's findings are directly relevant to Australian trial sponsors and ethics committees evaluating LLM-assisted recruitment tools. The TGA's emerging digital health guidance and the Australian Digital Health Agency's AI governance frameworks should incorporate bias auditing requirements analogous to those implied by this research. PBS and Medicare considerations are not directly applicable, but the equity implications for Aboriginal and Torres Strait Islander communities — who face compounded sociodemographic disadvantage — warrant specific attention in any Australian adaptation of LLM trial screening. Healthcare institutions, clinical research organisations, and trial sponsors considering deployment of LLM-based tools for clinical trial pre-screening or eligibility determination in adult populations. Particularly relevant for trials enrolling socially marginalised or housing-insecure populations.
Abstract
OBJECTIVE: To assess whether large language model (LLM)-based clinical trial screening judgments vary by patient sociodemographic characteristics. MATERIALS AND METHODS: We conducted a cross-sectional evaluation of Phase II-III US adult randomized controlled trial (RCT) protocols (2023-2024). Physician-validated clinical vignettes were evaluated in a control version and 33 sociodemographic identity variants differing only by labels. Nine LLMs assessed eligibility and related domains. Mixed-effects models estimated adjusted differences vs control. RESULTS: Across 58 protocols and 5.3 million evaluations, eligibility judgments were largely stable across identities. Race and ethnicity showed minimal effects after accounting for socioeconomic status. Homelessness produced the largest negative eligibility shift and pronounced effects in adherence, resources, and trust. DISCUSSION AND CONCLUSION: LLMs applied explicit eligibility criteria consistently, but disparities emerged in domains requiring inference about behavior or resources, underscoring the need for careful deployment to promote fair trial access.
References
- 1.Soffer, S., Omar, M., Efros, O., Apakama, D. U., Mudrik, A., Freeman, R., Nadkarni, G. N., & Klang, E. (2026). Sociodemographic bias in large language model clinical trial screening. Journal of the American Medical Informatics Association: JAMIA. https://doi.org/10.1038/s41746-025-01576-4
Related Research
Minimally invasive therapy & allied technologies : MITAT : official journal of the Society for Minimally Invasive Therapy
Ventral hernia repair in emergency settings. A machine learning model to predict post-operative complications.
3 Aug 2026
Acta odontologica Scandinavica
Predicting gingival embrasure risk after invisible orthodontics using multimodal data and machine learning
23 July 2026
International ophthalmology
The relationship between diabetic retinopathy and intestinal microbiota: a systematic review and meta-analysis
10 July 2026
This content is for educational purposes for healthcare professionals only and does not constitute clinical advice. Clinical decisions should be based on individual patient assessment, current guidelines, and appropriate specialist consultation. Editorial Standards · Privacy Policy · Terms of Service