Sources of Bias in Clinical Artificial Intelligence and Applications in Rheumatology
Clinical Snapshot
PICO Framework
| P — Population | Patients with rheumatic diseases and the clinical AI/machine-learning models developed to support their care |
| I — Intervention | Clinical artificial intelligence and machine-learning models applied across the rheumatology model lifecycle (design, training, deployment, monitoring) |
| C — Comparator | No formal comparator; narrative review comparing biased versus bias-mitigated model development approaches and frameworks |
| O — Outcomes | Identification and characterisation of sources of bias (preexisting, technical, emergent) in rheumatology AI models; subgroup performance failures; implications for equitable care delivery |
Bottom Line
This narrative review from UCSF rheumatologists provides a clinically useful taxonomy of bias in rheumatology AI, distinguishing preexisting bias (embedded in historical data), technical bias (from design and optimisation choices), and emergent bias (arising during real-world deployment). The central argument — that aggregate model performance metrics routinely obscure subgroup failures — is well-reasoned and directly relevant to clinical governance. The review is timely given accelerating AI deployment in rheumatology, where disease heterogeneity and existing care disparities create particular vulnerability to biased model outputs. However, the absence of a systematic search methodology, quantitative synthesis, or formal validation of the proposed framework limits its evidentiary weight. Clinicians and health system leaders should treat this as expert conceptual guidance rather than evidence-based practice recommendations. Its greatest practical value lies in prompting structured pre-deployment evaluation of AI tools, with explicit attention to subgroup performance across sex, ethnicity, socioeconomic status, and disease phenotype. For Australian practice, the framework aligns with TGA SaMD regulatory expectations and is particularly pertinent given documented disparities in rheumatic disease care among Indigenous Australians and culturally diverse populations.
Key Findings
Effect Size: Not applicable — no primary quantitative data reported
Primary Outcome: Conceptual characterisation of three categories of bias in rheumatology AI: preexisting bias (reflecting historical inequities in data), technical bias (arising from design and optimisation choices), and emergent bias (arising from real-world deployment and workflow interactions)
Nnt Or Sensitivity: Not applicable — narrative review; no diagnostic, therapeutic, or prognostic effect estimates generated
Clinical Application
The conceptual framework is immediately applicable as a structured checklist for evaluating AI tools prior to clinical deployment. Operationalising the recommendations (e.g., subgroup performance audits, transparency requirements, sustained post-deployment monitoring) requires institutional infrastructure and dedicated health informatics resources that may not be uniformly available across practice settings Directly relevant to Australian rheumatology practice. The Australian Commission on Safety and Quality in Health Care and the TGA's emerging regulatory framework for Software as a Medical Device (SaMD) both require consideration of equity and subgroup performance. The RACGP and Australian Rheumatology Association should consider this framework when developing guidance on AI tool adoption. Indigenous Australians experience disproportionate burden from rheumatic heart disease and inflammatory arthritis; AI models trained predominantly on non-Indigenous datasets risk perpetuating these disparities. PBS-funded biologics and targeted synthetic DMARDs involve complex prescribing decisions where biased AI support tools could exacerbate existing access inequities. The Australian Digital Health Agency's national digital health strategy provides a policy context for implementing the transparency and oversight principles advocated in this review. Rheumatologists, clinical informaticists, health system leaders, and quality improvement teams involved in evaluating, procuring, or deploying AI-assisted clinical decision support tools for patients with rheumatic diseases including rheumatoid arthritis, systemic lupus erythematosus, vasculitis, and other heterogeneous inflammatory conditions
Abstract
Rheumatology machine-learning models are limited by preexisting, technical, and emergent biases; the interaction of data constraints, design choices, and real-world clinical workflows, rather than from isolated technical errors. Across the model lifecycle, optimization objectives can encode patterns of care, access, and documentation, producing hidden subgroup failures that are obscured by aggregate performance metrics. Given the heterogeneity of rheumatic disease and current disparities in care delivery, addressing bias requires deliberate design choices before, during, and after a model is built, as well as a commitment to transparency, and sustained oversight.
References
- 1.Creasman, M., Garcia-Agundez, A., Yazdany, J., & Schmajuk, G. (2026). Sources of bias in clinical artificial intelligence and applications in rheumatology. Rheumatic Diseases Clinics of North America. Advance online publication. https://doi.org/10.1016/j.rdc.2026.03.005
Related Research
The Journal of rheumatology
Examining the Role of Wearables in Inflammatory Arthritis Care: A Narrative Literature Review
3 Aug 2026
Rheumatic diseases clinics of North America
Demystifying Artificial Intelligence: Key Concepts with Examples in Rheumatology
2 Aug 2026
Rheumatic diseases clinics of North America
Machine Learning-Enhanced Autoantibody Discovery and Diagnostics in Systemic Autoimmune Rheumatic Diseases
2 Aug 2026
This content is for educational purposes for healthcare professionals only and does not constitute clinical advice. Clinical decisions should be based on individual patient assessment, current guidelines, and appropriate specialist consultation. Editorial Standards · Privacy Policy · Terms of Service