Evidence-Based Medicine

Research Appraisals

Evidence-based critical appraisals of the latest medical research, systematically evaluated using Oxford CEBM methodology.

Showing 2 appraisals

observationalEvidence: Moderate
60CEBM

Proceedings of the National Academy of Sciences of the United States of America

Measuring disparate impact in human and machine decisions.

Empirical analyses have grown increasingly important in discrimination litigation with the greater availability of detailed data on individuals and decisions. A popular analytic strategy is to estimate disparities after adjusting for observed covariates, typically with a regression model, in hopes of ferreting out discriminatory intent. This approach, however, is ill-suited to auditing algorithms that are now commonly used to aid decisions, which typically do not include race or other legally protected factors as inputs. Motivated by legal understandings of disparate impact, we introduce an approach that aims to measure "unjustified" disparities in both human and machine decisions. Our method, which we call risk-adjusted regression, proceeds in three steps. In the first step, we combine all available information in a machine learning model to estimate the value, or inversely, the risk, of taking a certain action, such as approving a loan application or hiring a job candidate. Second, we measure disparities in decisions after adjusting for these risk estimates alone. Finally, in the third step, we assess the sensitivity of results to potential mismeasurement of risk. We demonstrate this approach on a detailed dataset of 2.2 million police stops of pedestrians in New York City, and show that traditional statistical tests of discrimination can substantially understate the magnitude of (risk-adjusted) racial disparities.

28 July 2026

Read appraisal →
otherEvidence: Strong
95CEBM

BMC medical research methodology

Variable selection for clinical prediction models in low-dimensional data - a simulation study comparing traditional regression and machine learning methods

PURPOSE: A wide range of methods exist for developing a clinical prediction model (CPM) and for performing variable selection. Our purpose was to develop a fair simulation study design and to investigate the properties, strengths, and weaknesses of different methods to predict a continuous outcome in low-dimensional data situations. METHODS: In this simulation study, we conducted a neutral comparison of traditional (linear regression with stepwise selection) and machine learning (regularized regression with elastic net, gradient boosting, random forest) variable selection strategies to derive a CPM. The generated datasets included a total of 15 variables, with 8 of those being predictor variables. Four data- and outcome-generating mechanisms with increasing complexity produced data structures typical for biomedicine covering linear associations and gradually introducing non-linear and non-additive elements into the data structure. RESULTS: All methods generally performed better with increasing sample size and less noise in the data. Gradient boosting with regression models and with trees as base learners, and the elastic net regularized regression included nearly all variables (i.e., both the predictor and non-predictor variables), especially with increasing sample size. The linear regression model with stepwise selection (LMSS) showed the best trade-off between correctly including the predictors and excluding the non-predictor variables in most of the scenarios, even when the functional form of continuous predictors deviated from linearity. In more complex data, variable selection using the Boruta or Hapfelmeier approach for random forest performed similar to LMSS. CONCLUSION: The sample size must be sufficiently large to enable the methods to reliably identify the predictor variables and to ensure that the developed CPMs are accurate and well-calibrated. LMSS revealed good properties and the random forest with the Boruta or Hapfelmeier approach are suitable alternatives if complex associations between predictors and outcomes are assumed.

9 July 2026

Read appraisal →