Research Appraisalother

Variable selection for clinical prediction models in low-dimensional data - a simulation study comparing traditional regression and machine learning methods

BMC medical research methodologyVey, Johannes A, Heinze, Georg, Kieser, Meinhard7 July 2026DOI

Clinical Snapshot

95CEBM
Evidence: Strongother

PICO Framework

P — PopulationSimulated biomedical datasets representing low-dimensional clinical data (15 variables total, 8 true predictors) with continuous outcomes, across varying sample sizes and data-generating mechanisms of increasing complexity
I — InterventionMachine learning variable selection methods: elastic net regularised regression, gradient boosting (regression and tree-based base learners), and random forest with Boruta and Hapfelmeier variable importance approaches
C — ComparatorTraditional linear regression with stepwise variable selection (LMSS)
O — OutcomesVariable selection performance (true positive rate for predictor inclusion, false positive rate for non-predictor inclusion), prediction accuracy, and model calibration across four data-generating mechanisms and multiple sample sizes

Bottom Line

This well-designed simulation study provides important methodological guidance for clinical researchers developing prediction models from low-dimensional biomedical data. Contrary to the prevailing enthusiasm for machine learning, the study demonstrates that traditional linear regression with stepwise variable selection (LMSS) achieves the best balance between correctly identifying true predictors and excluding noise variables across most realistic biomedical data scenarios. Gradient boosting and elastic net regularisation tend to include spurious variables — particularly as sample size increases — which risks overfitting and poor model transportability. Random forest with Boruta or Hapfelmeier variable importance approaches emerges as a credible alternative when non-linear or interaction effects are suspected. Critically, all methods require sufficiently large samples to perform reliably; small samples compromise both variable selection accuracy and model calibration regardless of method choice. For Australian clinical researchers, this study supports a pragmatic, evidence-based approach: use LMSS as the default for low-dimensional CPM development, reserve random forest approaches for data with suspected complex associations, and prioritise adequate sample size planning. The restriction to continuous outcomes and low-dimensional settings means findings should not be extrapolated to binary outcomes, survival models, or high-dimensional data without further evidence.

Evidence: Strong

Key Findings

  • P Value: Not applicable for simulation study design; comparative performance assessed across simulation replications

  • Effect Size: LMSS demonstrated the best trade-off between predictor inclusion and non-predictor exclusion in most scenarios. Gradient boosting (both regression and tree-based) and elastic net showed higher false positive rates, particularly at larger sample sizes. Random forest with Boruta or Hapfelmeier approaches performed comparably to LMSS in more complex data structures.

  • Primary Outcome: Variable selection performance: trade-off between true positive rate (correctly including predictor variables) and false positive rate (incorrectly including non-predictor variables) across four data-generating mechanisms and multiple sample sizes

  • Nnt Or Sensitivity: Not directly applicable; in methodological terms, LMSS showed superior specificity for variable selection (lower false positive rate) while maintaining adequate sensitivity (true positive rate) across most scenarios. Random forest with Boruta/Hapfelmeier provided a viable alternative with similar sensitivity-specificity balance in complex non-linear data structures.

  • Confidence Interval: Not reported in the abstract; Monte Carlo standard errors for simulation estimates not provided

Clinical Application

All evaluated methods (linear regression with stepwise selection, elastic net, gradient boosting, random forest with Boruta/Hapfelmeier) are implemented in widely available statistical software (R packages: glmnet, mboost, gbm, randomForest, Boruta). LMSS is the most accessible method for clinical researchers without advanced machine learning expertise. The random forest alternatives require additional familiarity with variable importance frameworks but are feasible in most academic clinical research settings. In Australia, clinical prediction model development is relevant across primary care (RACGP guidelines emphasise evidence-based risk stratification tools), hospital-based medicine, and health technology assessment for PBS/TGA submissions. The RACGP and NHMRC both emphasise the importance of well-calibrated, validated prediction tools. The finding that LMSS provides reliable variable selection in low-dimensional settings supports its continued use in Australian clinical research contexts where sample sizes are often modest (e.g., single-centre studies, rare disease registries). For researchers developing models for PBS reimbursement submissions or TGA-regulated clinical decision support tools, the calibration findings are directly relevant to regulatory requirements for model performance. The study's guidance on minimum sample size requirements is particularly pertinent for Australian researchers working with smaller datasets from regional or Indigenous health cohorts. Clinical researchers, biostatisticians, and data scientists developing clinical prediction models from low-dimensional biomedical datasets (typically 10–20 candidate variables) with continuous outcomes. Applicable to settings such as laboratory value prediction, physiological parameter modelling, and clinical scoring system development.

Abstract

PURPOSE: A wide range of methods exist for developing a clinical prediction model (CPM) and for performing variable selection. Our purpose was to develop a fair simulation study design and to investigate the properties, strengths, and weaknesses of different methods to predict a continuous outcome in low-dimensional data situations. METHODS: In this simulation study, we conducted a neutral comparison of traditional (linear regression with stepwise selection) and machine learning (regularized regression with elastic net, gradient boosting, random forest) variable selection strategies to derive a CPM. The generated datasets included a total of 15 variables, with 8 of those being predictor variables. Four data- and outcome-generating mechanisms with increasing complexity produced data structures typical for biomedicine covering linear associations and gradually introducing non-linear and non-additive elements into the data structure. RESULTS: All methods generally performed better with increasing sample size and less noise in the data. Gradient boosting with regression models and with trees as base learners, and the elastic net regularized regression included nearly all variables (i.e., both the predictor and non-predictor variables), especially with increasing sample size. The linear regression model with stepwise selection (LMSS) showed the best trade-off between correctly including the predictors and excluding the non-predictor variables in most of the scenarios, even when the functional form of continuous predictors deviated from linearity. In more complex data, variable selection using the Boruta or Hapfelmeier approach for random forest performed similar to LMSS. CONCLUSION: The sample size must be sufficiently large to enable the methods to reliably identify the predictor variables and to ensure that the developed CPMs are accurate and well-calibrated. LMSS revealed good properties and the random forest with the Boruta or Hapfelmeier approach are suitable alternatives if complex associations between predictors and outcomes are assumed.

References

  1. 1.Vey, J. A., Heinze, G., & Kieser, M. (2021). Variable selection for clinical prediction models in low-dimensional data - a simulation study comparing traditional regression and machine learning methods. BMC Medical Research Methodology. https://doi.org/10.1186/S12874-021-01374-Y
Share:XLinkedIn

This content is for educational purposes for healthcare professionals only and does not constitute clinical advice. Clinical decisions should be based on individual patient assessment, current guidelines, and appropriate specialist consultation. Editorial Standards · Privacy Policy · Terms of Service