Variable selection for clinical prediction models in low-dimensional data - a simulation study comparing traditional regression and machine learning methods
Clinical Snapshot
PICO Framework
| P — Population | Simulated biomedical datasets representing low-dimensional clinical data (15 variables total, 8 true predictors) with continuous outcomes, across varying sample sizes and data-generating mechanisms of increasing complexity |
| I — Intervention | Machine learning variable selection methods: elastic net regularised regression, gradient boosting (regression and tree-based base learners), and random forest with Boruta and Hapfelmeier variable importance approaches |
| C — Comparator | Traditional linear regression with stepwise variable selection (LMSS) |
| O — Outcomes | Variable selection performance (true positive rate for predictor inclusion, false positive rate for non-predictor inclusion), prediction accuracy, and model calibration across four data-generating mechanisms and multiple sample sizes |
Bottom Line
This well-designed simulation study provides important methodological guidance for clinical researchers developing prediction models from low-dimensional biomedical data. Contrary to the prevailing enthusiasm for machine learning, the study demonstrates that traditional linear regression with stepwise variable selection (LMSS) achieves the best balance between correctly identifying true predictors and excluding noise variables across most realistic biomedical data scenarios. Gradient boosting and elastic net regularisation tend to include spurious variables — particularly as sample size increases — which risks overfitting and poor model transportability. Random forest with Boruta or Hapfelmeier variable importance approaches emerges as a credible alternative when non-linear or interaction effects are suspected. Critically, all methods require sufficiently large samples to perform reliably; small samples compromise both variable selection accuracy and model calibration regardless of method choice. For Australian clinical researchers, this study supports a pragmatic, evidence-based approach: use LMSS as the default for low-dimensional CPM development, reserve random forest approaches for data with suspected complex associations, and prioritise adequate sample size planning. The restriction to continuous outcomes and low-dimensional settings means findings should not be extrapolated to binary outcomes, survival models, or high-dimensional data without further evidence.
Key Findings
P Value: Not applicable for simulation study design; comparative performance assessed across simulation replications
Effect Size: LMSS demonstrated the best trade-off between predictor inclusion and non-predictor exclusion in most scenarios. Gradient boosting (both regression and tree-based) and elastic net showed higher false positive rates, particularly at larger sample sizes. Random forest with Boruta or Hapfelmeier approaches performed comparably to LMSS in more complex data structures.
Primary Outcome: Variable selection performance: trade-off between true positive rate (correctly including predictor variables) and false positive rate (incorrectly including non-predictor variables) across four data-generating mechanisms and multiple sample sizes
Nnt Or Sensitivity: Not directly applicable; in methodological terms, LMSS showed superior specificity for variable selection (lower false positive rate) while maintaining adequate sensitivity (true positive rate) across most scenarios. Random forest with Boruta/Hapfelmeier provided a viable alternative with similar sensitivity-specificity balance in complex non-linear data structures.
Confidence Interval: Not reported in the abstract; Monte Carlo standard errors for simulation estimates not provided
Clinical Application
All evaluated methods (linear regression with stepwise selection, elastic net, gradient boosting, random forest with Boruta/Hapfelmeier) are implemented in widely available statistical software (R packages: glmnet, mboost, gbm, randomForest, Boruta). LMSS is the most accessible method for clinical researchers without advanced machine learning expertise. The random forest alternatives require additional familiarity with variable importance frameworks but are feasible in most academic clinical research settings. In Australia, clinical prediction model development is relevant across primary care (RACGP guidelines emphasise evidence-based risk stratification tools), hospital-based medicine, and health technology assessment for PBS/TGA submissions. The RACGP and NHMRC both emphasise the importance of well-calibrated, validated prediction tools. The finding that LMSS provides reliable variable selection in low-dimensional settings supports its continued use in Australian clinical research contexts where sample sizes are often modest (e.g., single-centre studies, rare disease registries). For researchers developing models for PBS reimbursement submissions or TGA-regulated clinical decision support tools, the calibration findings are directly relevant to regulatory requirements for model performance. The study's guidance on minimum sample size requirements is particularly pertinent for Australian researchers working with smaller datasets from regional or Indigenous health cohorts. Clinical researchers, biostatisticians, and data scientists developing clinical prediction models from low-dimensional biomedical datasets (typically 10–20 candidate variables) with continuous outcomes. Applicable to settings such as laboratory value prediction, physiological parameter modelling, and clinical scoring system development.
Abstract
PURPOSE: A wide range of methods exist for developing a clinical prediction model (CPM) and for performing variable selection. Our purpose was to develop a fair simulation study design and to investigate the properties, strengths, and weaknesses of different methods to predict a continuous outcome in low-dimensional data situations. METHODS: In this simulation study, we conducted a neutral comparison of traditional (linear regression with stepwise selection) and machine learning (regularized regression with elastic net, gradient boosting, random forest) variable selection strategies to derive a CPM. The generated datasets included a total of 15 variables, with 8 of those being predictor variables. Four data- and outcome-generating mechanisms with increasing complexity produced data structures typical for biomedicine covering linear associations and gradually introducing non-linear and non-additive elements into the data structure. RESULTS: All methods generally performed better with increasing sample size and less noise in the data. Gradient boosting with regression models and with trees as base learners, and the elastic net regularized regression included nearly all variables (i.e., both the predictor and non-predictor variables), especially with increasing sample size. The linear regression model with stepwise selection (LMSS) showed the best trade-off between correctly including the predictors and excluding the non-predictor variables in most of the scenarios, even when the functional form of continuous predictors deviated from linearity. In more complex data, variable selection using the Boruta or Hapfelmeier approach for random forest performed similar to LMSS. CONCLUSION: The sample size must be sufficiently large to enable the methods to reliably identify the predictor variables and to ensure that the developed CPMs are accurate and well-calibrated. LMSS revealed good properties and the random forest with the Boruta or Hapfelmeier approach are suitable alternatives if complex associations between predictors and outcomes are assumed.
References
- 1.Vey, J. A., Heinze, G., & Kieser, M. (2021). Variable selection for clinical prediction models in low-dimensional data - a simulation study comparing traditional regression and machine learning methods. BMC Medical Research Methodology. https://doi.org/10.1186/S12874-021-01374-Y
Related Research
Minimally invasive therapy & allied technologies : MITAT : official journal of the Society for Minimally Invasive Therapy
Ventral hernia repair in emergency settings. A machine learning model to predict post-operative complications.
3 Aug 2026
Journal of the American Medical Informatics Association : JAMIA
Sociodemographic bias in large language model clinical trial screening
2 Aug 2026
Acta odontologica Scandinavica
Predicting gingival embrasure risk after invisible orthodontics using multimodal data and machine learning
23 July 2026
This content is for educational purposes for healthcare professionals only and does not constitute clinical advice. Clinical decisions should be based on individual patient assessment, current guidelines, and appropriate specialist consultation. Editorial Standards · Privacy Policy · Terms of Service