Prediction-based variable selection for component-wise gradient boosting

Sophie Potts; Elisabeth Bergherr; Constantin Reinke; Colin Griesbach

doi:10.1515/ijb-2023-0052

Prediction-based variable selection for component-wise gradient boosting

Int J Biostat. 2023 Nov 27. doi: 10.1515/ijb-2023-0052. Online ahead of print.

Authors

Sophie Potts¹, Elisabeth Bergherr¹, Constantin Reinke², Colin Griesbach¹

Affiliations

¹ Chair of Spatial Data Science and Statistical Learning, University of Goettingen, Goettingen, Germany.
² Chair of Empirical Methods in Social Science and Demography, University of Rostock, Rostock, Germany.

PMID: 38000054
DOI: 10.1515/ijb-2023-0052

Abstract

Model-based component-wise gradient boosting is a popular tool for data-driven variable selection. In order to improve its prediction and selection qualities even further, several modifications of the original algorithm have been developed, that mainly focus on different stopping criteria, leaving the actual variable selection mechanism untouched. We investigate different prediction-based mechanisms for the variable selection step in model-based component-wise gradient boosting. These approaches include Akaikes Information Criterion (AIC) as well as a selection rule relying on the component-wise test error computed via cross-validation. We implemented the AIC and cross-validation routines for Generalized Linear Models and evaluated them regarding their variable selection properties and predictive performance. An extensive simulation study revealed improved selection properties whereas the prediction error could be lowered in a real world application with age-standardized COVID-19 incidence rates.

Keywords: gradient boosting; high-dimensional data; prediction analysis; sparse models; variable selection.