Cross-Validated Loss-Based Covariance Matrix Estimator Selection in High Dimensions

Philippe Boileau; Nima S Hejazi; Mark J van der Laan; Sandrine Dudoit

doi:10.1080/10618600.2022.2110883

Cross-Validated Loss-Based Covariance Matrix Estimator Selection in High Dimensions

J Comput Graph Stat. 2023;32(2):601-612. doi: 10.1080/10618600.2022.2110883. Epub 2022 Oct 7.

Authors

Philippe Boileau¹, Nima S Hejazi², Mark J van der Laan³, Sandrine Dudoit⁴

Affiliations

¹ Graduate Group in Biostatistics and Center for Computational Biology, UC Berkeley.
² Division of Biostatistics, Department of Population Health Sciences, Weill Cornell Medicine.
³ Division of Biostatistics, Department of Statistics, and Center for Computational Biology, UC Berkeley.
⁴ Department of Statistics, Division of Biostatistics, and Center for Computational Biology, UC Berkeley.

Abstract

The covariance matrix plays a fundamental role in many modern exploratory and inferential statistical procedures, including dimensionality reduction, hypothesis testing, and regression. In low-dimensional regimes, where the number of observations far exceeds the number of variables, the optimality of the sample covariance matrix as an estimator of this parameter is well-established. High-dimensional regimes do not admit such a convenience. Thus, a variety of estimators have been derived to overcome the shortcomings of the canonical estimator in such settings. Yet, selecting an optimal estimator from among the plethora available remains an open challenge. Using the framework of cross-validated loss-based estimation, we develop the theoretical underpinnings of just such an estimator selection procedure. We propose a general class of loss functions for covariance matrix estimation and establish accompanying finite-sample risk bounds and conditions for the asymptotic optimality of the cross-validation selector. In numerical experiments, we demonstrate the optimality of our proposed selector in moderate sample sizes and across diverse data-generating processes. The practical benefits of our procedure are highlighted in a dimension reduction application to single-cell transcriptome sequencing data.

Keywords: covariance matrix estimation; cross-validation; dimension reduction; high-dimensional statistics; loss-based estimation.

Grants and funding

R01 DC007235/DC/NIDCD NIH HHS/United States