Sparse data-driven random projection in regression for high-dimensional data
Rattachement africain : at. Niveau de preuve : code pays fourni par la source.
Le résumé fourni par la source
We examine the linear regression problem in a challenging high-dimensional setting with correlated predictors where the degree of sparsity of the coefficients is unknown and can vary from sparse to dense. In this setting, we propose a combination of probabilistic variable screening with random projection tools as a computationally efficient approach. In particular, we introduce a new data-driven random projection for dimension reduction in linear regression,which is motivated by a theoretical bound on the gain in expected prediction error over conventional random projections when using information about the true coefficient. The variables to be included in the projection are screened by considering the correlation of the predictors. To reduce the dependence on fine-tuning choices, we aggregate over an ensemble of linear models. A threshold parameter is introduced to obtain a higher degree of sparsity, which can be chosen together with the number of models in the ensemble by cross-validation.In extensive simulations, we compare the proposed method with other random projection tools and with well-known methods, and show that it is competitive in terms of prediction in a variety of scenarios with different sparsity and predictor covariance settings, while most competitors are targeted at either sparse or dense settings.Finally, we illustrate the method on two data applications.
Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.
Le contrôle bibliographique ouvert
DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.
- Titre Crossref
- Sparse data-driven random projection in regression for high-dimensional data
- Date Crossref
- 09/05/2025
- Éditeur
- International Association for Statistical Computing
- Type
- journal-article
Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude, et il ne compte pas comme une seconde source scientifique indépendante.
Les institutions déclarées
Une affiliation ne permet pas de déduire la nationalité d’un auteur.