Aller au contenu principal
Accès ouvert déclaré2026article

Noise-Dependent Robustness of XGBoost, LightGBM, and CatBoost

0Citations signalées
4Institutions associées
1Pays d’affiliation

Résumé fourni par la source

Gradient-boosted decision trees (GBDTs) are among the leading methods for tabular data; however, their comparative robustness to corrupted labels remains unclear. This study resolves that uncertainty with a controlled benchmark protocol in which library identity is the only free variable, so that an observed difference can be attributed to the implementation rather than to tuning, data splits, or noise realizations, and in which capacity, noise floor, and dataset-selection confound checks are mandatory before any ranking is reported. We benchmarked Extreme Gradient Boosting (XGBoost), Light Gradient-Boosting Machine (LightGBM), and Categorical Boosting (CatBoost) under symmetric, asymmetric pair-flip, and instance-dependent label noise across 15 datasets from the OpenML Curated Classification benchmark suite 2018 (OpenML-CC18), four noise rates, and three random seeds. The clean data-tuned configurations were fixed across the noise conditions. Predictive performance is the macro-averaged F1 score (macro-F1) on a clean test set, and degradation is the absolute drop in macro-F1 relative to each library’s own clean-data score on the same dataset, so the recoveries reported below are absolute percentage points of macro-F1. The ranking depends on both the noise rate and the noise model: no library separated at a 10% rate under symmetric or asymmetric noise; separation emerged from 20% under symmetric noise and only at 40% under asymmetric noise; and under instance-dependent noise, it weakened as the rate rose. At 40% symmetric and asymmetric noise, CatBoost showed significantly lower per-dataset degradation than LightGBM (Friedman tests, both p<0.001), with mean ranks of 1.20 versus 2.47 and 1.20 versus 2.67, respectively, and with XGBoost intermediate in both cases (2.33 and 2.13) and separable from CatBoost under symmetric noise only. At 40% instance-dependent noise, the ranking disappeared: no significant library ranking was detected (p=0.63), and performance gaps were smaller than twice the pooled seed standard deviation, which shows that GBDT robustness conclusions are noise-model-dependent. Rankings also varied by evaluation dimension: LightGBM had the most stable feature importances in observed means, calibration rankings depended on the noise model, and CatBoost’s training loss was the strongest mislabel-detection signal. Under symmetric noise only, small-loss reweighting significantly improved all three libraries, while early stopping recovered up to 8.5 pp for LightGBM; mitigations were not evaluated under asymmetric or instance-dependent noise. For practitioners, this means the library and the mitigation should be chosen together with the noise process that is expected: prefer CatBoost when label-conditional noise is likely and accuracy is the objective; expect no library to buy robustness under feature-dependent noise at high rates; apply early stopping or small-loss reweighting under symmetric noise, where both give significant gains and early stopping helps LightGBM most; and select on calibration, mislabel detectability, or importance stability instead when the deployment depends on those, because the ranking differs by dimension.

Institutions

Sujets associés

Machine Learning and Data ClassificationExplainable Artificial Intelligence (XAI)Imbalanced Data Classification Techniques

BNTIC News n’est pas le producteur de ces données. Métadonnées interrogées à la demande auprès de OpenAlex (CC0). Sources et limites.