Retrospective development and external validation of machine learning models for COVID-19 severity prediction in lung and hematological malignancies: evidence of subtype-specific feature selection bias in joint training
Rattachement africain : de, fr, us, gb. Niveau de preuve : code pays fourni par la source.
Le résumé fourni par la source
Patients with lung or hematological malignancies face an elevated risk of severe COVID-19 outcomes. While machine learning frameworks offer predictive utility, training a single joint clinical-biological model on heterogeneous cancer populations can introduce structural feature selection bias, limiting clinical generalizability. Utilizing a multicenter retrospective cohort (n = 340) for model development and an independent external cohort (103) for validation, we evaluated an end-to-end machine learning pipeline. Initially, a joint 7-variable signature—comprising age, ECOG performance status, lymphopenia, BMI, albumin, creatinine, and SARS-CoV-2 serology—was analyzed using logistic regression (LR) and random forest (RF) algorithms. Phenotypic deconstruction via cluster-level ablation and cancer-type-stratified optimal feature searches were subsequently executed to evaluate population-specific feature mechanics. The primary joint LR model achieved an optimism-corrected internal AUC of 0.75 (95% CI: 0.69–0.79) and an external validation AUC of 0.72 (95% CI: 0.62–0.82). However, cluster ablation and subgroup stratification revealed a critical biological divergence between the cancer types. The joint model's feature selection prioritized variables that acted as predictive noise for specific populations; notably, the apparent predictive contribution of albumin and creatinine in the joint model diminished or reversed after subtype-specific stratification, suggesting these variables largely reflect mixed-population effects rather than robust subtype-specific biological signal; their removal improved lung cancer model performance (albumin ΔAUC +0.043, creatinine +0.036) and did not materially benefit hematological patients either. The two cancer-type-optimal signatures shared age, ECOG performance status and lymphopenia, but diverged in their remaining components — BMI and albumin for hematological patients versus diabetes and immunodepression status for lung patients — indicating that the joint model's discriminative variables are only partially transferable across subtypes. While a unified 7-variable framework provides acceptable predictive baseline performance, joint training on mixed cancer populations introduces systematic feature selection bias. These exploratory findings suggest that cancer-type-stratified models may better capture subtype-specific biological signatures than a single joint model; however, independent prospective validation with pre-specified subtype stratification is required before any clinical deployment. Not applicable.
Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.
Le contrôle bibliographique ouvert
DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.
- Titre Crossref
- Retrospective development and external validation of machine learning models for COVID-19 severity prediction in lung and hematological malignancies: evidence of subtype-specific feature selection bias in joint training
- Date Crossref
- 27/08/2026
- Éditeur
- Springer Science and Business Media LLC
- Type
- journal-article
Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude, et il ne compte pas comme une seconde source scientifique indépendante.
Les institutions déclarées
Une affiliation ne permet pas de déduire la nationalité d’un auteur.