Analysis code and figure source data for: How documentation practice inflates machine-learning prognostic models in COVID-19 — a leakage audit and admission-time model in 2433 hospitalised patients in Uzbekistan
Résumé fourni par la source
Analysis code and figure source data supporting a retrospective cohort study of 2433patients hospitalised with RT-PCR-confirmed SARS-CoV-2 infection in Tashkent,Uzbekistan, between 11 April and 10 August 2020. The study audits where the apparent discrimination of a machine-learning prognosticmodel comes from. Adding predictor blocks one at a time under identical folds,cross-validated AUROC rose from 0.896 with a conservatively specified core set to0.962 with all available predictors, and fell to 0.841 when the population wasrestricted to symptomatic patients. Once its membership functions were estimatedinside training folds only, a label-informed latent feature space performed worsethan mean substitution, multiple imputation and native gradient boosting. This deposit contains: the Python scripts that produced the cohort, the latentfeature space, the cross-validation, the provenance audit and the temporalvalidation; an R script and an Excel workbook that regenerate all six figures; thecomplete 619-item variable dictionary; the full performance table for all 25strategy-by-algorithm combinations; and confusion matrices across the thresholdrange. The patient records themselves are NOT included. They are individual-level clinicalrecords from a single hospital, abstracted under an ethics approval (Ministry ofHealth of the Republic of Uzbekistan, reference 6/1-1444) that does not provide foropen release, and are available from the corresponding author on reasonable request.The file closest to patient-level data, the Fig4_predictions worksheet, containsmodel output and a binary outcome only, with no clinical, demographic or identifyingfields.