Aller au contenu principal
Accès ouvert déclaré 2026 software

Replication materials for How Should Missing Data Be Handled When Fitting Conditional Inference Trees?

0Citations signalées, ce qui n’est pas une note de qualité
1Institutions déclarées
1Pays d’affiliation déclarés

Rattachement africain : us. Niveau de preuve : code pays fourni par la source.

Le résumé fourni par la source

Replication materials for the manuscript "How Should Missing Data Be Handled When Fitting Conditional Inference Trees?" (Sherlock). This archive contains the Monte Carlo simulation code and complete numerical results for a study comparing six approaches to handling missing data when fitting conditional inference trees (ctree; Hothorn, Hornik & Zeileis, 2006): listwise deletion, surrogate splits, missingness incorporated in attributes (MIA), single imputation, uncorrected stacking of multiply imputed datasets, and stacking with a sample-size correction (Stack/M). Simulation design Ten simulation studies spanning 189 conditions. Each dataset contained eight predictors — four signal-carrying (binary, continuous, five-level ordered, four-level unordered) and four pure-noise predictors matched type for type, so that splits on noise predictors index variable-selection bias by measurement type. Five data-generating processes (null, main effects, interaction, ordinal/nominal, weak signal) were crossed with three missingness mechanisms (MCAR, MAR, MNAR) at rates of 15%, 30% and 45%, across sample sizes from n = 250 to n = 25,000. Imputations were generated with MICE, with M = 30 throughout except in Study 5, which varies M from 5 to 50. Studies 8–10 repeat the comparison with predictors correlated through a latent Gaussian factor model. Studies 1–5, 8 and 9 used 500 replications per condition; Studies 6, 7 and 10 used 200. Recorded per replicate: spurious subgroup, recovery of the complete true structure, Brier score or MSE on an independently generated test set of 2,000 fully observed cases, terminal-node count, rows available to the approach, and whether a tree was produced. Reproducibility was assessed by the adjusted Rand index between the terminal-node partitions of trees refitted on independent imputation draws. Contents ctreeMI-simulation-code.zip — R scripts implementing the data-generating processes, missingness mechanisms, imputation, fitting and analysis. See the top-level README for a description of each script and a mapping from scripts to the tables and figures in the manuscript. results-simulation.zip — complete numerical results underlying all tables and figures. Reproducibility R 4.6.1 (2026-06-24); ctreeMI 1.0.0, partykit 1.2-29, mice 3.19.0. Each replicate's seed is a deterministic function of its condition and replicate index and is set within the worker process, so results do not depend on the number of processor cores used or on the order in which jobs completed. The Stack/M correction extracts node-level test statistics from partykit internals, so partykit 1.2-29 should be used to reproduce the published results exactly. The applied examples use the brandsma and boys datasets distributed with the mice package. No restricted data are included. Related resources ctreeMI on CRAN: https://CRAN.R-project.org/package=ctreeMI Stack/M correction: Sherlock et al. (2026), Multivariate Behavioral Research, doi:10.1080/00273171.2026.2661244

Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.

Le contrôle bibliographique ouvert

La source scientifique ouverte est momentanément indisponible.

Les institutions déclarées

Une affiliation ne permet pas de déduire la nationalité d’un auteur.

BNTIC News n’est pas le producteur de ces données. Les publications sont interrogées à la demande dans Crossref, OpenAIRE, DOAJ, Europe PMC, HAL, DataCite, AfricArXiv, ROR et la Banque mondiale, sans clé d’accès. OpenAlex reste optionnel. Aucun service payant n’est nécessaire et aucune donnée externe n’est enregistrée en base. Consulter les sources et leurs limites.