A moderated nonlinear factor analysis workflow for detecting and estimating multiple sources of differential item functioning and evaluating their impact on scale scores
Le résumé fourni par la source
Measurement invariance is essential for valid psychological comparisons, yet the most widely used approaches consider one source of non-invariance at a time. This risks oversimplifying measurement heterogeneity and potentially biasing comparisons when item functioning is associated with multiple correlated continuous and categorical characteristics. Moderated nonlinear factor analysis (MNLFA) can model multiple sources simultaneously, including covariate effects on latent traits (impact) and differential item functioning (DIF). We present an open-source regularised MNLFA workflow in R for DIF detection, unpenalised model re-estimation, and expected a posteriori (EAP) score generation for creating scale scores. The workflow uses regDIF to identify candidate DIF effects and OpenMx to re-estimate the selected model, avoiding regularisation-induced shrinkage in estimates, before evaluating whether DIF adjustment changes latent trait estimates. We demonstrate the workflow using the Short Warwick–Edinburgh Mental Wellbeing Scale in an integrative data analysis (IDA) framework, pooling six UK adolescent studies comprising 155,360 participants and 221,418 observations. Across three calibration samples, regularised selection identified a largely consistent set of item-level DIF effects related to age, sex, and study membership. Covariate-effect estimates remained stable across model specifications, and impact-plus-DIF and impact-only EAP scores showed near-perfect agreement (all r ≈ .99; root-mean-square differences = .008 – .016). Thus, adjustment for the detected DIF had negligible influence on latent wellbeing estimates. These findings show why item-level DIF should be distinguished from practically meaningful test-level bias. More broadly, the workflow provides a route for evaluating measurement invariance, harmonising measures, and generating comparable latent scores in heterogeneous or integrated datasets.
Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.