A Simulation Study Comparing Multiple Imputation and Complete Case Analysis for Handling Missing Preschool Body Mass Index
Discuss this preprint
Start a discussion What are Sciety discussions?Listed in
This article is not in any list yet, why not save it to one of your lists.Abstract
Introduction
Missing data frequently occurs in health databases and can bias analyses if not correctly dealt with.
Objectives
Using real-world data where missing values were present in as much as 30% of our sample, we compared complete-case and multiple-imputation methods for recovering true parameters of a multivariable logistic regression model for the association between maternal glucose levels during pregnancy and child excess weight at preschool age.
Methods
This study utilized a cohort of 130,424 children from two Canadian urban health zones with complete preschool-age body mass index (BMI) measurements linked to administrative health data. We introduced missingness through deletion following three distinct mechanisms: missing completely at random (MCAR), at random (MAR), and not at random (MNAR). We employed complete-case and multiple-imputation methods to handle the introduced missingness. The associations between five categories of maternal glucose levels during pregnancy with child excess weight at pre-school age were determined from logistic regression models using the full observed data (true values), observed data without deletions (complete-case estimates), and imputed data (multiple-imputation estimates). Accuracy of complete-case and multiple-imputation estimates were evaluated against true values. Finally, we conducted a sensitivity analysis for the MNAR mechanism using pattern-mixture models with an additive shift.
Results
Under MCAR and MAR, multiple-imputation generally yielded larger bias and relative bias, but smaller or similar mean square error and reached higher significance than complete-case analysis. Both methods showed high significance (≥ 0.96) for most effects and high coverage (≥ 0.99) consistently. Under MNAR, both complete-case and multiple-imputation showed poor performance regarding bias and statistical significance. Sensitivity analysis indicated performance varied by specific effect.
Conclusions
When faced with missing data, researchers should assess missingness mechanisms, report both complete-case and multiple-imputation estimates under MCAR/MAR while accounting for power-versus-bias tradeoffs, and employ pattern-mixture sensitivity analyses to test robustness when MNAR is plausible.