The Influence of Missing Data Handling Methods on CES-D Analysis in the China Health and Retirement Longitudinal Study (CHARLS)

Read the full article See related articles

Discuss this preprint

Start a discussion What are Sciety discussions?

Listed in

This article is not in any list yet, why not save it to one of your lists.
Log in to save this article

Abstract

Most evaluations of missing data analyses rely on literature reviews or comparisons of hypothesized methods applied to a given dataset, yet the real-world impact of different strategies remains unknown. Using the widely analyzed CHARLS dataset, we reviewed missing data handling methods in 219 studies using the CES-D-10, and compared the three most prevalent procedures: complete-case analysis (CCA), impute zero (IZ), and multiple imputation (MI). Although CES-D-10 items were not missing completely at random, 90% of publications used CCA, resulting in lower mean CES-D-10 scores but higher depression prevalence, especially among adults aged 75 and older. The IZ CES-D-10 score constructed in the harmonized CHARLS dataset systematically underestimated both prevalence and mean scores, distorting age trajectories and attenuating sex differences in later life. Many CHARLS-based depression studies are therefore likely affected by systematic bias due to suboptimal treatment of missing CES-D-10 items, particularly those studies using the harmonized dataset. Moreover, because the IZ approach has similarly been used to construct CES-D scores in other major aging studies (e.g., HRS, ELSA, and KLoSA), this bias may extend beyond CHARLS and undermine the validity of depression findings across multiple international cohorts.

Article activity feed