Real-world performance of large-scale propensity score adjustment strategies: Matching, weighting, and stratification

Read the full article See related articles

Discuss this preprint

Start a discussion What are Sciety discussions?

Listed in

This article is not in any list yet, why not save it to one of your lists.
Log in to save this article

Abstract

OBJECTIVE Propensity score (PS) models are commonly used for addressing confounding in observational studies. Researchers are increasingly incorporating large numbers of covariates into PS models, using techniques like large-scale propensity score (LSPS) adjustment. It remains unclear if the inclusion of many covariates affects the subsequent choice of adjustment strategy. METHODS In this paper, we evaluate commonly used adjustment strategies using large-scale models, including matching, stratification, and inverse probability of treatment weighting (IPTW), to estimate treatment effects under a new-user cohort design. Evaluations employ real-world data across four national healthcare databases and incorporate 3,840 treatment effect estimates per model specification for 24 different treatment cohort comparisons against 160 negative control outcomes. We assess adjustment strategies for covariate balance between treatment cohorts and type-I and type-II error, mean-squared error, and precision of the adjusted effect estimate. RESULTS Across all model specifications, IPTW with symmetric Crump trimming and 1:1 matching perform best overall, achieving strong balance and low bias. Additionally, empirical calibration of effect estimates can substantially reduce differences between strategies. DISCUSSION No single strategy is clearly superior in all settings. We recommend incorporating the top-performing strategies at least as sensitivity analyses in studies even if researchers anticipate that some other strategy may excel due to expectation of their study's operating characteristics. We also recommend use of diagnostics like covariate balance to check strategy performance. CONCLUSION IPTW with Crump trimming and 1:1 matching are strong default choices for large-scale PS adjustment, but strategy selection should be guided by study context and validated with balance and error diagnostics. We broadly recommend empirical calibration to reduce sensitivity to model choice.

Article activity feed