Whose Truth Is Ground Truth?: Consequences of Label Choice on Machine Learning Models Predicting Depression
Discuss this preprint
Start a discussion What are Sciety discussions?Listed in
This article is not in any list yet, why not save it to one of your lists.Abstract
Machine learning (ML) models developed using electronic health record (EHR) data frequently rely on provider-documented diagnoses as ground truth to develop models, despite substantial variation in how mental health conditions are diagnosed and recorded. To empirically demonstrate the impact of diagnostic label source, we examined how use of patient-reported, provider-coded, or patient–provider concordant depression labels impact the performance and feature importance of tree-based ML models. We analyzed EHR data from 644,387 adults (2012–2019) in the OCHIN ADVANCE network. Depression outcomes were defined using (1) patient-reported PHQ-9 scores, (2) provider-coded ICD diagnoses, and (3) concordance. Predictors included demographics, vital signs, chronic conditions, and social determinants of health. Classification and regression trees (CART) were trained using 70/30 splits and 10-fold cross-validation. Performance was assessed using sensitivity, specificity, AUC, and F1 score. SHapley Additive exPlanations (SHAP) quantified feature importance. Model performance was modest (AUC = .62–.64); sensitivity was highest for provider-indicated depression (0.33). Feature importance differed substantially by labeling method: anxiety was the only consistently influential predictor across label strategies. Demographic features were highly influential in provider-labeled models but less important in patient-reported or concordant label strategies. We empirically demonstrated how depression label choice alters ML model performance and feature interpretation. Higher AUC values from provider-derived labels were offset by amplified demographic feature importance, raising concerns about the interaction of label choice and bias. Our findings underscore the importance of transparent reporting, comparing labeling strategies, and including patient-reported outcomes in model development to support equitable mental health AI.