Clinical trial success prediction from open registry text using classical machine learning
Listed in
This article is not in any list yet, why not save it to one of your lists.Abstract
Drug candidates that fail in clinical trials may expose participants to unsafe or ineffective interventions and divert limited development capacity from more promising programs. This study assessed whether public registry text could be used to estimate clinical trial outcomes. The Clinical Trial Outcome Dataset (CTOD), developed by Gao and colleagues, provides open labels for about 125,000 clinical trials. Logistic regression, random forest, and gradient boosting models were built to use only ClinicalTrials.gov text, with temporal hold-out and evaluation across Phases 1, 2, and 3. To reduce leakage, fields directly related to outcome were removed, and 187 pre-2019 trials containing post-outcome text were excluded before the 27 models were refit. Random forest had the highest mean ROC-AUC (0.668) and mean Brier skill (0.069), while the strongest configuration reached ROC-AUC 0.738 in Phase 1. Against 7,260 manually adjudicated hard-case trials, phase-specific random-forest ROC-AUCs were 0.668 to 0.686. All-indication training improved ROC-AUC for neuroscience in all nine paired comparisons. Deduplication, removal of organizational fields, and targeted cleanup of residual outcome-related language did not materially change the results. Together, these findings support the use of open registry text in a low-compute workflow for benchmarking trial outcomes.