Genetic Architecture and Sample Size Impact Relative Performance of Nonlinear Machine Learning and Standard Polygenic Risk Scores
Discuss this preprint
Start a discussion What are Sciety discussions?Listed in
This article is not in any list yet, why not save it to one of your lists.Abstract
Standard polygenic risk scores (PRSs) are constructed based on additive genome-wide association study summary statistics. Nonlinear machine learning methods have been increasingly applied to construct PRSs from individual-level data, aiming to improve predictive performance over standard PRSs through modelling nonadditive genetic effects. However, their superiority across studies has been inconsistent. The conditions under which they provide meaningful improvements remain unclear. We combined theory, simulations and a real-world application to investigate when two widely used nonlinear machine learning methods, random forest and XGBoost, outperform standard PRSs. Theoretically, we showed that standard PRSs can implicitly capture some genetic variance attributable to nonadditive genetic effects through their contributions to marginal SNP effects. Although nonlinear models have a higher theoretical potential, their bias-variance trade-off can limit predictive gains at finite sample sizes. Simulations showed that XGBoost outperformed standard PRSs only when the genetic architecture involves a large proportion of interaction genetic variance concentrated across relatively few interactions and training sample sizes are large. Random forest consistently underperformed standard PRSs. In risk prediction of ischemic heart disease using UK Biobank data, XGBoost showed little improvement in predictive performance over standard PRSs, whereas random forest again performed worse. Together, these findings suggest that nonlinear machine learning do not uniformly outperform standard PRSs; rather, their relative performance depends jointly on genetic architecture and training sample size. Our study helps to reconcile the inconsistent results reported across previous studies and provides a framework for identifying settings in which more complex PRS models are likely to be beneficial.