Is linear regression sufficient for normative psychological assessment? A comparison between OLS, GAM, and GAMLSS across seven public reference datasets
Listed in
This article is not in any list yet, why not save it to one of your lists.Abstract
Background: Normative Z-scores are essential for interpreting individual cognitive performance relative to healthy reference populations. Traditional norms often compare individuals within predefined groups, such as age bands. This is useful and easy to interpret, but it can lose information because variables such as age change continuously and because age, sex, and education may contribute jointly to performance. Regression-based norms estimate the expected score for each individual from their specific covariates and then express the difference between the observed and expected score as a normative Z-score. Therefore, the key methodological question is not simply whether a more complex model can be fitted, but whether that added complexity changes the Z-score in a meaningful way. Methods: We implemented a comparative framework and fitted ordinary least squares (OLS), generalized additive models (GAM), and generalized additive models for location, scale, and shape (GAMLSS) to the seven public reference datasets distributed in the NormData R package. The datasets span processing speed, verbal memory, verbal fluency, academic achievement, personality, and anxiety. For each dataset we compared each regression model with the classic subgroup norm, and then compared the regression models with one another. Metrics included Pearson r, residual RMSE, and the percentage of participants whose model-based Z-score agreed with the classic Z-score. Results: All three models reproduced the classic Z-scores closely (r = 0.964-1.000). OLS and GAM produced nearly identical Z-scores throughout (r = 0.996-1.000), indicating that the observed age trajectories were smooth enough for a quadratic term to capture. GAMLSS diverged from OLS mainly in datasets with non-constant residual variance (Fluency, sigma-ratio = 1.79; TMAS, sigma-ratio = 1.50), where modeling the scale parameter recovered dispersion differences that the classic subgroup method encodes implicitly. Conclusions: Across seven heterogeneous reference datasets, OLS provides a sufficient normative Z-score computation when the sample size is large enough and residual variance is approximately constant. GAM is useful when the mean trajectory is strongly nonlinear; GAMLSS is useful when residual variance changes across covariate profiles. The practical message is deliberately simple: add model complexity only when the data show what that complexity contributes to the individual Z-score.