Pre-Transformer Models for Longevity Science and Deep Ageing Clocks: An Empirical Analysis

Read the full article

Discuss this preprint

Start a discussion What are Sciety discussions?

Listed in

This article is not in any list yet, why not save it to one of your lists.
Log in to save this article

Abstract

The transformer now dominates artificial intelligence, and its diffusion into computational biology raises a concrete question for the ageing field: on the data that ageing research actually has, do attention-based models outperform the pre-transformer methods—penalized regression, kernel machines, tree ensembles, and the pre-2017 deep networks (multilayer perceptrons, convolutional and recurrent nets, autoencoders)— that built the ageing-clock paradigm? Comparative claims have rested largely on the structure of the data rather than on head-to-head experiments. We supply the experiment. We benchmark eight model classes for biological-age prediction across six public cohorts spanning five modalities: blood DNA-methylation (GSE40279), blood biochemistry (NHANES), whole-blood transcriptome (GTEx), gut microbiome (American Gut), wearable accelerometry (NHANES), and brain MRI (Cam-CAN). Every model is tuned and evaluated under one nested cross-validation protocol on identical splits, and we score not only accuracy but data efficiency, compute, interpretability, cross-cohort transfer, and—critically—the biological validity of each model’s age-acceleration residual against all-cause mortality. Pre-transformer models win outright on five of the six cohorts and essentially tie on the sixth. On tabular molecular omics, penalized regression and gradient boosting are best or tied-best (methylation MAE 3.4 yr, transcriptome 8.1 yr); on genuinely structured modalities the modality-native pre-transformer deep net wins (accelerometry ConvLSTM MAE 12.7 yr, brain-MRI 3D-CNN 4.7 yr). From-scratch transformers never win and often overfit; a pretrained transformer edges ahead only on the single large-𝑛 cohort (𝑛=34,000), by a practically negligible 0.10 yr, and only past a training size of ≈23,000 labelled samples—larger than almost any human ageing cohort—while costing three to four orders of magnitude more compute. Aggregate accuracy is, in fact, a near-tie; the pre-transformer case is decided elsewhere. First, minimizing chronological-age error does not maximize biological signal: the model with the lowest MAE (the pretrained transformer) carries the weakest mortality association, and both transformers rank at the bottom for risk stratification (HR ≈1.13 per SD versus ≈1.28 for elastic net), an empirical echo of “too much of a good thing.” Second, penalized regression transfers across cohorts best of all (+1.2 yr degradation), while the from-scratch and non-pretrained deep models degrade more than twice as much (+2.8–4.0 yr). For the low-sample, high-dimensional, interpretability-sensitive, and regulator-facing regime that defines ageing research, the pre-transformer toolkit is not merely adequate but the rational default; transformers are a targeted addition, not a replacement.

Article activity feed