Structural, Functional and Cognitive Validation of a Deep-Learning MCI-to-AD Conversion Model in OASIS-3
Listed in
This article is not in any list yet, why not save it to one of your lists.Abstract
Background
Deep-learning models predict mild cognitive impairment (MCI)-to-Alzheimer’s disease (AD) conversion from structural MRI with high accuracy. However, they are typically validated within a single cohort and judged on discrimination alone. Whether their risk scores are biologically grounded, and whether they generalise to independent data, remains unclear.
Methods
We applied the ADNI-trained Temporal Adaptive Fusion Network (TAF-Net), without retraining, to 101 MCI participants from OASIS-3 (25 converters, 76 stable) and tested whether its conversion-risk scores were grounded in independent structural, functional, and cognitive markers of AD. Regional atrophy rates were derived from longitudinal FreeSurfer, cognition from longitudinal MMSE and CDR-Sum-of-Boxes, and baseline resting-state functional connectivity from fMRIPrep, in structural and functional subsamples of 41 and 40 participants. Associations used rank-based statistics with false-discovery-rate correction.
Results
Higher TAF-Net risk tracked faster atrophy in medial-temporal AD-signature regions but not in AD-spared cortex — an anatomically specific coupling that survived adjustment for global atrophy. In external validation, risk discriminated converters (AUC = 0.72), comparable to native atrophy and strongly concordant with it; atrophy statistically accounted for the model’s prognostic signal. Discrimination transferred but calibration did not: the two lowest tertiles of risk were assigned near-zero probability yet converted at 15%. Risk also tracked the multi-year rate of cognitive decline and, at baseline, was associated with reduced within-network functional-connectivity integrity, concentrated in salience and default-mode hubs; longitudinal functional analyses were underpowered.
Conclusions
A conversion model trained on one cohort produced risk scores that, in an independent cohort, were grounded in the structural, functional, and cognitive hallmarks of Alzheimer’s disease — supporting biological validity and external generalisation. The score, however, largely re-expresses the neurodegenerative substrate captured by structural atrophy. Rank ordering transferred across cohorts but absolute risk did not, so the score requires recalibration before its values can be interpreted as individual probabilities.