Generalizability of EEG-Based Dementia Classifiers: A Multicenter study of Alzheimer’s, MCI, And FTD
Discuss this preprint
Start a discussion What are Sciety discussions?Listed in
This article is not in any list yet, why not save it to one of your lists.Abstract
EEG-based machine learning shows promise for neurodegenerative disease classification, but robustness to sample imbalance, center heterogeneity, and validation leakage remains a key concern for clinical translation. We developed a new framework to assess diagnostic performance, calibration, and cross-center generalizability of EEG multifeatured classifiers across CN (cognitively normal), MCI (mild cognitive impairment), AD (Alzheimer’s disease), and FTD (frontotemporal dementia), while addressing imbalance, statistical uncertainty, and validation rigor across six centers. Supervised classifiers were evaluated at aggregated- and subject-level repeated cross-validation and leave-one-center-out (LOCO) schemes, and calibration was implemented via Platt scaling within strictly nested folds. CN vs AD classification showed robust performance and cross-center generalizability, with consistent AUC and calibration across cross-validation and leave-one-center-out analyses. In contrast, CN versus MCI showed moderate, heterogeneous performance and limited cross-center generalizability, with chance-level results in some cohorts, while MCI versus AD showed moderate discrimination in a single available center. FTD contrasts showed modest or limited performance due to sparse samples. Predicted probabilities were stable across validation regimes for AD, but less consistent for MCI and FTD, and correlated robustly with cognitive impairment severity only for AD. Feature importance analyses identified disease-specific signatures, including alpha-band degradation and slow-wave increases in AD, with weaker and more heterogeneous patterns in prodromal and differential dementia contrasts (FTD vs AD). EEG classifiers provided robust discrimination for CN vs AD but showed limited and heterogeneous performance for MCI and FTD across centers. These results emphasize the need for balanced sampling, strict validation of clinical and EEG protocols, and uncertainty quantification to support reliable clinical deployment.