Multi-source domain generalization with few-shot calibration for cross-dataset EEG state classification under proxy labels
Listed in
This article is not in any list yet, why not save it to one of your lists.Abstract
Cross-dataset generalization of EEG-based classification under weak, proxy-derived labels remains an open problem for altered-states research. We present a reproducible eight-dataset alignment pipeline that maps eight heterogeneous EEG corpora (712,832 windows; 697,906 with valid labels) to a common 14-channel EPOC+ montage with 63-dimensional spectral features, and we recover the real 1–9 arousal self-assessments for MAHNOB-HCI from session.xml. Random Forest classifiers are trained on seven source domains and evaluated on the held-out target under both zero-shot (ZS) and 20%-participant few-shot calibration. The benchmark exposes two methodological pitfalls rather than a performance result: (i) per-class recall shows that all targets but DEAP collapse to a single majority class, and (ii) a within-dataset upper-bound experiment (Table 3) shows that of eight proxy label sets, one is learnable within-dataset, two are marginal, and five sit at or below three-class chance even when trained and tested on the same dataset, so the cross-dataset failure is a label-validity problem rather than a transfer-method problem. Across the eight targets (20 seeds, 8,000 windows each), zero-shot accuracy averages 36.85% (95% CI 34.40–39.30) and calibrated 43.76% (41.77–45.75), but zero-shot balanced accuracy stays at 33.01–35.62% (Cohen’s κ ≤ 0.068), i.e. at chance. The +6.91pp mean change is driven almost entirely by one target, ds006437 (6.31% → 60.60%); after Holm–Bonferroni correction only ds006437 and ds004572 remain significant, the latter with a practically null effect (+0.39pp). The collapse persists under SMOTE oversampling, an EEGNet-v4 baseline, and CORAL/AdaBN feature alignment, locating the bottleneck in proxy-label validity and feature-space class overlap rather than classifier capacity.