BioBERT-Derived Clinical Text Representations and Sparse Principal Components of Multi-Omics Data for Breast Cancer Survival Prediction: A Leakage-Controlled Cross-Validated Benchmark
Listed in
This article is not in any list yet, why not save it to one of your lists.Abstract
Background
Reported performance for multi-modal cancer prognostic models is frequently obtained from a single train–test partition and without comparison against a simple clinical baseline. We evaluated an integrated framework combining language-model representations of structured clinical records with high-dimensional multi-omics data under a leakage-controlled protocol.
Methods
Clinical variables from 589 TCGA-BRCA patients were converted to natural language sentences and encoded with BioBERT; DNA methylation, RNA, miRNA and protein data were reduced by Sparse principal component analysis. Clinical-only, multi-omics-only and integrated feature sets were evaluated with Cox proportional hazards, XGBoost and random survival forest (RSF) using five repetitions of stratified five-fold cross-validation with pooled out-of-fold predictions. Models were compared against a marginal Kaplan–Meier null and a conventional age-and-stage Cox model, with differences assessed by paired bootstrap.
Results
Among 589 patients with 88 deaths (14.9%; median follow-up 33.7 months), the integrated RSF model performed best (Harrell’s C 0.685, 95% CI 0.620–0.754; Uno’s C 0.748), exceeding the conventional clinical Cox model (0.637, 0.559–0.707). The ordering integrated > clinical-only > multi-omics-only held in all nine feature-set by algorithm comparisons and at all three component settings tested. The integrated model significantly outperformed every multi-omics-only model ( P = 0.008–0.044) but not clinical-only models (ΔC = 0.037, P = 0.128). Brier scores approached a null reference at 2 and 3 years and improved modestly at 5 years.
Conclusion
Integration produced the highest discrimination and exceeded a conventional clinical model, but the gain over clinical features within the same pipeline was not statistically resolvable at this event count. Reporting simple baselines, null-model calibration and cross-validated estimates should be standard in multi-modal prognostic modelling.