BioBERT-Derived Clinical Text Representations and Sparse Principal Components of Multi-Omics Data for Breast Cancer Survival Prediction: A Leakage-Controlled Cross-Validated Benchmark

Read the full article See related articles

Listed in

This article is not in any list yet, why not save it to one of your lists.
Log in to save this article

Abstract

Background

Reported performance for multi-modal cancer prognostic models is frequently obtained from a single train–test partition and without comparison against a simple clinical baseline. We evaluated an integrated framework combining language-model representations of structured clinical records with high-dimensional multi-omics data under a leakage-controlled protocol.

Methods

Clinical variables from 589 TCGA-BRCA patients were converted to natural language sentences and encoded with BioBERT; DNA methylation, RNA, miRNA and protein data were reduced by Sparse principal component analysis. Clinical-only, multi-omics-only and integrated feature sets were evaluated with Cox proportional hazards, XGBoost and random survival forest (RSF) using five repetitions of stratified five-fold cross-validation with pooled out-of-fold predictions. Models were compared against a marginal Kaplan–Meier null and a conventional age-and-stage Cox model, with differences assessed by paired bootstrap.

Results

Among 589 patients with 88 deaths (14.9%; median follow-up 33.7 months), the integrated RSF model performed best (Harrell’s C 0.685, 95% CI 0.620–0.754; Uno’s C 0.748), exceeding the conventional clinical Cox model (0.637, 0.559–0.707). The ordering integrated > clinical-only > multi-omics-only held in all nine feature-set by algorithm comparisons and at all three component settings tested. The integrated model significantly outperformed every multi-omics-only model ( P = 0.008–0.044) but not clinical-only models (ΔC = 0.037, P = 0.128). Brier scores approached a null reference at 2 and 3 years and improved modestly at 5 years.

Conclusion

Integration produced the highest discrimination and exceeded a conventional clinical model, but the gain over clinical features within the same pipeline was not statistically resolvable at this event count. Reporting simple baselines, null-model calibration and cross-validated estimates should be standard in multi-modal prognostic modelling.

Article activity feed