ChemoCalib: multiblock PLS calibration of genome-scale metabolic models improves flux prediction over expression-only integration
Discuss this preprint
Start a discussion What are Sciety discussions?Listed in
This article is not in any list yet, why not save it to one of your lists.Abstract
Motivation
Constraint-based metabolic modeling faces a calibration gap: genome-scale metabolic models (GEMs) integrated with transcriptomics alone rely on expression-to-flux heuristics (E-Flux, GIMME, iMAT, MOMENT) that ignore cross-omics co-variance structure and lack statistical mechanisms for propagating omics uncertainty into reaction bounds, yielding flux predictions with limited agreement to 13 C metabolic flux analysis (MFA) measurements.
Results
We present Chemo-Calib, a multiblock PLS (MB-PLS) framework that calibrates GEM reaction bounds from the shared latent structure of metabolomics, transcriptomics, and proteomics data. On 11 E. coli 13 C-MFA reference conditions spanning the Keio fluxome and Holm 2010 datasets, ChemoCalib constrained FBA on iJO1366 achieves a Spearman ρ = 0.461 overall (up to 0.523 in PPP) and Pearson r of 0.49–0.58 across central carbon pathways, with statistically significant improvement over expression-only baselines including E-Flux2 and SPOT ( p < 0.05, Holm-corrected). The latent-to-constraint mapping employs GPR-aware VIP aggregation (Algorithm 1) to project multi-omics latent scores onto genome-scale reaction bounds without heuristic thresholding. An optional in-silico active learning loop (relegated to Supplementary Material) further tightens calibration through virtual experiment selection.
Availability
ChemoCalib is open-source (MIT) at https://github.com/chemocalib/chemocalib with Docker support, a 5-minute tutorial, and pre-computed iJO1366 benchmarks. Preprint available at bioRxiv; code archived at Zenodo DOI: 10.5281/zenodo.21645890.