Ontology-Guided Pathway Activity Identifies a Cell-Intrinsic Defense Response Program Associated with MEK Inhibitor Sensitivity
Discuss this preprint
Start a discussion What are Sciety discussions?Listed in
This article is not in any list yet, why not save it to one of your lists.Abstract
Predicting cancer drug response from gene expression requires models that expose which biological pathways drive cell-line-specific sensitivity. We introduce Gene-Ontology Pathway Attention (GOPA), whose core module computes deterministic attention weights softmax( x c · B̂ ) from expression and the column-normalized gene-term annotation matrix with no learned parameters. Applying GOPA to 542 drugs across the GDSC panel under leave-cell-line-out evaluation, we find that defense response pathways predict sensitivity to kinase inhibitors in GDSC, with the strongest and most confound-resistant signal in MEK/MAPK inhibitors. This association shows cross-assay support in PRISM (11 of 11 overlapping drugs; all p < 0.002) and retains 72% signal strength after controlling for five confounds, but was not reproduced in the gCSI panel (different response metric, smaller sample, MEK inhibitors absent), indicating the finding may be MAPK-pathway-specific and assay-dependent. On the 16-drug benchmark, XGBoost achieves the lowest RMSE (1.185); GOPA (1.226) is the strongest neural model. On the full 542-drug panel, target-encoded XGBoost matches GOPA on RMSE (1.322 vs. 1.327); GOPA achieves higher residual Pearson correlation (0.469; 95% CI [0.453, 0.485]) than target-encoded XGBoost (0.391; [0.379, 0.404]); paired difference +0.078 [0.062, 0.093]; GOPA wins on 363 of 539 drugs (67.3%); Wilcoxon p = 1.1 × 10 −23 ). When pathway representations are evaluated with matched downstream learners, simple gene-set projections achieve equivalent prediction, indicating that GOPA’s value lies in its deterministic, population-comparable pathway summaries rather than representational superiority. A controlled geometry comparison shows Poincaré-ball embeddings preserve GO graph distances better than Euclidean ( ρ = 0.732 vs. 0.474) while Euclidean embeddings achieve stronger ancestor retrieval; neither geometry improves prediction (ΔRMSE = +0.008; p = 0.18).