Interpretable variant effect prediction from genomic foundation model embeddings
Discuss this preprint
Start a discussion What are Sciety discussions?Listed in
This article is not in any list yet, why not save it to one of your lists.Abstract
Scientific foundation models learn high-dimensional representations from diverse data modalities, yet what they encode and how to extract that knowledge remain open questions. Here we show that probing the internal representations of Evo 2, a 7-billion-parameter genomic foundation model, enables accurate and interpretable genetic variant effect prediction. We introduce a covariance-based probe that captures second-order structure from Evo 2 sequence embeddings to predict variant pathogenicity across variant types and functional consequences, matching or exceeding specialized predictors within their domains. To ground these predictions in known biological mechanisms, we train a complementary panel of probes on existing annotations to detect which genomic properties are disrupted by a variant. This categorized evidence is then integrated with each variant's genomic context through a language model to generate variant-specific mechanistic hypotheses. Our pathogenicity predictions correlate with experimental measures of variant function, clinical penetrance, and biobank disease associations while the mechanistic hypotheses are consistent with expert reviews, known mechanism classes, and downstream molecular readouts. We release pathogenicity scores, disruption profiles, and contextualized interpretations for 4.2 million variants from the ClinVar database as an open resource through the Evo Variant Effect Explorer (EVEE). More broadly, this structured probing approach offers a general framework for interrogating foundation models across scientific disciplines and grounding their outputs in existing domain concepts.