Tree-aware conditional language modeling recovers mutational patterns of viral evolution
Discuss this preprint
Start a discussion What are Sciety discussions?Listed in
This article is not in any list yet, why not save it to one of your lists.Abstract
The evolutionary accessibility of a protein mutation depends on the sequence background in which it arises and its lineage history, yet most protein language models estimate sequence plausibility without explicitly considering the ordered sequence changes through which descendants arise. We developed evoPLM-Tree, a tree-aware conditional autoregressive language model that predicts descendant protein sequences from ancestral sequences together with phylogenetically derived evolutionary features. We demonstrated our approach using SARS-CoV-2 spike protein, pairing sequences from early Omicron lineages according to their positions on a mutation-annotated phylogeny, and evaluating model performance on sequence pairs from later lineages. Prompt-masking experiments showed that incorporating phylogenetic context substantially increased reliance on the supplied input information compared with a sequence-only model. Generated descendant sequences accurately reproduced the positional distribution of mutations observed during viral evolution, with strong correlations between predicted and observed mutation-frequency profiles for both the receptor-binding domain (Spearman’s ρ = 0.823) and the full spike protein ( ρ = 0.736). Although prediction accuracy for individual substitutions decreased with increasing evolutionary distance, the model consistently captured aggregate mutational patterns across the spike protein. Model-assigned mutation probabilities were also enriched among substitutions experimentally tolerated in deep mutational scanning assays of Omicron BA.2 receptor-binding domain expression (1.19-fold enrichment) and ACE2 binding (1.04-fold enrichment), despite the model being trained solely on observed ancestor–descendant sequence pairs and associated phylogenetic context features. These results demonstrate that explicitly providing protein language models with phylogenetic context during sequence generation can recover lineage-specific mutational patterns and yields probabilistic predictions consistent with experimentally measured functional constraints. evoPLM-Tree provides a framework for modeling protein evolution along phylogenetic lineages and prioritizing plausible future mutations from genomic surveillance data.
Author summary
Viruses evolve by accumulating mutations along branching lineages, so the changes a virus is likely to acquire next depend on its current sequence. Computational models that predict protein sequences usually ignore this history: they judge whether a sequence looks plausible in general, not whether it is a plausible descendant of a particular ancestor. We asked whether giving a language model explicit information about evolutionary history would help it reproduce how a viral protein actually changes. Using the SARS-CoV-2 spike protein, for which millions of real sequences have been arranged into a detailed evolutionary tree, we trained a model on pairs of ancestor and descendant sequences together with simple measurements taken from the tree, such as how many mutational steps separate the two. Our resulting model, evoPLM-Tree, benefitted substantially more on the information it was given than a model trained on sequences alone, reproduced where mutations occur across spike, and assigned higher probability to substitutions that laboratory experiments show the protein tolerates. As surveillance data are captured for other pathogens, the same approach could help anticipate mutations in newly emerging viruses.