Empirical substitution models of SARS-CoV-2 protein evolution for phylogenetic inference

Read the full article

Discuss this preprint

Start a discussion What are Sciety discussions?

Listed in

This article is not in any list yet, why not save it to one of your lists.
Log in to save this article

Abstract

Probabilistic phylogenetic inference requires substitution models of molecular evolution, which should be as realistic as possible to yield accurate predictions. At the protein level, empirical substitution models are well established in phylogenetics. However, the number of available empirical substitution models remains limited, and only a few focus on rapidly evolving viruses. Indeed, no substitution models have yet been developed specifically for SARS-CoV-2 proteins, despite their potential to improve evolutionary predictions. Here, we present a set of empirical substitution models for SARS-CoV-2 proteins, including the main protease, papain-like protease, spike protein, and the complete proteome. We estimated these models using maximum likelihood from thousands of protein sequences and validated with independent datasets. Overall, the estimated empirical substitution models outperformed currently available empirical models in terms of phylogenetic likelihood and revealed distinct evolutionary patterns among SARS-CoV-2 proteins. Additionally, using forward-in-time evolutionary simulations, we found that SARS-CoV-2 proteins evolved under these empirical substitution models generated protein variants with folding stabilities more realistic than those produced under other empirical substitution models. We conclude that evolutionary inferences based on protein-specific substitution models are generally more accurate and biologically realistic than those based on generalist models, and recommend the empirical substitution models presented here for evolutionary predictions of SARS-CoV-2 proteins.

Article activity feed