Unobserved Sequence Space Has Many Functional Proteins

Read the full article See related articles

Discuss this preprint

Start a discussion What are Sciety discussions?

Listed in

This article is not in any list yet, why not save it to one of your lists.
Log in to save this article

Abstract

The distribution of functional proteins across amino acid sequence space, and the proportion of functional space covered by existing proteins, remains unknown 1,2 . Illuminating this distribution is integral to understanding protein evolution and advancing protein design. The recent explosion of AI/ML protein design tools presents an opportunity to explore protein sequence space distant from extant proteins, but these tools remain poorly validated. Here, we determine that portions of protein sequence space, despite being unobserved in nature, contains many functional proteins that cannot be predicted accurately in silico . We measure experimental fitness of highly diverse proteins across 3 families and assemble the largest known dataset of diverse, functionally labeled natural protein orthologs and new-to-nature proteins. For each family, we observe many functional, new-to-nature sequences with low amino acid identity to existing orthologs. Sequence-based scoring metrics, especially Potts models and protein language models, provide accurate but inconsistent and highly correlated function predictions. Empirical protein fitness landscapes are rugged, and predictions of function do not consistently capture either the local shape or global trends of the empirical fitness landscapes. Finally, we find extensive functional sequence space between existing proteins in each family, providing experimental support for the hypothesis that natural protein sequences explored by evolution represent a miniscule fraction of all possible functional sequences.

Article activity feed