Efficient exploration of sequence space enables rapid generation of functional genome editors
Discuss this preprint
Start a discussion What are Sciety discussions?Listed in
This article is not in any list yet, why not save it to one of your lists.Abstract
The problem of how protein sequences translate into defined functions remains largely unsolved despite decades of progress. New methods to efficiently explore protein sequence space will help to shed light on these sequence-function relationships, particularly for complex protein function. Here, we describe an approach to create novel, functional proteins through the integration of deep mutational scanning, structural analysis, and evolutionary mining within prompts for a generative protein language model (PLM). We demonstrate the utility of this approach with the generation of novel compact RNA-guided nucleases. This approach is highly efficient, resulting in active nucleases with ∼40% sequence divergence relative to natural proteins and activity equivalent to or exceeding by up to ∼3X that of other compact nucleases at multiple endogenous loci in human cells. The approach described here is rapidly deployable and produces new sequences that will serve as scaffolds for further exploration of complex protein functionality, as well as substrates for novel genome engineering applications.