How are evolutionarily young and old proteins distributed in sequence space?
Discuss this preprint
Start a discussion What are Sciety discussions?Listed in
This article is not in any list yet, why not save it to one of your lists.Abstract
Protein sequence space is vast due to the combinatorial diversity of 20 amino acids. However, evolution has generated a limited set of “old” canonical protein families sharing evolutionary ancestry, structures and functions. It remains unclear how canonical sequences are placed in sequence space, how recently evolved “young” proteins compare to them, and whether random, young, and canonical sequences can interconvert along evolutionarily plausible paths, and which biophysical properties distinguish or link these sequences. Here, we analyse naturally occurring de novo proteins from yeast and flies, which originate from non-coding DNA and thus have experienced limited evolutionary selection. They serve as a model for examining the relationships between young de novo and intergenic proteins, older canonical proteins, and their randomized counterparts. Because de novo and randomized sequences lack detectable homology, we use an alignment-free k-mer-based distance approach. Randomization shifts distance distributions toward expected random behaviour in all classes, but natural, non-randomized sequence classes remain distinct, indicating non-random residue organization. Each class exhibits characteristic k-mer patterns, with de novo proteins clearly separated from both canonical and all randomized sequences. Sequences bridging these classes are frequently predicted to contain transmembrane helices. De novo proteins are thus not random samples of sequence space. Instead, they occupy constrained yet evolutionarily accessible regions defined by residue order and biophysical constraints, suggesting a plausible pathway for the emergence and diversification of new proteins.
Significance Statement
Despite the vast combinatorial potential of amino acids, evolution has produced only a limited repertoire of canonical proteins with conserved structure and function. How evolutionarily young proteins relate to older canonical proteins, and whether the sequence space between them is traversable, remain unclear. Here, we decompose canonical proteins, intergenic sequences, and recently emerged yeast and fly de novo proteins, together with randomized controls, into short, interpretable fragments (k-mers) and compare them using alignment-free distances. De novo proteins are markedly distinct from both randomized and canonical sequences. Notwithstanding their evolutionary distance, sequences are connected by stepwise paths comprising bridge sequences, often enriched for low-complexity motifs and transmembrane helices, connecting disordered and structured regions of sequence space.