Evaluating Protein Language Model Embeddings for Structural Similarity in the Protein-Sequence Twilight Zone

Read the full article See related articles

Discuss this preprint

Start a discussion What are Sciety discussions?

Listed in

This article is not in any list yet, why not save it to one of your lists.
Log in to save this article

Abstract

Evaluating protein sequence similarity remains challenging in the protein-sequence twilight zone (20–35% sequence identity), where traditional methods often fail. In this study, we evaluate whether mean-pooled embeddings from four protein language models: ESM-1b, ESM-2, ProtT5, and ProstT5 can estimate pairwise structural similarity without performing sequence alignment. The benchmark dataset includes 20,445 PISCES protein pairs with sequence identity ≤30%, representing the protein-sequence twilight zone, with TM-align-derived TM min used as the structural ground truth. Protein embeddings are compared using cosine similarity, Euclidean- and Manhattan-derived similarities, an RBF kernel, and dot product. Among these similarity metrics, cosine similarity performs best across all four models. Moreover, ProstT5 achieves the highest Spearman correlation with TM min , followed by ESM-2, ProtT5, and ESM-1b, while all four PLMs outperform BLASTP overall. Furthermore, the advantage of PLM embeddings is most pronounced for protein pairs with the lowest sequence identity. ProstT5 also provides the best discrimination between structurally similar and dissimilar protein pairs. Moreover, it offers a favorable balance between similarity performance and the computational requirements of residue-level embedding generation and storage. Overall, these findings support PLM embeddings as an effective alignment-free approach for detecting structural relationships among proteins in the twilight zone.

Article activity feed