De novo design of ligand binding proteins using large language models alone

This article has been Reviewed by the following groups

Read the full article

Listed in

Log in to save this article

Abstract

Protein design has rapidly advanced with the advent of sequence- and structure-based machine learning models. However, reasoned design, which applies physicochemical principles and rules derived from sequence-structure-function relationships, has not seen the same benefits from generative machine learning models. Here, we test the ability of common large language models (LLMs; e.g. Claude, ChatGPT, and Gemini) to consider design principles to generate de novo proteins that bind metals and lipophilic small molecules without copying existing sequences. Common LLMs alone are able to 1) generate protein sequences to adopt a desired fold and bind the target ligand and 2) explain the principles that motivate the design choices. Following structure prediction and filtering, we selected a small set of designs (6 to 12 designs per query) for experimental validation, affording metal binders in one round of LLM-based design (25% hit rate) and perfluorooctanoic acid binders in two rounds (25% hit rate in the second round of design). Importantly, the LLMs produce detailed justification to accompany the de novo designed sequences, providing a conceptual framework on which designs can be evaluated. While the successful designs have some deviations from the prompted parameters and LLM-articulated design rationale, these campaigns provide a case study that highlights the utility of LLMs in making protein design more comprehensible and accessible to users without sophisticated design expertise.

Article activity feed

  1. Be explicit about selection cutoffs, and consider testing designs that fall below them. You note that the confidence scores (average Chai-1 pTM of 0.49 and ipTM of 0.25 for Zn²⁺ binders) sit well below the >0.9 cutoffs typical of design campaigns. The 12 LMBPs were chosen by "self-consistency and manual inspection," but the paper doesn't say what RMSD, pTM, or ipTM values the selected designs have. I would mark the selected designs in Fig. 2B,C and report their scores in a table, or state the exact cutoffs. Also, I would have ordered some designs below conventional thresholds. You already advanced low-scoring sequences, and these are only in silico scores. Testing whether they predict expression, folding, and binding in vitro would be useful data for zero-shot design, and it would show whether filtering helped at all.

  2. Fig. 2A says "chatGPT-o3" while the text says ChatGPT-5. The Fig. 3 legend writes "LBP-hex" where the text writes "LPB-hex," and the text writes "LMPB. Make sure the labeling is consistent.

  3. Thanks for this — it's a clear, honest case study, and we appreciate that you reported the designs that didn't match the prompt. A few suggestions. I would give the four evaluation criteria a main figure. You assess designs on sequence novelty, designability, homodimer propensity, and metal binding capacity. While confidence is plotte din the main figure, selected designes are not highlighted in those plots. Because the paper's central claim is that LLMs produce novel, designable sequences, I'd like to see how each criterion was quantified in one place. A figure comparing the models could include: the length distribution for each LLM (mean and spread) amino acid composition by model secondary structure content by model BLASTp and PDB similarity statistics that support the novelty claim This would back up the observations in the text, such as Gemini favoring charged, repetitive sequences and ChatGPT favoring Asn, Thr, and Ser, with numbers.