HANSEN: An Integrated Structural and Functional Proteome Resource for Structure-Guided Drug Discovery in Mycobacterium leprae
Discuss this preprint
Start a discussion What are Sciety discussions?Listed in
This article is not in any list yet, why not save it to one of your lists.Abstract
Structural biology has advanced antimicrobial discovery by enabling drug-target identification and validation and supporting structure-guided inhibitor design. However, Mycobacterium leprae ( M. leprae ), the obligate intracellular bacillus that causes leprosy (Hansen’s disease), remains structurally under-characterised. Only 10 Protein Data Bank (PDB) entries represent seven unique proteins within a proteome encoded by 1,603 protein-coding genes. To address this gap, we present HANSEN ( https://hansen-leprosy.medschl.cam.ac.uk/home ), an integrated structural and functional resource containing computationally predicted three-dimensional models across the M. leprae proteome. Monomeric and oligomeric models were generated using complementary structure-prediction methods, including AlphaFold 3, Boltz-1, Chai-1, and Boltz-2. Models were annotated with predicted Local Distance Difference Test (pLDDT) scores and predicted aligned error (PAE) values. Ligand-binding pockets were predicted using AF2BIND, P2Rank, and fpocket, and ligands from the best-matching PDB templates were modelled within oligomeric complexes. Residue-level B-cell epitope propensity was estimated using DiscoTope-3.0, and ProteomeLM-derived essentiality scores were calculated for each protein. These features were integrated into a relational web database with interactive visualisation through Mol*. We also ranked all 1,603 proteins using a Target Priority Score ranging from 0 to 100. The score combines ProteomeLM-derived essentiality with pocket and AF2BIND predictions, functional annotations, and Boltz-2-associated measures of model quality and tractability. The essentiality model used a logistic-regression head trained on Mycobacterium tuberculosis (M. tuberculosis) Tn-seq labels. It achieved an AUROC of 0.84 in homology-grouped M. tuberculosis cross-validation and a transfer AUROC of 0.78 against the orthologue-aligned M. leprae reference set. Proteins were assigned to four tiers, ranging from high priority to exploratory candidates. Together, HANSEN provides a practical resource for generating and prioritising experimentally testable hypotheses for M. leprae target discovery and structure-guided drug development.
Teaser
Predicted structures and druggability annotations for the whole M. leprae proteome in an open resource.
Key points
-
HANSEN provides a proteome-wide structural resource for M. leprae , integrating predicted monomeric and oligomeric models with confidence metrics and functional annotations.
-
HANSEN integrates UniProt annotations, ligand and cofactor associations, structure-based predictions of small-molecule binding pockets, B-cell epitope propensity, gene-essentiality estimates and multi-parameter target prioritisation within a single, protein-centric interface for the proteome of M. leprae .
-
The resource further incorporates a dedicated analytical module for Oxford Nanopore MinION amplicon-sequencing data, enabling the identification of mutations within drug-resistance-determining regions that confer antimicrobial resistance (AMR) in M. leprae .
-
Benchmarking against available experimental structures, together with cross-method concordance analyses, supports the use of pLDDT, PAE and agreement between prediction methods as complementary indicators of model reliability.
-
Integrated target prioritisation produced a ranked set of candidate proteins, including established mycobacterial drug targets, to support experimental hypothesis generation for leprosy drug discovery.