MHChron: diversity-balanced dataset design for robust peptide–MHC binding prediction across MHC class I and II
Discuss this preprint
Start a discussion What are Sciety discussions?Listed in
This article is not in any list yet, why not save it to one of your lists.Abstract
Accurate prediction of peptide–MHC (pMHC) binding is central to immunogenicity assessment, yet many existing predictors are trained and evaluated on narrow allele sets and restricted peptide-lengths. Here, we present MHChron, a unified pMHC binding prediction framework predicated on systematic data curation, meticulous engineering of dataset balance and diversity, and rigorous evaluation through careful splits controlling for data leakage. We assemble one of the most diverse pMHC training dataset reported to date, integrating publicly available binding data across a broad allele coverage (class I n=214, class II n=98) and peptide length range (from 8 to 36 residues). Using a focused and carefully sampled subset of this dataset, we train complementary sequence-based and structure-aware models and test them under increasingly stringent generalisation regimes. Both models achieve consistently strong performance, outperforming the evaluated state-of-the-art predictors despite being trained on numerically fewer data points. Notably, the structure-aware model did not consistently surpass the sequence-based model, except under the most demanding setting of extrapolation to unseen allele clusters, suggesting that performance gains stem primarily from dataset diversity and rigorous evaluation rather than architectural complexity. Sequence-based MHChron is released with reproducible installation and an automated whole-protein screening pipeline, enabling broad and practical use.