Genetic decoding reveals druggable biology implicitly learned by a medical-history foundation model
Discuss this preprint
Start a discussion What are Sciety discussions?Listed in
This article is not in any list yet, why not save it to one of your lists.Abstract
Foundation models trained on electronic healthcare records (EHRs) have gained traction with the aim to transform personalised medicine. However, their interpretability is bound to redescribing the records the models were trained on, missing implicitly learned concepts and biases. Here, we show that human genetics provides an orthogonal layer to surface implicitly learned biological concepts and otherwise hidden risk factors. Re-implementing the generative transformer Delphi-2M in >500,000 UK Biobank participants, we performed genome-wide association testing on its 120 learned embeddings and identified 434 genome-wide-significant signals across 151 independent loci and 98 embeddings, revealing a heritable structure that feature-attribution methods cannot recover. Effector-gene mapping implicated cholesterol metabolism and an IL-1-family epithelial-alarmin pathway, supported by strong (>50-fold) enrichment for variants previously associated with blood lipids, body-mass index, and asthma. Loci recovered the targets of essentially all approved lipid-lowering and severe-asthma therapies, and another twelve drugs not obvious from genetic results based on single ICD-10 GWAS. Yet, embeddings poorly explained variation in pleiotropic risk factors, while still retaining most of their predictive value. Substantial improvements in predictive performance were hence confined to a minority of common diseases by adding specific diagnostic or organ-derived markers. Our findings suggest that human genetics might be most powerful as an orthogonal explanatory or regularising layer to train the next generation of EHR-based foundation models that likely benefit most from the addition of targeted biomarkers to advance personalised medicine.