Early identification of rare disease using deep phenotyping of electronic health records

Read the full article See related articles

Listed in

This article is not in any list yet, why not save it to one of your lists.
Log in to save this article

Abstract

Patients with rare genetic disorders often wait years for a correct diagnosis, highlighting the urgent need for early and efficient triage for those at risk. Electronic health records (EHRs) are rich in useful information, yet the extent to which this data is sufficient to generate prediagnostic signals is uncertain. We trained machine learning models on billing codes and phenotypes extracted from clinical notes and applied these across eight conditions in a longitudinal cohort of roughly 3 million patient records from the Mayo Clinic, seeking early identification of conditions. The number of patients identified early varied by condition, ranging from 10% to 89% at 99% specificity. Median lead times exceeded one year prior to the first diagnostic billing code for most conditions, including Fabry disease (1,408 days), hereditary angioedema (2,755 days), neurofibromatosis type 1 (1,067 days), and hereditary hemorrhagic telangiectasia (2,320 days). We show that ML models trained on EHR data can prioritize patients for clinician review and confirmatory genetic testing and that the integration of diverse phenotypic data sources provides superior predictive value.

Article activity feed