A multi-view similarity network fusion framework for syndrome discovery from aggregated health records
Discuss this preprint
Start a discussion What are Sciety discussions?Listed in
This article is not in any list yet, why not save it to one of your lists.Abstract
Syndrome discovery, the identification of clinically meaningful groupings of signs and symptoms, is a foundational but labor-intensive task in syndromic surveillance, and the COVID-19 pandemic exposed the rigidity of expert-curated definitions in the face of novel threats. Unsupervised, data-driven methods are well-suited to this problem but remain underused. We propose an unsupervised framework based on Similarity Network Fusion (SNF) that operates on only five variables: diagnosis code, sex, age group, epidemiological week and year, and encounter count. Each diagnosis code was represented through three complementary views corresponding to the fundamental questions of syndromic surveillance: what condition is recorded (clinical, via SapBERT embeddings), who is affected (demographic, via chi-square distances), and when it occurs (temporal, via Move-Split-Merge). The fused affinity matrix is partitioned by spectral clustering and exported directly in the Open Syndrome Definition (OSD) format for downstream integration. To our knowledge, this is the first framework of SNF applied to the task of syndrome discovery. We use the framework in 72.9 million primary care encounters across ten Brazilian municipalities. Treating each city as an experiment with no shared training signal yields 47 candidate syndromes, 72% of which are rated fully valid by an expert blind to the procedure. By requiring no predefined targets, the framework discovers candidate syndromes at scale, including ones never explicitly sought, and emits them in a deployable format, shortening the path from emerging signal to usable definition.
Author summary
Public health systems are under growing pressure to catch outbreaks early, but the syndrome definitions that trigger surveillance alerts, the combinations of signs and symptoms that flag a possible threat, are still built largely by hand and too slowly to keep up. We developed a data-driven framework that groups diagnosis codes into candidate syndromes by asking three questions: what condition a code represents, who is affected, and when it occurs. Applied to 72.9 million primary care encounters across ten Brazilian cities, the framework recovered known syndromes, including arboviruses and influenza-like illness, without being given their definitions, and produced 47 candidate syndromes that held across all cities, 72% of which a domain expert rated valid. Beyond infectious disease, it surfaced coherent groupings in areas that surveillance rarely monitors, such as mental health. Each candidate is exported in the Open Syndrome format, ready for existing surveillance systems and the wider community to use.