PMBB Geno-Pheno Toolkit: A suite of scalable, reproducible pipelines for cross-biobank association analyses
Discuss this preprint
Start a discussion What are Sciety discussions?Listed in
This article is not in any list yet, why not save it to one of your lists.Abstract
Summary
Electronic health record (EHR)-linked biobanks generate unprecedented genomic and phenotypic datasets, but their scientific utility is constrained by data fragmentation across institutional silos and incompatible computing infrastructures, forcing researchers to rewrite ad-hoc scripts for each new environment. We present the PMBB Geno-Pheno Toolkit, a suite of modular Nextflow pipelines for biobank-scale association analyses. This note focuses on the toolkit’s SAIGE family of pipelines — supporting genome-wide (GWAS), exome-wide (ExWAS), and phenome-wide (PheWAS) association testing — together with the companion GWAMA and ExWAS meta-analysis pipelines that enable cross-biobank replication. All components are containerized (Docker/Apptainer) and orchestrated with Nextflow, allowing the same workflows to run unmodified on local HPC clusters, cloud platforms, and the All of Us Research Workbench. Complementary toolkit pipelines for PLINK-based GWAS, polygenic scoring, LD-based clumping, and phenotype harmonization are also available and briefly noted.
Availability
The PMBB Geno-Pheno Toolkit is freely available at https://github.com/PMBB-Informatics-and-Genomics/pmbb-geno-pheno-toolkit under MIT open-source license.