gVCF2CNV: a scalable pipeline for CNV detection from whole-genome sequencing data

Read the full article See related articles

Discuss this preprint

Start a discussion What are Sciety discussions?

Listed in

This article is not in any list yet, why not save it to one of your lists.
Log in to save this article

Abstract

Motivation

Copy-number variants (CNVs) contribute to human disease and population trait variation. CNV detection from large whole-genome sequencing cohorts remains computationally demanding, as most methods require BAM or CRAM files. Genomic VCF (gVCF) files are smaller, routinely generated by standard variant-calling workflows, and contain the read depth and allelic information needed for CNV detection. However, gVCF files are not directly compatible with established CNV callers that rely on Log R Ratio (LRR) and B Allele Frequency (BAF) signals.

Results

We present gVCF2CNV, a Nextflow pipeline that converts gVCF files into Log R Ratio and B Allele Frequency signals compatible with established CNV callers. Applied to 12,509 individuals from the SPARK cohort, gVCF2CNV generated signals at an average of 2.7 million SNV positions per individual and completed signal extraction in 4 hours using 192 CPUs. CNV calling with PennCNV and QuantiSNP identified candidate CNVs across a broad size range, with trio-based Mendelian precision reaching approximately 80% or higher for deletions of at least 30 kb and duplications of at least 5 kb. Application to 414,824 individuals from the All of Us cohort was completed in 96 hours, demonstrating feasibility at biobank scale. These results show that gVCF files can serve as a scalable input for CNV detection in large WGS cohorts.

Availability and Implementation

gVCF2CNV is available at https://github.com/JacquemontLab/gVCF2CNV , implemented as a Nextflow pipeline with Perl and Python components, supported on Linux.

Contact

mame.seynabou.diop@umontreal.ca .

Article activity feed