Spectronaut-nf: A Nextflow Pipeline for Parallel Processing of DIA Data with Spectronaut
Listed in
This article is not in any list yet, why not save it to one of your lists.Abstract
Summary
Contemporary proteomics methods can now generate large-scale DIA datasets of thousands of files that demand substantial computational resources for efficient analysis. Spectronaut is a widely used platform for DIA data processing; however, large-scale searches are often constrained by computational performance and long execution times when run on single workstations. Here, we present Spectronaut-nf, a Nextflow-based pipeline that enables scalable and parallelized execution of Spectronaut analyses across high-performance computing (HPC) environments. The workflow divides directDIA analysis into modular stages, including spectral library generation, DIA searching, and merging results, allowing efficient distribution of tasks across multiple compute nodes. Benchmarking using 72 diaPASEF raw files using typical hardware demonstrated that Spectronaut-nf completed searches in 23.77 hours, compared with 39.09 hours on a Windows workstation and 67.04 hours on a single-node Linux HPC setup. Stress testing with 1,037 diaPASEF raw files further demonstrated the scalability and robustness of the workflow for large proteomics datasets. Across platforms, protein and peptide identifications remained consistent, with only minimal variability attributable to platform-specific differences. Overall, Spectronaut-nf provides a flexible, scalable, and efficient framework for high-throughput DIA proteomics analysis in HPC environments.