ProcessNets: Towards an efficient approach for ensemble analysis of biological networks

Read the full article See related articles

Listed in

This article is not in any list yet, why not save it to one of your lists.
Log in to save this article

Abstract

Integrative analysis of the properties of multiple networks pertaining to a biological system is a key problem in systems biology. But computing the properties of thousands of large networks poses a challenge. Most techniques that partially address this challenge, such as parallelism and optimized algo-rithms, treat the analysis of each network as an independent task. However, the networks arising from the same biological system are often similar to each other, and can we harness this similarity to achieve speedup? Towards this end, we propose a phylogeny-like data structure, nc -tree, to compactly represent an ensemble of related biological networks; and a versatile frame-work, P rocess N ets , that runs incremental algorithms over the nc -tree to speedup property computations on the input networks. Our asymptotic analysis of space and time complexity, coupled with empirical evaluations of diverse ensembles of simulated and real-world GTEx gene coexpression networks, demonstrate that P rocess N ets leads to significant space and time advantages when computing different network science measures on ensembles of similar networks. For instance, we achieve a compression factor of 2.1 × and 8.1 × in storing two representative ensembles of GTEx Muscle Skeletal networks with 1000 nodes/genes, and a time speedup of 2.86 × to 3.97 × over the fastest baseline for computing degree centrality in ensembles of networks with at least 14, 500 nodes derived from subsampled GTEx datasets of ten tissues. These results are promising and encourage application of P rocess N ets to analyze biological network ensembles derived from rapidly accumulating consortium/biobank-based datasets.

Article activity feed