A Divide-and-Conquer Approach to Large-Scale Evolutionary Analysis of Single-Cell DNA Data

Yushu Liu
Luay Nakhleh

Read the full article

Discuss this preprint

Start a discussion What are Sciety discussions?

Listed in

This article is not in any list yet, why not save it to one of your lists.

Abstract

Single-cell sequencing technology is producing large datasets, often containing thousands or even tens of thousands of single-cell genomic data points from an individual patient. Evolutionary analyses of these data sets help uncover and order genetic variants in the data as well as elucidate mutation trees and intra-tumor heterogeneity (ITH) in the case of cancer data sets. To enable such large-scale analyses computationally, we propose a divide-and-conquer approach that could be used to scale up computationally intensive inference methods. The approach consists of four steps: 1) partitioning the dataset into subsets, 2) constructing a rooted tree for each subset, 3) computing a representative genotype for each subset by utilizing its inferred tree, and 4) assembling the individual trees using a tree built on the representative genotypes. Besides its flexibility and enabling scalability, this approach also lends itself naturally to ITH analysis, as the clones would be the individual subsets, and the “assembly tree” could be the mutation tree that defines the clones. To demonstrate the effectiveness of our proposed approach, we conducted experiments employing a range of methods at each stage. In particular, as clustering and dimensionality reduction methods are commonly used to tame the complexity of large datasets in this area, we analyzed the performance of a variety of such methods within our approach.

Version published to 10.1101/2024.04.28.591536 on bioRxiv
Apr 30, 2024

Understanding Pathways in Bioinformatics, Genomics, and Health Applications

This article has 1 author:
1. Diptarup Mallick
This article has no evaluationsLatest version Jan 19, 2026
Somatic and germline mutational processes across the tree of life

This article has 11 authors:
1. Peter Campbell
2. Sangjin Lee
3. Yichen Wang
4. Heaton Haynes
5. Emily Mitchell
6. Mark Maddison
7. Liam Crowley
8. Patrick Adkins
9. Nova Mieszkowska
10. Mark Blaxter
11. Richard Durbin
This article has no evaluationsLatest version Jan 19, 2026
Integrated bulk RNA and single-cell RNA sequencing to identify and validate exercise-related genes for predicting the prognosis of invasive ductal carcinoma

This article has 4 authors:
1. YouXin Tang
2. Peng Zhang
3. Yuan Yuan
4. JunXi Gao
This article has no evaluationsLatest version Jan 16, 2026

Discuss this preprint

Listed in

Abstract

Article activity feed

Related articles

Understanding Pathways in Bioinformatics, Genomics, and Health Applications

Somatic and germline mutational processes across the tree of life

Integrated bulk RNA and single-cell RNA sequencing to identify and validate exercise-related genes for predicting the prognosis of invasive ductal carcinoma