Faster model-based estimation of ancestry proportions
This article has been Reviewed by the following groups
Discuss this preprint
Start a discussion What are Sciety discussions?Listed in
- Evaluated articles (Peer Community in Evolutionary Biology)
Abstract
Ancestry estimation from genotype data in unrelated individuals has become an essential tool in population and medical genetics to understand demographic population histories and to model or correct for population structure. The ADMIXTURE software is a widely used model-based approach to account for population stratification, however, it struggles with convergence issues and does not scale to modern human datasets or the large number of variants in whole-genome sequencing data. Likelihood-free approaches optimize a least square objective and have gained popularity in recent years due to their scalability. However, this comes at the cost of accuracy in the ancestry estimates in more complex admixture scenarios. We present a new model-based approach, fastmixture , which adopts aspects from likelihood-free approaches for parameter initialization, followed by a mini-batch expectation-maximization procedure to model the standard likelihood. In a simulation study, we demonstrate that the model-based approaches of fastmixture and ADMIXTURE are significantly more accurate than recent and likelihood-free approaches. We further show that fastmixture runs approximately 30 × faster than ADMIXTURE on both simulated and empirical data from the 1000 Genomes Project such that our model-based approach scales to much larger sample sizes than previously possible.
Article activity feed
-
-
-
Your submission has now been reviewed by three experts in the field. They are all positive about your study but raise important points.
While it is notable that the new implementation is supposidely faster, an assessment of which improvement is most significant would be of interest. More importantly, the comparison against competing methods could involve more complex scenarios to really appreciate the potential novel contribution of fastMixture to the field. The github repository must include clear information on the version control requirements and a toy example to run.
These changes are essential to prove that fastMixture is going to replace Admixture in the future, as stated in your study.
Please ensure that you address all the points raised by the reviewers or justify why those changes are not needed or outside of the scope of …
Your submission has now been reviewed by three experts in the field. They are all positive about your study but raise important points.
While it is notable that the new implementation is supposidely faster, an assessment of which improvement is most significant would be of interest. More importantly, the comparison against competing methods could involve more complex scenarios to really appreciate the potential novel contribution of fastMixture to the field. The github repository must include clear information on the version control requirements and a toy example to run.
These changes are essential to prove that fastMixture is going to replace Admixture in the future, as stated in your study.
Please ensure that you address all the points raised by the reviewers or justify why those changes are not needed or outside of the scope of the study.
-
