Optimizing genomic selection: A comparison of SNP selection strategies for reduced-density panels in beef cattle
Listed in
This article is not in any list yet, why not save it to one of your lists.Abstract
The exponential increase in the number of genotyped animals, combined with the availability of high-density SNP chips has introduced computational challenges for routine genomic evaluations, particularly during the construction of the genomic relationship matrix. Although higher-density SNP panels can facilitate the identification of causal mutations, their use substantially increases computational requirements without a proportional gain in genomic prediction performance. To optimize computational efficiency while maintaining accuracy of genomic predictions, this study compared five SNP selection strategies (i.e., random sampling, random sampling with inclusion of informative SNPs, linkage disequilibrium (LD)-based pruning, a Shannon entropy–based machine learning approach, and F ST -based prioritization) to develop reduced-density panels for Nellore cattle. Using high-density (HD) genotype data comprising 437,650 SNPs from 304,782 animals (after quality control) as reference, three reduced-density panels (25K, 45K, and 65K SNPs) panels were tested across five traits (i.e., Age at first calving, Stayability, Weaning weight, Yearling weight, Muscling) with diverse genetic architectures. Genomic estimated breeding values (GEBVs) derived from these reduced panels were compared to those obtained from the HD reference panel using Pearson’s correlations, under both genomic best linear unbiased prediction (GBLUP) and single-step GBLUP (ssGBLUP) methods. In the GBLUP model, prediction accuracy generally improved with increased marker density. Random selection with and without the informative SNPs consistently yielded the highest accuracies, whereas the F ST -based approach showed the lowest agreement with the HD reference across all densities. In contrast, ssGBLUP demonstrated strong robustness to marker reduction, producing uniformly high correlations (≈1.00) across all SNP densities and selection strategies. These findings indicate that optimized low-density SNP panels maintain prediction accuracy comparable to HD panels, offering a cost-effective tool for large-scale genomic evaluations.
Author Summary
Genomic selection has transformed cattle breeding by allowing producers to identify animals with superior genetic potential using DNA information. However, modern genomic evaluations often rely on very large genetic datasets that require substantial computing power and increase genotyping costs, particularly in large breeding populations such as Nellore cattle in Brazil. In this study, we evaluated whether reduced-density marker panels could maintain the same level of prediction accuracy as high-density panels commonly used in genomic evaluations. We compared five different strategies for selecting informative genetic markers and tested panels containing different numbers of markers across economically important traits. We found that reduced-density panels, particularly those developed using random or linkage-based selection methods, produced genomic predictions highly similar to those obtained with high-density panels. In addition, prediction methods that combined genomic and pedigree information remained highly robust even with fewer markers. Our findings suggest that reduced-density panels can support accurate and cost-effective genomic evaluations, allowing breeding programs to evaluate more animals more frequently while reducing computational demands.