Interpretable machine Learning Links Wheat 55K SNP Variation to Breeding-trait Prediction and Candidate Locus Prioritization

Read the full article

Listed in

This article is not in any list yet, why not save it to one of your lists.
Log in to save this article

Abstract

Genomic prediction provides a route to accelerate wheat improvement by estimating breeding values directly from dense molecular markers, yet its practical value depends on both predictive reliability and biological interpretability. In many applied breeding populations, sample sizes are modest, phenotyping is limited to a small number of agronomically important traits, and marker effects are difficult to interpret beyond black-box prediction. Here, we developed an interpretable genomic prediction framework for three wheat breeding traits, plant height, plot yield and spike number per mu, using 55K SNP array data and matched field phenotypes. The analysis integrated SNP quality control, population structure assessment, repeated cross-validation, permutation testing, model comparison and candidate SNP prioritization within a unified workflow. After quality control, 2,310 SNPs from 193 genotyped wheat materials were retained, of which 180 materials had matched phenotypic records. Population structure analysis revealed substantial genomic stratification, with PC1 explaining 28.20% of SNP variation and showing a strong association with spike number per mu. Across 20 repeats of five-fold cross-validation, the best-performing models were trait-dependent. ExtraTrees achieved the highest accuracy for plant height and plot yield, with mean Pearson correlations of 0.3996 and 0.4103, respectively, whereas an RKHS-RBF model achieved high predictive accuracy for spike number per mu, with a mean Pearson correlation of 0.8721. Permutation tests confirmed that prediction was significantly better than random expectation for all three traits. Candidate SNP prioritisation identified lead markers on chr2B for plant height and plot yield, and a strong chr4B signal for spike number per mu. The chr4B lead SNP at 32,250,770 bp was located within TraesCS4B02G044500 and lay in a candidate interval that physically overlapped a reported Rht-B1-associated selective sweep region, suggesting a possible relationship with plant architecture, spike-related traits or yield components. These results demonstrate that interpretable machine-learning-based genomic prediction can extract trait-relevant signals from a moderate-sized wheat SNP-chip dataset. The framework provides candidate markers for subsequent KASP assay development, expanded population validation and functional investigation, while also defining the limitations of inference from a single-environment breeding dataset

Article activity feed