首页|期刊导航|作物学报|Haseman-Elston回归与QR分解结合的全基因组选择新方法

Haseman-Elston回归与QR分解结合的全基因组选择新方法OA

A novel genomic selection method combining Haseman-Elston regression and QR decomposition

中文摘要英文摘要

高通量测序与表型技术的快速发展产生了越来越多的大型基因型与表型数据集,而传统的全基因组选择算法难以处理这类数据,计算效率问题已成为全基因组选择的关注重点.为解决该问题,本研究提出一种计算高效的全基因组选择新方法(HQGS),首先利用Haseman-Elston回归估算遗传方差,再对核心群体进行正交三角分解(QR分解)以获得高维矩阵的近似解,这大大提高了计算效率.在模拟数据分析中,设置了不同的核心群体大小、标记个数、真实遗传率、训练群体大小水平用于评估 HQGS、基因组最佳线性无偏预测(GBLUP)、随机森林(RF)和支持向量机(SVM).从预测准确性看,HQGS和GBLUP在绝大部分情况下是相似的,它们全部显著优于RF,大部分情况下优于SVM;从计算效率看,HQGS显著优于GBLUP、RF和SVM,其中,GBLUP的计算效率最低.在真实数据分析中,小麦群体 1 包含的 14 个性状中,HQGS有 3 个性状的预测准确性显著优于GBLUP,8 个性状与GBLUP相似,6 个性状的预测准确性显著优于RF,5 个性状与RF相似,5 个性状的预测准确性显著优于SVM,5 个性状与SVM相似.在小麦群体2包含的4个环境下的产量数据中,HQGS在3个环境下的预测准确性优于GBLUP,在1个环境下优于RF,在2 个环境下与RF相似,在 2 个环境下优于SVM,在 1 个环境下与SVM相似.本研究提出了一种计算高效的全基因组选择新方法,有效避免了大型遗传关系矩阵求逆,在保持预测准确性的情况下显著提高了计算效率,为处理大型数据提供了更加高效可靠的新途径.

The rapid development of high-throughput sequencing and phenotyping technologies has generated increasing large-scale genotype and phenotype datasets.The conventional genomic selection(GS)algorithm struggle to handle such data,and computational efficiency has become a key concern in genomic selection.In order to address this problem,this study pro-poses a novel computationally efficient Haseman-Elston regression+QR decomposition genomic selection(HQGS)method.First,Haseman-Elston regression is used to estimate genetic variance,followed by orthogonal triangular decomposition(QR decompo-sition)of the core population to obtain an approximate solution for high-dimensional matrices,which greatly improves computa-tional efficiency.In simulated data analysis,different levels of core population size,number of markers,true heritability,and training population size were set to evaluate HQGS,genomic best linear unbiased prediction(GBLUP),random forest(RF),and support vector machine(SVM).In terms of prediction accuracy,HQGS and GBLUP were similar in most cases,both significantly outperforming RF in all cases and outperforming SVM in most cases.In terms of computational efficiency,HQGS significantly outperformed GBLUP,RF,and SVM,with GBLUP being the least efficient.In real data analysis,for 14 traits in wheat population 1,HQGS showed significantly better prediction accuracy than GBLUP for 3 traits,was similar to GBLUP for 8 traits,signifi-cantly outperformed RF for 6 traits,was similar to RF for 5 traits,significantly outperformed SVM for 5 traits,and was similar to SVM for 5 traits.For yield data under four environments in wheat population 2,HQGS outperformed GBLUP in three environ-ments,outperformed RF in one environment,was similar to RF in two environments,outperformed SVM in two environments,and was similar to SVM in one environment.This study presents a novel computationally efficient genomic selection method that effectively avoids the inversion of large genetic relationship matrices,significantly improving computational efficiency while maintaining prediction accuracy,providing a more efficient and reliable new approach for handling large datasets.

刘海岚

四川农业大学玉米研究所,四川 成都 611130

全基因组选择矩阵求逆基因组亲缘关系矩阵Haseman-Elston回归QR分解

genomic selectioninverse of matrixgenomic relationship matrixHaseman-Elston regressionQR decomposition

《作物学报》 2026 (8)

2317-2326,10

本研究由国家自然科学基金项目(32271984)资助.This study was supported by the National Natural Science Foundation of China(32271984).

10.3724/SP.J.1006.2026.61006

评论