首页|期刊导航|山西农业大学学报(自然科学版)|基于深度森林的水稻表型预测方法

基于深度森林的水稻表型预测方法OA

Study on phenotypic prediction method of rice based on deep forest

中文摘要英文摘要

[目的]水稻基因组预测有利于精准设计育种,但目前还存在基因组数据高维、非线性、可解释性差等问题,因此,开发出既能保障预测精度与泛化能力、又能兼顾生物学可解释性的新型算法架构可为预测难题提供解决方法.[方法]本研究提出并构建了基于迭代自适应深度森林的水稻表型预测方法(Iterative Adaptive Deep Forest,IADF).该方法由4个部分协同工作的模块组成:首先是单核苷酸多态性(Single Nucleotide Polymorphism,SNP)数据处理模块,对SNP数据质量进行控制,减少噪声;其次是深度森林(Deep Forest)模块,利用其独特的级联森林结构,无需梯度驱动即可有效捕获基因间复杂的非线性互作;然后是迭代特征加权机制(Iterative Feature Reweighting Mechanism,IF-RM)模块,作为智能引导框架,通过学习-反馈-筛选的正反馈循环,模拟探索过程,动态地将模型的计算资源聚焦于贡献度最高的特征子集;最后是自动化超参数优化(Automated Hyperparameter Optimization)模块,利用Optuna框架,对深度森林引擎和迭代特征加权机制的关键参数进行系统性协同寻优,确保整个集成模型的性能达到最佳.[结果]IADF模型在穗角、穗强度、旗叶角度及千粒质量4个关键农艺性状上的皮尔逊相关系数(Pearson Correlation Coef-ficient,PCC)均显著优于LightGBM、XGBoost及DeepCCR等基准模型.IADF的平均预测精度达到0.53,相较于次优模型提升约29%,其中在千粒质量性状上的预测PCC最高达0.70.此外,均方根误差(RMSE)分析进一步证实了该模型在降低预测偏差方面的稳健性.消融实验证实了SNP筛选、迭代特征加权机制及超参数优化模块协同工作的必要性.此外,参数分析显示输入特征数量与模型性能呈非线性关系,在保留3000~3500个关键SNP时模型性能达到最佳.[结论]所提方法的预测精度在提高PCC和降低RMSE方面,相比lightgbm、XGBoost等模型均有不同程度的提升,为实现精准、高效的水稻基因育种提供了重要的技术路径.

[Objective]Rice genome prediction facilitates precision breeding design;however,current approaches face challenges such as high-dimensional genomic data,nonlinearity,and poor interpretability.Therefore,developing a novel algorithmic framework that balances prediction accuracy,generalization capability,and biological interpretability offers a solution to these challenges.[Methods]This study proposes and implements an iterative adaptive deep forest-based rice phenotype prediction method(Iterative Adaptive Deep Forest,IADF).The method consists of four collaboratively functioning modules:a single nucleotide polymorphism(SNP)data processing module that controls SNP data quality and reduces noise.Second,the Deep Forest module utilized its unique cascaded forest structure to effectively capture complex nonlinear gene-gene interactions with-out gradient driving.Third,the Iterative Feature Reweighting Mechanism(IFRM)module served as an intelligent guidance framework,simulating the exploration process through a learning-feedback-screening positive feedback loop,and dynamically focused the model's computational resources on the feature subset with the highest contribution.Finally,the Automated Hy-perparameter Optimization module employed the Optuna framework to systematically and collaboratively optimize the key pa-rameters of the Deep Forest engine and the IFRM,ensuring the optimal performance of the entire integrated model.[Results]The experimental results demonstrated that the IADF model exhibited significantly superior Pearson Correlation Coefficients(PCC)to benchmark models such as LightGBM,XGBoost,and DeepCCR across four key agronomic traits:panicle angle,panicle strength,flag leaf angle,and thousand-grain weight.The IADF model achieved an average prediction accuracy of 0.53,representing an approximate 29%improvement over the second-best model,with the highest prediction PCC reaching 0.70 for the thousand-grain weight trait.Furthermore,root mean square error(RMSE)analysis further validated the model's robustness in reducing prediction bias.Ablation experiments confirmed the necessity of the synergistic collaboration among SNP filtering,the iterative feature weighting mechanism,and the hyperparameter optimization modules.Additionally,parame-ter analysis revealed a nonlinear relationship between the number of input features and model performance,with optimal perfor-mance achieved when retaining 3000-3500 key SNPs.[Conclusion]The prediction accuracy of the proposed method was im-proved to varying degrees in terms of increasing PCC and reducing RMSE compared with LightGBM,XGBoost,and other models,which provides an important technical path for realizing accurate and efficient rice molecular breeding.

朱金圆;刘玉东;胡莉莉;王庆勇

安徽农业大学 信息与人工智能学院,安徽 合肥 230036安徽农业大学 信息与人工智能学院,安徽 合肥 230036安徽农业大学 信息与人工智能学院,安徽 合肥 230036安徽农业大学 信息与人工智能学院,安徽 合肥 230036

信息技术与安全科学

水稻表型预测深度森林迭代特征加权超参数

RicePhenotypic predictionDeep forestIterative feature weightingHyper-parameters

《山西农业大学学报(自然科学版)》 2026 (3)

13-25,13

国家自然科学基金(62301006,32472007)国家重点研发计划项目(2023YFD1802200)安徽省教育厅高校科研项目(2025AH-GXZK40390)

10.13842/j.cnki.issn1671-8151.202512002

评论