首页|期刊导航|肿瘤预防与治疗|多机器学习模型在慢性肝病和肝癌预测中的效能比较及最优模型筛选

多机器学习模型在慢性肝病和肝癌预测中的效能比较及最优模型筛选OA

Comparison of the Efficacy of Multiple Machine Learning Models for Pre-diction of Chronic Liver Disease and Hepatocellular Carcinoma and Screening of the Optimal Model

中文摘要英文摘要

目的:基于血清学测定、影像学检测和肝活检的慢性肝病和肝癌诊断方法具有灵敏度低、有创且易受主观因素影响等特点.因此迫切需要寻求无创、准确性好、灵敏度高的慢性肝病和肝癌检测方法.本研究旨在探索基于机器学习的无创性慢性肝病和肝癌预测的最优模型.方法:研究选取 2013~2023 年新疆医科大学第一附属医院的24 215 例患者和 50 个特征数据,经 autoencoder、matrix completion、聚类内部均值补全、missForest 四种方法补全,将单变量统计检验、随机森林、特征递归消除、相关性分析筛选出的 36 个关键特征作为随机森林、逻辑回归、XGBoost 三种模型的输入,通过准确率、召回率、ROC 曲线及 AUC 等指标评估模型性能,交叉验证评估模型泛化能力,最后利用沙普利加和解释(SHapley Additive exPlanation,SHAP)对重要特征进行解释分析.结果:Matrix completion 补全方法结合 XGBoost 模型表现最佳,准确率为 0.769,其中脂肪肝、肝炎、肝硬化、肝癌的 AUC 分别为 0.93、0.87、0.94 和0.92,并且 SHAP 分析结果显示血小板计数、血清胆碱酯酶、乙肝表面抗原定量等是预测的关键特征.结论:基于Matrix completion 补全方法的 XGBoost 机器学习模型是预测慢性肝病的最优模型,利用 SHAP 模型增强该模型的可解释性,能够识别出慢性肝病和肝癌的重要特征,为早诊提供参考.

Objective:The diagnostic methods for chronic liver disease and hepatocellular carcinoma based on serological assays,imaging tests and liver biopsies are characterized by low sensitivity,invasiveness and susceptibility to subjective fac-tors.Thus,there is an urgent need to develop a non-invasive,accurate,and highly sensitive detection method for chronic liver disease and hepatocellular carcinoma.The purpose of this study is to explore the optimal machine learning-based model for non-invasive prediction of chronic liver disease and hepatocellular carcinoma.Methods:A total of 24,215 patients with chronic liver disease and 50 feature variables were en-rolled from the First Affiliated Hospital of Xinjiang Medical University between 2013 and 2023.Four methods,including au-toencoder,matrix completion,cluster internal mean imputation and missForest,were adopted for missing data imputation.Subsequently,36 key features screened by univariate statistical test,random forest,recursive feature elimination and correla-tion analysis were taken as the input of three models:random forest,logistic regression and XGBoost.Model performance was evaluated by accuracy,recall,ROC curves and AUC values.Cross-validation was used to assess the generalization abili-ty of the models.Finally,SHapley Additive exPlanation(SHAP)was applied to interpret and analyze the important fea-tures.Results:The combination of matrix completion imputation method and XGBoost model achieved the optimal perform-ance,with an accuracy of 0.769.Specifically,the AUC values for fatty liver,hepatitis,liver cirrhosis,and liver cancer were 0.93,0.87,0.94,and 0.92,respectively.Furthermore,SHAP analysis revealed that platelet count,serum cholines-terase,and quantitative hepatitis B surface antigen were among the key features for predicting chronic liver disease.Conclu-sion:The XGBoost machine learning model based on the matrix completion imputation method is the optimal model for pre-dicting chronic liver disease.By enhancing the model's interpretability with SHAP,the model can identify critical features related to chronic liver disease,providing reference for early diagnosis.

张韬;买热比娅·马合木提;高颖

830000 乌鲁木齐,新疆医科大学第一附属医院 感染病·肝病中心二科830000 乌鲁木齐,新疆医科大学第一附属医院 感染病·肝病中心二科830000 乌鲁木齐,新疆医科大学第一附属医院 感染病·肝病中心二科

医药卫生

慢性肝病肝癌XGBoost模型预测SHAP分析

Chronic liver diseaseHepatocellular carcinomaXGBoostModel predictionSHAP analysis

《肿瘤预防与治疗》 2026 (5)

349-358,10

This study was supported by grants from Science and Technology Committee of Xinjiang Uygur Autonomous Region(No.2022E02115). 新疆维吾尔自治区区域协同创新专项-科技援疆计划(编号:2022E02115)

10.3969/j.issn.1674-0904.2026.05.002

评论