首页|期刊导航|广东医学|慢性肾脏病患者骨质疏松危险因素与多机器学习临床预测列线图构建

慢性肾脏病患者骨质疏松危险因素与多机器学习临床预测列线图构建OA

Risk factors for osteoporosis in patients with chronic kidney disease and development of a clinical nomogram based on multiple machine learning algorithms

中文摘要英文摘要

目的 基于多种机器学习算法构建慢性肾脏病(CKD)患者继发骨质疏松的临床预测模型,并绘制列线图,为临床提供个体化的风险评估工具.方法 回顾性收集2016年1月至2025年12月于中国人民解放军南部战区总医院收治的242例CKD患者的临床资料,其中69例(28.5%)经双能X线吸收测定法(DXA)诊断为骨质疏松.采用LASSO回归与Boruta算法共同筛选预测变量.分别构建逻辑回归(LR)、决策树(DT)、极限梯度提升(XGBoost)、梯度提升机(GBM)、支持向量机(SVM)、神经网络(NN)及K近邻(KNN)共7种机器学习模型.按7∶3比例将数据分为训练集与测试集,采用十折交叉验证及留一法交叉验证(LOOCV)评估模型性能,并报告乐观校正后的AUC.以LR和XGBoost为主要分析模型,其余作为探索性分析.绘制校准曲线与决策曲线分析(DCA),采用SHAP方法解释最优模型,并基于最佳模型绘制列线图.结果 LASSO与Boruta交集筛选出4个核心变量:体重、丙氨酸氨基转移酶(ALT)、碱性磷酸酶(ALP)、低密度脂蛋白胆固醇(LDL).测试集结果显示,LR的AUC为0.790 4,准确率0.726,敏感度0.842,特异度0.685;经 LOOCV 乐观校正后 AUC 为 0.804(95%CI:0.755~0.853).XGBoost 的测试集 AUC 为 0.700 3,校正后AUC为0.751.其余模型(DT、GBM、SVM、NN、KNN)存在不同程度的过拟合或性能偏低.校准分析显示LR的预测概率存在系统性压缩(校准斜率0.32),绝对概率尚不能直接用于临床决策,但模型具有良好的风险区分能力.DCA提示LR在一定阈值范围内具有净获益.基于LR构建的列线图(纳入体重、ALT、ALP、LDL)的AUC与完整模型一致(0.792),显著提升了临床可操作性.SHAP分析显示体重、LDL、ALP是主要的正向预测因子.结论 基于体重、ALT、ALP和LDL构建的逻辑回归模型及列线图对CKD患者骨质疏松具有较好的区分能力,且变量简易、可解释性强,适用于基层医院的风险分层.但由于样本量有限及校准存在偏差,模型输出的绝对概率需谨慎解读,未来需开展多中心前瞻性研究进一步验证.

Objective To develop and validate a clinical prediction model for secondary osteoporosis in patients with chronic kidney disease(CKD)using multiple machine learning algorithms and to construct a nomogram for individu-alized risk assessment.Methods A retrospective analysis was performed on the clinical data of 242 patients with CKD admitted to the General Hospital of Southern Theater Command of the Chinese People's Liberation Army between January 2016 and December 2025.Among them,69 patients(28.5%)were diagnosed with osteoporosis by dual-energy X-ray absorptiometry(DXA).Candidate predictors were jointly selected using least absolute shrinkage and selection operator(LASSO)regression and the Boruta algorithm.Seven machine learning models were developed,including logistic regres-sion(LR),decision tree(DT),extreme gradient boosting(XGBoost),gradient boosting machine(GBM),support vec-tor machine(SVM),neural network(NN),and k-nearest neighbor(KNN).The dataset was randomly divided into training and test sets at a ratio of 7∶3.Model performance was evaluated using 10-fold cross-validation and leave-one-out cross-validation(LOOCV),with optimism-corrected area under the receiver operating characteristic curve(AUC)reported.LR and XGBoost were predefined as the primary analytical models,whereas the remaining algorithms were con-sidered exploratory.Calibration curves and decision curve analysis(DCA)were used to assess model calibration and clini-cal utility.SHapley Additive exPlanations(SHAP)were applied to interpret the optimal model,and a nomogram was sub-sequently established.Results Four key predictors were consistently identified by both LASSO regression and the Boruta algorithm:body weight,alanine aminotransferase(ALT),alkaline phosphatase(ALP),and low-density lipoprotein cholesterol(LDL-C).In the test sets cohort,the LR model achieved an AUC of 0.790 4,with an accuracy of 0.726,sensitivity of 0.842,and specificity of 0.685.After optimism correction using LOOCV,the AUC increased to 0.804(95%CI:0.755-0.853).The XGBoost model yielded a validation AUC of 0.700 3,with an optimism-corrected AUC of 0.751.The remaining models(DT,GBM,SVM,NN,and KNN)demonstrated varying degrees of overfitting or inferi-or predictive performance.Calibration analysis indicated systematic compression of predicted probabilities in the LR model(calibration slope=0.32),suggesting that although the model exhibited satisfactory discrimination,its absolute risk esti-mates should be interpreted with caution.DCA demonstrated a favorable net clinical benefit of the LR model across a clini-cally relevant range of threshold probabilities.A nomogram incorporating body weight,ALT,ALP,and LDL-C was estab-lished based on the LR model and achieved an AUC of 0.792,comparable to that of the complete model while substantially improving clinical applicability.SHAP analysis identified body weight,LDL-C,and ALP as the most influential positive predictors.Conclusion The logistic regression model and nomogram incorporating body weight,ALT,ALP,and LDL-Cdemonstrated good discriminative ability for predicting osteoporosis in patients with CKD.Owing to its simplicity and in-terpretability,the model may serve as a practical tool for risk stratification,particularly in primary healthcare settings.Nevertheless,given the limited sample size and imperfect calibration,the predicted absolute probabilities should be inter-preted cautiously.Further prospective multicenter studies are warranted to externally validate and refine the model.

杨柳;朱凯旋;李佳

中国人民解放军南部战区总医院内分泌科(广东 广州 510010)中国人民解放军南部战区总医院内分泌科(广东 广州 510010)中国人民解放军南部战区总医院内分泌科(广东 广州 510010)

医药卫生

慢性肾脏病骨质疏松机器学习预测模型神经网络列线图

chronic kidney diseaseosteoporosismachine learningprediction modelneural networknomo-gram

《广东医学》 2026 (7)

1003-1014,12

国家重点研发计划项目(2021YFC2501701)国家卫生健康委能力建设和继续教育中心2025年度慢病管理研究课题(GWJJMB202510024006)

10.13820/j.cnki.gdyx.20261276

评论