基于可解释性的Transformer模型在早期尘肺病辅助诊断中的临床价值评估OA
Evaluation of clinical value of interpretable Transformer-based models in computer-aided di-agnosis of early-stage pneumoconiosis
[背景]早期尘肺病诊断高度依赖医师经验,存在主观性强、一致性低等问题.Vision Trans-former(ViT)架构在早期尘肺病识别中的临床价值与可解释性尚需系统验证. [目的]系统评估基于 ViT的深度学习模型在早期尘肺病 DR胸片识别中的诊断效能,并利用梯度加权类激活映射(Grad-CAM)算法增强模型决策过程的透明度与可解释性,为开发智能辅助诊断工具提供可靠的依据. [方法]本研究回顾性收集了 1 326例接尘职工的胸部高千伏 X线胸片.首先采用 U-Net模型自动分割肺野区域,该模型训练中采用了结合 Dice损失与交叉熵损失的混合损失函数及数据增强技术.然后利用分割后的肺野感兴趣区域(ROI),训练了四种 ViT衍生模型(ViT、Sim-pleViT、TwinsSVT、CrossFormer),并比较其对早期尘肺病的二分类(正常 vs.异常)诊断效能.以三名资深放射医师的共同诊断结果为金标准,数据集按 7∶3比例随机划分为训练集(928例)与验证集(398例).所有分类模型均采用 Adam优化器,训练 3 000轮,初始学习率为0.001,并采用余弦衰减策略动态调整学习率.模型效能的比较采用 DeLong检验,诊断一致性通过Kappa系数评估,并采用Grad-CAM算法对最优模型进行特征可视化分析. [结果]U-Net肺野自动分割模型在验证集上性能优异,其 Dice系数为 0.971,平均交并比(mIOU)为 0.944.在 分 类 任 务 中,CrossFormer模型性能最优,其 AUC值为 0.962(95%CI:0.945~0.979),准确率、灵敏度、特异度和 F1分数分别为 92.0%、92.1%、91.8%和 0.916,诊断性能优于其他对比模型(DeLong检验,均 P<0.01).CrossFormer模型的诊断性能与资深放射医师相当(Kappa=0.834,P<0.001),其准确率(91.7%)与医师 A(92.5%)和医师 B(92.2%)处于同一水平.McNemar检验表明诊断错误无系统性偏差(P>0.05).Grad-CAM可视化分析显示,模型的注意力区域与尘肺病的典型影像分布高度一致,主要集中于双肺中上野及肺门周围,与小阴影的实际分布区域高度一致. [结论]CrossFormer模型在早期尘肺病识别中展现出优异的诊断效能,结合 Grad-CAM的可解释性分析,证实了模型决策的合理性,为开发辅助医师诊断尘肺病的智能工具提供了可靠的模型选择依据与技术方案参考.
[Background]The diagnosis of early-stage pneumoconiosis relies heavily on physician experi-ence,leading to substantial subjectivity and interobserver variability.The clinical value and inter-pretability of the Vision Transformer(ViT)architecture for the early identification of pneumoco-niosis require systematic validation. [Objective]To systematically evaluate the diagnostic performance of ViT-based deep learning models for identifying early-stage pneumoconiosis on digital radiography(DR)chest radiographs,and to enhance the transparency of model decision-making process using the gradient-weighted class activation mapping(Grad-CAM)algorithm,thereby providing a reliable basis for developing computer-aided diagnosis tools. [Methods]This study retrospectively collected high-kilovolt chest radiographs from 1 326 dust-exposed workers.The lung fields were au-tomatically segmented using a U-Net model trained with a hybrid loss function combining Dice loss and cross-entropy loss,together with data augmentation techniques.Subsequently,using the segmented lung field regions of interest(ROIs),four ViT-derived models(ViT,SimpleViT,TwinsSVT,and CrossFormer)were trained,and their diagnostic efficacies for the binary classification(normal vs.abnormal)of early-stage pneumoconiosis were compared.The consensus diagnosis of three senior radiologists served as the reference standard.The dataset was randomly divided into a training set(928 cases)and a validation set(398 cases)at a 7∶3 ratio.All classification models were trained using the Adam optimizer for 3 000 epochs with an initial learning rate of 0.001,and a cosine decay strategy was employed to dy-namically adjust the learning rate.Model performance was compared using DeLong's test,diagnostic consistency was assessed via the Kappa coefficient,and the optimal model was further evaluated by feature visualization analysis using Grad-CAM. [Results]The U-Net lung field segmentation model demonstrated excellent performance on the validation set,achieving a Dice coefficient of 0.971 and a mean intersection over union(mIOU)of 0.944.In the classification task,the CrossFormer model performed best,yielding an area under the curve(AUC)of 0.962(95%CI:0.945,0.979),with accuracy,sensitivity,specificity,and F1-score of 92.0%,92.1%,91.8%,and 0.916,respectively.Its diagnostic performance was superior to the other evaluated models(DeLong's test,P<0.01).The diagnostic performance of the CrossFormer model was highly comparable to that of senior radiologists(Kappa=0.834,P<0.001),with its accuracy(91.7%)being on par with Physician A(92.5%)and Physician B(92.2%).McNemar's test indicated no systematic bias in diagnostic errors(P>0.05).Grad-CAM visualization revealed that the model's attention maps aligned closely with the typical radiological distribution of pneumoconiosis,mainly involving the upper and middle lung zones and perihilar regions,corresponding to the actual distribution of small opacities. [Conclusion]The CrossFormer model demonstrates excellent diagnostic performance for the early identification of pneumoconiosis.Combined with Grad-CAM-based interpretability analysis,the model shows clinically plausible decision-making patterns,providing a reliable model selection basis and technical reference for the development of computer-assisted diagnostic tools for pneumoconiosis.
何岭;罗燕;雷钧艳;尹颀;杨德明;张长彪;王村建
重庆市涪陵区疾病预防控制中心职业病防治科,重庆 408000重庆市涪陵区疾病预防控制中心职业病防治科,重庆 408000重庆市涪陵区疾病预防控制中心职业病防治科,重庆 408000重庆市涪陵区疾病预防控制中心职业病防治科,重庆 408000重庆市涪陵区疾病预防控制中心职业病防治科,重庆 408000重庆市涪陵区人民医院放射影像科(介入治疗室),重庆 408000重庆市涪陵区疾病预防控制中心职业病防治科,重庆 408000
医药卫生
尘肺病Vision Transformer深度学习计算机辅助诊断胸部X线摄影
pneumoconiosisVision Transformerdeep learningcomputer-aided diagnosischest radiography
《环境与职业医学》 2026 (7)
844-851,8
2024年涪陵区科卫联合医学科研项目(2024KWLH023)
评论