基于深度学习的卵巢肿块良恶性分类及泛化性能研究OA
Deep learning-based classification of benign and malignant ovarian masses and generalization performance
目的 评估深度学习模型在典型与随机临床图像上的诊断性能,并将其与不同经验水平的超声医生进行对比,以客观衡量人工智能辅助诊断在真实场景中的应用潜力.方法 回顾性纳入2022年12月~2024年7月福建医科大学附属泉州第一医院(福建省泉州市第一医院)的168例患者共936张超声图像,按7:1:2的比例将其分为训练、验证和测试集.采用6种CNN模型(VGG19_bn、DenseNet121、Swin Transformer-Tiny、ConvNeXt-Tiny、MobileNetV2、ResNet101)评估在医生选取的典型图像上的性能.为检验泛化能力,进一步进行完全随机图像抽样,重复上述实验.将表现最优模型的诊断结果与高、中、初级3位超声医生的诊断效能进行量化比较.结果 在 典 型 图 像 上,VGG19_bn模型性能最佳(AUC=0.906).在随机图像测试中,DenseNet121表现出最强的稳健性(AUC=0.888),其诊断性能全面超越高级超声医生(AUC=0.738),尤其在平衡敏感度(0.786)与特异度(0.857)方面优势明显.医生诊断呈现高敏感度、低特异度的特点,且不同级别医生的敏感度存在波动.结论 经过充分训练的深度学习模型不仅能在受控条件下取得良好的诊断效能,在模拟真实临床复杂情境的随机图像测试中也表现出稳定的泛化能力,其综合诊断指标与超声医师的诊断水平具有可比性.
Objective To evaluate the diagnostic performance of deep learning models on both representative and randomly selected clinical ultrasound images,and to compare their performance with that of radiologists with different levels of experience,in order to objectively assess the potential of artificial intelligence-assisted diagnosis in real-world clinical settings.Methods A total of 936 ultrasound images from 168 patients,retrospectively collected from December 2022 to July 2024 at the Quanzhou First Hospital Affiliated to Fujian Medical University(Quanzhou First Hospital,Fujian),were included in this study.The dataset was divided into training,validation,and test sets in a ratio of 7:1:2.Six convolutional neural network models(VGG19_bn,DenseNet121,Swin Transformer-Tiny,ConvNeXt-Tiny,MobileNetV2,and ResNet101)were employed to evaluate diagnostic performance on representative images selected by radiologists.To further assess model generalization,a fully random image sampling strategy was applied,and the experiments were repeated.The diagnostic performance of the best-performing model was quantitatively compared with that of three radiologists with senior,intermediate,and junior levels of experience.Results On representative images,the VGG19_bn model achieved the best performance(AUC=0.906).In the random image testing scenario,DenseNet121 demonstrated the strongest robustness(AUC=0.888),outperforming the senior ultrasound radiologist(AUC=0.738).Notably,DenseNet121 showed a superior balance between sensitivity(0.786)and specificity(0.857).In contrast,radiologist diagnosis tended to exhibit high sensitivity but relatively low specificity,with variability in sensitivity observed across different experience levels.Conclusion Well-trained deep learning models not only achieve strong diagnostic performance under controlled conditions but also demonstrate stable generalization ability in randomly sampled,clinically realistic scenarios.Their overall diagnostic performance is comparable to that of radiologists,highlighting their potential value in real-world clinical applications.
洪雅婷;阮依丹;李苹;刘卓晟;柳培忠;冯龙翔;吴秀明;蔡诗恬
福建医科大学附属泉州第一医院(福建省泉州市第一医院)妇产科,福建 泉州 362021华侨大学 医学院,福建 泉州 362021福建医科大学附属泉州第一医院(福建省泉州市第一医院)妇产科,福建 泉州 362021华侨大学 医学院,福建 泉州 362021华侨大学 工学院,福建 泉州 362021华侨大学 医学院,福建 泉州 362021福建医科大学附属泉州第一医院(福建省泉州市第一医院)超声科,福建 泉州 362021福建医科大学附属泉州第一医院(福建省泉州市第一医院)妇产科,福建 泉州 362021
卵巢肿块超声图像深度学习良恶性分类卷积神经网络
ovarian massesultrasonographydeep learningbenign-malignant classificationconvolutional neural networks
《分子影像学杂志》 2026 (4)
502-507,6
福建省科技创新联合基金(2024Y9435、2024Y9434)福建医科大学启航基金(2024QH1315)
评论