首页|期刊导航|Visual Intelligence|Pre-training on high-resolution X-ray images:an experimental study

Pre-training on high-resolution X-ray images:an experimental studyOA

中文摘要

Existing X-ray image based pre-trained vision models are typically trained on a relatively small-scale dataset(less than 500,000 samples)with limited resolution(e.g.,224×224).However,the key to the success of self-supervised pre-training of large models lies in massive training data,and the maintenance of high-resolution X-ray images contributes to effective solutions for some challenging diseases.In this paper,we proposed a high-resolution(1280×1280)X-ray image based pre-trained baseline model on our newly collected large-scale dataset containing more than 1 million X-ray images.Our model employs the masked auto-encoder framework,wherein the tokens that have been processed with a high rate are used as input,and the masked image patches are reconstructed by means of the Transformer encoder-decoder network.More importantly,a novel context-aware masking strategy has been introduced.This strategy utilizes the breast contour as a boundary for adaptive masking operations.We validate the effectiveness of our model through its application in two downstream tasks,namely X-ray report generation and disease detection.Extensive experiments demonstrate that our pre-trained medical baseline model can achieve comparable to,or even exceed,those of current state-of-the-art models on downstream benchmark datasets.

Xiao Wang;Yuehang Li;Wentao Wu;Jiandong Jin;Yao Rong;Bo Jiang;Chuanfu Li;Jin Tang

School of Computer Science and Technology,Anhui University,Hefei 230601,ChinaSchool of Computer Science and Technology,Anhui University,Hefei 230601,ChinaSchool of Artificial Intelligence,Anhui University,Hefei 230601,ChinaSchool of Artificial Intelligence,Anhui University,Hefei 230601,ChinaSchool of Computer Science and Technology,Anhui University,Hefei 230601,ChinaSchool of Computer Science and Technology,Anhui University,Hefei 230601,ChinaFirst Affiliated Hospital of Anhui University of Chinese Medicine,Hefei 230022,ChinaSchool of Computer Science and Technology,Anhui University,Hefei 230601,China

信息技术与安全科学

High-resolution X-ray imagePre-trained big modelsMasked auto-encoder(MAE)Medical report generation

《Visual Intelligence》 2025 (1)

P.98-112,15

supported by the National Natural Science Foundation of China(Nos.62102205 and 62302006)the Natural Science Foundation of Anhui Province(No.2308085QF221).

10.1007/s44267-025-00080-3

评论