首页|期刊导航|Intelligent Oncology|Decision-making performance of large language models vs.human physicians in challenging lung cancer cases:A real-world case-based study

Decision-making performance of large language models vs.human physicians in challenging lung cancer cases:A real-world case-based studyOA

中文摘要

Background:Despite the promise shown by large language models(LLMs)for standardized tasks,their multidimensional performance in real-world oncology decision-making remains unevaluated.This study aims to introduce a framework for evaluating LLMs and physician decisions in challenging lung cancer cases.Methods:We curated 50 challenging lung cancer cases(25 local and 25 published)classified as complex,rare,or refractory.Blinded three-dimensional,five-point Likert evaluations(1–5 for comprehensiveness,specificity,and readability)compared standalone LLMs(DeepSeek R1,Claude 3.5,Gemini 1.5,and GPT-4o),physicians by experience level(junior,intermediate,and senior),and AI-assisted juniors;intergroup differences and augmentation effects were analyzed statistically.Results:Of 50 challenging cases(18 complex,17 rare,and 15 refractory)rated by three experts,DeepSeek R1 achieved scores of 3.95±0.33,3.71±0.53,and 4.26±0.18 for comprehensiveness,specificity,and readability,respectively,positioning it between intermediate(3.68,3.68,3.75)and senior(4.50,4.64,4.53)physicians.GPT-4o and Claude 3.5 reached intermediate physician–level comprehensiveness(3.76±0.39,3.60±0.39)but junior-to-intermediate physician–level specificity(3.39±0.39,3.39±0.49).All LLMs scored higher on rare cases than intermediate physicians but fell below junior physicians in refractory-case specificity.AIassisted junior physicians showed marked gains in rare cases,with comprehensiveness rising from 2.32 to 4.29(84.8%),specificity from 2.24 to 4.26(90.8%),and readability from 2.76 to 4.59(66.0%),while specificity declined by 3.2%(3.17 to 3.07)in refractory cases.Error analysis showed complementary strengths,with physicians demonstrating reasoning stability and LLMs excelling in knowledge updating and risk management.Conclusions:LLMs performed variably in clinical decision-making tasks depending on case type,performing better in rare cases and worse in refractory cases requiring longitudinal reasoning.Complementary strengths between LLMs and physicians support case-and task-tailored human–AI collaboration.

Ning Yang;Kailai Li;Baiyang Liu;Xiting Chen;Aimin Jiang;Chang Qi;Wenyi Gan;Lingxuan Zhu;Weiming Mou;Dongqiang Zeng;Mingjia Xiao;Guangdi Chu;Shengkun Peng;Hank ZHWong;Lin Zhang;Hengguo Zhang;Xinpei Deng;Quan Cheng;Bufu Tang;Anqi Lin;Juan Zhou;Peng Luo

Department of Oncology,General Hospital of Southern Theater Command,Guangzhou Guangdong 510030,ChinaDepartment of Oncology,Zhujiang Hospital of Southern Medical University,Guangzhou Guangdong 510282,ChinaDepartment of Oncology,General Hospital of Southern Theater Command,Guangzhou Guangdong 510030,ChinaGraduate School,Guangzhou University of Chinese Medicine,Guangzhou Guangdong 510006,ChinaDepartment of Urology,Changhai Hospital,Naval Medical University(Second Military Medical University),Shanghai 200433,ChinaDepartment of Clinical Oncology,The University of Hong Kong,Hong Kong SAR 999077,ChinaDepartment of Joint Surgery and Sports Medicine,Zhuhai People''s Hospital(Zhuhai Clinical Medical College of Jinan University),Zhuhai Guangdong 519000,ChinaDepartment of Oncology,Zhujiang Hospital of Southern Medical University,Guangzhou Guangdong 510282,ChinaDepartment of Oncology,Zhujiang Hospital of Southern Medical University,Guangzhou Guangdong 510282,China Department of Urology,Shanghai General Hospital,Shanghai Jiao Tong University School of Medicine,Shanghai 200080,ChinaDepartment of Oncology,Nanfang Hospital,Southern Medical University,Guangzhou Guangdong 510515,China Cancer Center,The Sixth Affiliated Hospital,South China University of Technology,Guangzhou Guangdong 510006,ChinaHepatobiliary Surgery Department,People''s Hospital of Quzhou(Quzhou Affiliated Hospital of Wenzhou Medical University),Quzhou Zhejiang 324000,ChinaDepartment of Urology,The Affiliated Hospital of Qingdao University,Qingdao Shandong 266003,ChinaDepartment of Radiology,Sichuan Provincial People''s Hospital/Affiliated Hospital of University of Electronic Science and Technology of China,Chengdu Sichuan 610072,ChinaLi Ka Shing Faculty of Medicine,The University of Hong Kong,Hong Kong SAR 999077,ChinaSchool of Public Health and Preventive Medicine,Monash University,Melbourne VIC 3000,Australia Suzhou Industrial Park Monash Research Institute of Science and Technology,Suzhou Jiangsu 215000,ChinaCollege&Hospital of Stomatology,Anhui Medical University,Key Laboratory of Oral Diseases Research of Anhui Province,Hefei Anhui 230032,ChinaDepartment of Urology,State Key Laboratory of Oncology in Southern China,Sun Yat-sen University Cancer Center,Guangdong Provincial Clinical Research Center for Cancer,Guangzhou Guangdong 510060,ChinaDepartment of Neurosurgery,Xiangya Hospital,Central South University,Changsha Hunan 410008,China National Clinical Research Center for Geriatric Disorders,Xiangya Hospital,Central South University,Changsha Hunan 410008,ChinaDepartment of Radiation Oncology,Zhongshan Hospital Affiliated to Fudan University,Shanghai 200032,ChinaDepartment of Oncology,Zhujiang Hospital of Southern Medical University,Guangzhou Guangdong 510282,ChinaDepartment of Oncology,General Hospital of Southern Theater Command,Guangzhou Guangdong 510030,China Graduate School,Guangzhou University of Chinese Medicine,Guangzhou Guangdong 510006,China The First School of Clinical Medicine,Southern Medical University,Guangzhou Guangdong 510515,ChinaDepartment of Oncology,Zhujiang Hospital of Southern Medical University,Guangzhou Guangdong 510282,China

医药卫生

Large language modelsClinical evaluationDecision-makingLung cancer

《Intelligent Oncology》 2026 (1)

P.15-24,10

10.1016/j.intonc.2026.100039

评论