Can large language models generate geospatial codeOA
Can large language models generate geospatial code?
As large language models increasingly exhibit hallucinations such as refusal to respond,generation of non-executable code,and poor readability in geospatial code generation tasks,establishing a systematic and quantifiable evaluation framework has become essential for advancing their application in GIS.This paper introduces the GeoCode-Eval framework,the first comprehensive evaluation framework for LLMs in geospatial code generation.Grounded in three dimensions-cognition and memory,understanding and interpretation,and innova-tion and creation-the framework addresses eight competency levels,including platform and tool cognition,functional knowledge,dataset recognition,information extraction,and various code-related tasks.To support this,the GeoCode-Bench benchmark was developed,consisting of 5,000 multiple-choice questions,1,500 true/false questions,1,500 fill-in-the-blank questions,and 1,000 coding tasks.Using six indicators,namely executability,accuracy,readability,loca-tion correctness,content correctness,and summary completeness,the study evaluates twelve representative models spanning four categories,including DeepSeek-Coder-V2 and GeoCode-GPT(7B).A combination of analytical methods,including entropy weighting,the coefficient of variation,skewness,and kurtosis,is applied to examine model capability distribution,indicator distribution,code type characteristics,and error type patterns.Results show consistent perfor-mance in tool cognition and code summarization,while significant performance gaps persist in code generation,completion,and correction.Common errors include data type and syntax issues.This study provides a quantifiable foundation for the evaluation of capabilities and future optimization of LLMs in geospatial code generation,thereby extending the application boundaries of LLMs in GIS and offering valuable insights into the development of evaluation methodologies for LLM applications in other vertical domains.
Shuyang Hou;Longgang Xiang;Huayi Wu;Zhangxiao Shen;Jianyuan Liang;Haoyue Jiao;Anqi Zhao;Yaxian Qing;Dehua Peng;Zhipeng Gui;Xuefeng Guan
State Key Laboratory of Information Engineering in Surveying,Mapping,and Remote Sensing,Wuhan University institution institution,Wuhan,ChinaState Key Laboratory of Information Engineering in Surveying,Mapping,and Remote Sensing,Wuhan University institution institution,Wuhan,ChinaState Key Laboratory of Information Engineering in Surveying,Mapping,and Remote Sensing,Wuhan University institution institution,Wuhan,ChinaState Key Laboratory of Information Engineering in Surveying,Mapping,and Remote Sensing,Wuhan University institution institution,Wuhan,ChinaState Key Laboratory of Information Engineering in Surveying,Mapping,and Remote Sensing,Wuhan University institution institution,Wuhan,ChinaSchool of Resource and Environmental Sciences,Wuhan University,Wuhan,ChinaState Key Laboratory of Information Engineering in Surveying,Mapping,and Remote Sensing,Wuhan University institution institution,Wuhan,ChinaState Key Laboratory of Information Engineering in Surveying,Mapping,and Remote Sensing,Wuhan University institution institution,Wuhan,ChinaSchool of Remote Sensing and Information Engineering,Wuhan University,Wuhan,ChinaSchool of Remote Sensing and Information Engineering,Wuhan University,Wuhan,ChinaState Key Laboratory of Information Engineering in Surveying,Mapping,and Remote Sensing,Wuhan University institution institution,Wuhan,China
Geospatial code generationlarge language modelsGISDeepSeekLLM evaluation
Geospatial code generationlarge language modelsGISDeepSeekLLM evaluation
《地球空间信息科学学报(英文版)》 2026 (2)
中插1,753-787,36
The work was supported by the National Natural Science Foundation of China[Grant Nos.41930107 and 41971349].
评论