首页|期刊导航|中国防汛抗旱|基于社交媒体与大模型的城市内涝灾情信息提取研究

基于社交媒体与大模型的城市内涝灾情信息提取研究OA

Urban waterlogging disaster information extraction based on social media and large language models

中文摘要英文摘要

城市下垫面环境复杂、基础设施薄弱,在面临较大降雨时易发生内涝灾害,快速识别与提取灾情信息是提升城市防洪响应效率和灾害管理能力的重要基础.社交媒体因其实时性与高参与度,已成为城市内涝灾情信息获取的重要数据来源.针对社交媒体灾情文本信息规模大、语义复杂且难以直接服务决策等问题,以郑州"7·20"特大暴雨为例,基于城市内涝灾害链构建了15类精细化分类体系,研究提出了一种大语言模型驱动的灾情信息提取方法,并引入BERT-TextCNN模型作为基准进行对比分析.结果表明:①大语言模型具备较高的少样本学习效能.大语言模型DeepSeek、ChatGPT及GLM在少样本量下即可实现灾情分类,且性能随数据规模呈正相关增长,其中ChatGPT表现最优,加权平均F1值最高达0.89.②深度学习基准模型在多样本下表现稳健.BERT-TextCNN在充足标注数据支撑下展现出可靠的分类能力,加权平均F1值为0.86.③大语言模型具有显著的数据效率优势.DeepSeek与ChatGPT在450条标注样本的情况下,其综合性能即可优于基于5 600条样本训练的BERT-TextCNN模型,显著降低了人工标注成本.本研究可为城市内涝灾情信息提取及应急决策支持提供参考.

Urban underlying surfaces are complex and infrastructure is inadequate,making cities prone to waterlogging disasters during heavy rainfall.Rapid identification and extraction of disaster information is crucial for improving urban flood response efficiency and disaster management capabilities.Social media,due to its real-time nature and high public participation,has become an important data source for obtaining urban waterlogging disaster information.To address challenges such as the large scale,semantic complexity,and difficulty in directly supporting decision-making of social media disaster text information,this study takes the Zhengzhou"7·20"rainstorm as an example and constructs a 15-category fine-grained classification system based on the urban waterlogging disaster chain.A large language model(LLM)-driven disaster information extraction method is proposed,with the BERT-TextCNN model introduced as a benchmark for comparative analysis.The results show that:①LLMs exhibit high few-shot learning efficiency.The LLMs DeepSeek,ChatGPT,and GLM can achieve disaster classification with small sample sizes,and their performance shows a positive correlation with data scale.Among them,ChatGPT performs best,achieving a macro-and micro-averaged F1 score of up to 0.89.②The deep learning benchmark model performs robustly with larger sample sizes.BERT-TextCNN demonstrates reliable classification capability with sufficient labeled data,achieving a weighted F1 score of 0.86.③LLMs have a significant data efficiency advantage.With only 450 labeled samples,DeepSeek and ChatGPT achieve comprehensive performance superior to the BERT-TextCNN model trained on 5 600 samples,significantly reducing the cost of manual annotation.This study can provide a reference for urban waterlogging disaster information extraction and emergency decision support.

杨晨;李松涛;刘颖;张会;李硕

华北水利水电大学数字孪生水利高等研究院,郑州 450046华北水利水电大学水利学院,郑州 450046华北水利水电大学水资源学院,郑州 450046华北水利水电大学数字孪生水利高等研究院,郑州 450046华北水利水电大学数字孪生水利高等研究院,郑州 450046

天文与地球科学

城市内涝社交媒体信息提取BERT-TextCNN大语言模型

urban waterloggingsocial mediainformation extractionBERT-TextCNNlarge language models

《中国防汛抗旱》 2026 (5)

7-14,8

河南省科技攻关项目(252102321017)河南省高校重点科研项目计划基础研究专项(26ZX019).

10.16867/j.issn.1673-9264.2026189

评论