基于多模态视觉语言模型的老旧社区环境特征解析OA
Decoding Environmental Characteristics of Old Urban Residential Compounds Via Multi-modal Vision-Language Models:A Case Study of the City Core of Chongqing
[目的]针对老旧社区环境特征识别中传统方法效率低、受主观因素影响较大的局限,提出一种融合多模态视觉语言模型(VLM)与带地理信息的众源图像的创新方法框架,识别与提取老旧社区环境特征,并对其进行量化与归类.[方法]以重庆市中心城区为研究区,通过构建众源图像数据集,利用多模态视觉语言模型提取图像语义转化为文本,结合BERTopic主题建模技术提取聚类,将聚类结果映射到地理空间分析特征分布和共现结果.[结果]提取 50 个聚类得到涵盖空间骨架、微观细节、绿化形式及场所氛围的 17 组环境感知特征主题,识别出 7 类典型的特征空间共现模式,为空间定向改造提供建议.[结论]老旧社区环境具有复杂的异质性特征,并受历史风貌与公共生活的深度影响;多模态分析框架能够有效对老旧社区环境图像进行语义理解,实现低成本、高通量且精细化的环境特征挖掘.未来可以进一步结合社会学数据构建完整的更新闭环,真正服务于"好房子、好小区、好社区、好城区"的系统性建设.
[Objective]To address the limitations of traditional methods in identifying environmental characteristics of old urban residential compounds,including low efficiency and susceptibility to subjective factors,this study proposes an innovative methodological framework integrating Multi-modal Vision-Language Models(VLM)with Geotagged Crowdsourced Imagery to have the environmental characteristics of old urban residential compounds decoded and interpreted,which is then quantified and classified.[Method]With the city core of Chongqing taken as the empirical study area,a dataset of crowdsourced images is constructed.VLMs are utilized to transcode visual semantics into textual descriptions,which are then processed using BERTopic modeling for clustering analysis.The resulting clusters are mapped onto geographic space to analyze the spatial distribution and co-occurrence patterns of environmental characteristics.[Result]The extraction of 50 clusters yields 17 thematic groups of environmental perception features,covering spatial skeletons,micro-scale details,greenery forms,and place ambiance.Furthermore,7 typical spatial co-occurrence patterns of these characteristics are identified,providing suggestions for spatially targeted renovation.[Conclusion]The environments of old urban residential compounds exhibit complex heterogeneity,profoundly influenced by historical features and public life.The multi-modal analysis framework effectively enables semantic understanding of environmental images,achieving low-cost,high-throughput,and fine-grained feature mining.Future research can further integrate sociological data to build a complete loop of renewal,truly serving the systematic construction of"good houses,good neighborhoods,good communities,and good urban districts".
李彦锦;罗丹;肖竞;李玮
重庆大学建筑城规学院,重庆 400045重庆大学建筑城规学院,重庆 400045重庆大学建筑城规学院,重庆 400045长春中海地产有限公司,长春 130000
多模态视觉语言模型老旧社区众源图像空间特征重庆市中心城区
multimodal vision-language modelold urban residential compoundcrowd-sourced imageryspatial featurecity core of Chongqing
《中国城市林业》 2026 (1)
9-19,11
国家自然科学基金重点项目(52238003)重庆市社会科学青年项目(2021NDQN64)
评论