首页|期刊导航|高电压技术|MM-Power:面向电力巡检场景的视觉大模型多模态评测基准

MM-Power:面向电力巡检场景的视觉大模型多模态评测基准OA

MM-Power:A Multi-modal Benchmark for Evaluating Vision-language Models in Power Grid Inspection

中文摘要英文摘要

随着视觉大模型在通用多模态任务上的快速发展,其在电力巡检等专业场景中的应用潜力日益受到关注.然而,现有研究对于此类模型在工业细粒度感知、视觉推理和抗幻觉方面的真实能力边界仍缺乏系统认识.针对通用多模态基准难以覆盖电力巡检中专业目标繁多、小目标缺陷突出、成像视角复杂以及误报代价较高等问题,该文提出面向电力巡检任务的多模态评测基准MM-Power.该基准从细粒度感知、视觉语义一致性与幻觉评估、视觉逻辑推理3个层面构建12项任务,共包含2 708道选择题,覆盖可见光与红外两种模态、输电/变电/配电3类业务场景、云台/无人机/机器人/手持摄像机/手持红外测温仪5类采集视角,以及130类设备状态、缺陷与异常.进一步地,该文选取16个具有代表性的开源与闭源视觉大模型,在中英文提示词条件下开展统一评测.结果表明:英文提示词下最佳模型总体准确率仅为77.34%,中文提示词下最佳模型总体准确率为75.54%;模型在存在性判断、负样本鉴别和空间关系推理等任务上普遍表现较弱,提示词语言还会对部分模型的工业场景表现产生明显影响.研究结果表明,当前通用视觉大模型距离可靠的电力巡检应用仍有较大差距,MM-Power可为电力领域多模态模型的评测、优化与部署提供统一基准.

General-purpose vision-language models have shown strong performance on widely used multimodal bench-marks,yet a systematic understanding of their true capability boundaries in industrial fine-grained perception,visual reasoning,and hallucination resistance remains lacking.In power grid inspection,generic multimodal benchmarks often fail to cover the large variety of domain-specific targets,prominent small-target defects,complex imaging perspectives,and the high cost of false alarms.To address the mismatch between existing generic benchmarks and the demands of re-al-world power inspection,this paper presents MM-Power,a multimodal benchmark tailored to power grid inspection scenarios.MM-Power contains 2 708 multiple-choice questions organized into three tracks and 12 tasks,covering fi-ne-grained perception,visual-semantic consistency and hallucination assessment,and visual reasoning.The benchmark spans visible-light and infrared modalities,transmission/substation/distribution scenarios,five acquisition perspectives in-cluding pan-tilt cameras,unmanned aerial vehicles,robots,handheld cameras,and handheld infrared thermometers,and 130 categories of equipment states,defects,and anomalies.Based on a unified evaluation of 16 representative proprietary and open-weight vision-language models with both English and Chinese prompts,the best model achieves only 77.34%overall accuracy in English and 75.54%in Chinese.Significant weaknesses are observed in existence judgments,nega-tive-sample identification,and spatial relation reasoning,while prompt language also has a noticeable impact on the industrial-scenario performance of some models.These findings indicate that current vision-language models are still far from reliable deployment in practical power inspection,and MM-Power provides a unified benchmark for future evalua-tion,optimization,and deployment of multimodal models in the power domain.

闫云凤;齐冬莲;邓以恒;蔡芳仪;陈毅;龚世超;熊周智;曾子墨;蔡舒瑶;林嘉扬

浙江大学电气工程学院,杭州 310027浙江大学电气工程学院,杭州 310027浙江大学电气工程学院,杭州 310027西安交通大学电气工程学院,西安 710049浙江大学海洋学院,舟山 316021浙江大学电气工程学院,杭州 310027浙江大学电气工程学院,杭州 310027同济大学汽车与能源学院,上海 201804浙江大学电气工程学院,杭州 310027浙江大学电气工程学院,杭州 310027

视觉大模型电力巡检多模态评测幻觉评估视觉推理工业智能

vision-language modelpower grid inspectionmultimodal benchmarkhallucination evaluationvisual rea-soningindustrial intelligence

《高电压技术》 2026 (7)

2998-3009,12

国家自然科学基金(62476242)浙江省"尖兵""领雁"研发攻关计划项目(2025C01058).Project supported by National Natural Science Foundation of China(62476242),Pioneer R&D Program of Zhejiang Province(2025C01058).

10.13336/j.1003-6520.hve.20260524

评论