首页|期刊导航|华南理工大学学报(自然科学版)|多模态商品摘要生成的要素评估与偏好优化

多模态商品摘要生成的要素评估与偏好优化OA

Component Evaluation and Preference Optimization for Multimodal Product Summary Generation

中文摘要英文摘要

多模态商品摘要生成任务旨在结合商品的图文信息,生成简洁准确、突出核心卖点的摘要文本.然而,现有方法仍面临两大挑战:一是ROUGE等传统基于词汇重叠的评价指标,无法精准衡量摘要对商品关键信息的完整表达能力;二是主流监督微调范式难以挖掘用户对商品要素重要程度的隐性偏好,致使生成摘要与实际使用需求存在偏差.为此,该文首先提出了基于要素的摘要评价指标(CSE),从要素命中准确率(CHA)、要素数量比(CQR)2个维度综合评估摘要对商品关键信息的表达效果;在此基础上进一步构建基于偏好优化的多模态商品摘要生成模型PAMPS.该模型通过监督微调、摘要重采样、要素评价引导的偏好样本对构建、直接偏好优化4个训练阶段,实现模型输出与用户商品要素表达偏好的对齐.基于大规模中文电商数据集CEPSUM开展对比实验,结果显示:PAMPS模型在传统ROUGE指标上实现明显性能提升,其中Qwen2-VL-DPO-ROUGE相较于Qwen2-VL-SFT,ROUGE-1、ROUGE-2、ROUGE-L指标分别平均提升0.25、0.44和1.21个百分点,整体摘要生成质量更优;在所构建的CSE评价体系下,Qwen2-VL-DPO-CSE模型的CHA指标提升效果尤为突出,平均相对提升约4%,证明面向商品要素的偏好优化策略能够显著增强模型对商品核心要素的捕捉与表达能力.综合实验结果充分验证了所提方法对提升多模态商品摘要生成质量的有效性与工程实用价值.

The task of multimodal product summary generation aims to produce concise and accurate summary texts that highlight the core selling points of products by integrating both textual and visual information.However,existing methods still face two major challenges.First,traditional evaluation metrics such as ROUGE,which are based on lexical overlap,fail to precisely measure the completeness of a summary in expressing key product information.Second,mainstream supervised fine-tuning paradigms struggle to capture users'implicit preferences regarding the relative importance of different product information components,resulting in a misalignment between generated summaries and actual usage requirements.To address these issues,this paper proposes a component-based summarization evaluation metric,termed CSE(Component-based Summary Evaluation),which comprehensively evaluates the expression of key information from two perspectives:Component Hit Accuracy(CHA)and Component Quantity Ratio(CQR).Furthermore,a preference optimization-based multimodal product summary generation model,PAMPS,is designed.The model aligns its outputs with users'preferences for product component expression through four training stages:supervised fine-tuning,summary resampling,component evaluation-guided preference pair construction,and direct preference optimization.Experimental results on the large-scale Chinese e-commerce dataset CEPSUM show that PAMPS achieves significant improvements on ROUGE metrics.In particular,compared with Qwen2-VL-SFT,Qwen2-VL-DPO-ROUGE improves ROUGE-1,ROUGE-2,and ROUGE-L by 0.25,0.44,and 1.21 percentage points on average,respectively,demonstrating superior overall summary generation quality.Under the proposed CSE evaluation framework,the Qwen2-VL-DPO-CSE model demonstrates particularly prominent improvements in the CHA metric,with an average relative improvement of about 4%,indicating that component-oriented preference optimization can effectively enhance the model's ability to capture and express core product features.The comprehensive experimental results fully validate the effectiveness and practical applicability of the proposed method in improving the quality of multimodal product summary generation.

宋雪萌;李芷墨;侯博涵;尹建华

南方科技大学 计算机科学与工程系,广东 深圳 518055北京大学 计算机学院,北京 100871山东大学 计算机科学与技术学院,山东 青岛 266237山东大学 计算机科学与技术学院,山东 青岛 266237

信息技术与安全科学

多模态大模型摘要评估偏好优化

multimodal large modelsummarization evaluationpreference optimization

《华南理工大学学报(自然科学版)》 2026 (8)

36-48,13

国家自然科学基金项目(62376137,62172261)山东省自然科学基金项目(ZR2022YQ59) Supported by the National Natural Science Foundation of China(62376137,62172261)and the Natural Science Foundation of Shandong Province(ZR2022YQ59)

10.12141/j.issn.1000-565X.250375

评论