文本导向的多任务多模态情感感知分析模型OA
Text-Guided Multi-task Multimodal Emotion Perception Analysis Model
针对现有多模态情感分析模型在对文本模态信息的重视及上下文利用方面存在的不足,本文提出一种融合文本增强和跨模态交互注意力机制的多任务多模态情感感知分析模型.通过引入基于GPT的文本增强技术,提升情感词识别能力,加强文本模态;采用跨模态交互注意力机制,融合视觉、音频与文本信息;应用同方差不确定性损失函数,优化多任务权重调整.在CMU-MOSI和CMU-MOSEI数据集上,将本模型与基线模型进行比较,ACC2 和F1 指标分别达到87.2%和 85.8%,优于基线模型,证明了本模型的有效性,并显著提升了多模态情感感知分析性能.
To address the limitations of existing multimodal sentiment analysis models in underemphasizing the text modality and utilizing context,a multi-task multimodal emotion perception analysis model that integraces text augmentation and cross-modal interactive attention mechanisms is proposed.Specifically,we introduce a GPT-based text augmentation technique to enhance the recognition of emotional words and reinforce the text modality.A cross-modal interactive attention mechanism is then employed to fuse visual,audio and textual information effectively.Moreover,a homoscedastic uncertainty loss function is applied to optimize weight adjustments across multiple tasks.On the CMU-MOSI and CMU-MOSEI datasets,the proposed model achieves ACC2 and F1 scores of 87.2%and 85.8%respectively,outperforming the baseline model.This demonstrates the effectiveness of the proposed model and its significant improvement in multimodel perception analysis performance.
臧洁;李翔;卢睿;廖慧之;任赛赛;卢珊
辽宁大学 信息学部,辽宁 沈阳 110036辽宁大学 信息学部,辽宁 沈阳 110036辽宁警察学院 网络安全学院,辽宁 大连 116036辽宁大学 信息学部,辽宁 沈阳 110036辽宁大学 信息学部,辽宁 沈阳 110036辽宁大学 信息学部,辽宁 沈阳 110036
信息技术与安全科学
多模态情感分析情感词感知文本信息增强多任务学习
multimodal sentiment analysisemotional word perceptiontext information enhancementmulti-task learning
《辽宁大学学报(自然科学版)》 2026 (1)
51-61,11
辽宁省科技厅应用基础研究(2023JH2/101300134)2025辽宁大学研究生优质课程建设与教改研究项目(YJG202501075)
评论