基于全局局部交互模型的多模态情感分析OA
Multimodal Sentiment Analysis Based on Global-local Interaction Model
多模态方面级情感分析(Multimodal Aspect-Based Sentiment Analysis,MABSA)作为情感计算领域的关键研究方向,致力于融合文本、图像、音频等多种模态信息,用以实现对特定方面情感的精细化分析.在当前的多模态方面级情感分析研究中,存在着图像噪声干扰以及过度依赖局部特征等问题,进而影响了分析的准确性和全面性.针对这些局限,本文提出了一种创新的全局局部交互情感分析模型(Global-Local Interactive Emotion Analysis Model,GLIEAM).一方面,采用视觉Transformer(Vision Transformer,ViT)模型与生成式预训练Transformer(Generative Pre-trained Transformer,GPT)模型串联生成图像描述并将其与原始文本特征进行拼接,有效强化了信息融合效果,从而更全面地捕捉多模态数据中的情感线索.另一方面,为解决图像噪声问题,结合小波变换和非局部均值方法对图像进行去噪处理;同时,利用卷积神经网络(Convolutional Neural Network,CNN)和T2T视觉Transformer(Tokens-to-Token Vision Transformer,T2T-ViT)分别提取局部和全局图像特征,避免了对局部特征的过度依赖,实现了对图像特征的全面、均衡提取.通过在基准数据集上进行实验,结果表明该方法显著优于现有方法,在Twitter-15上准确率达到78.46%,Twitter-17数据集上准确率达到75.21%,尤其在低资源场景下展现出卓越的性能.
Multimodal aspect-based sentiment analysis(MABSA)is a critical research direction in the field of affective computing,aiming to integrate multimodal information,such as text,images,and audio to achieve fine-grained analysis of sentiment toward spe-cific aspects.Current research in MABSA faces challenges such as image noise interference and excessive reliance on local features,which compromise the accuracy and comprehensiveness of the analysis.To address these limitations,this paper proposes an innova-tive global-local interactive emotion analysis model(GLIEAM).On the one hand,the model employs a tandem architecture of vi-sion transformer(Vision Transformer,ViT)and generative pre-trained transformer(GPT)to generate image descriptions,which are then concatenated with original text features,significantly enhancing information fusion and enabling a more comprehensive capture of emotional cues in multimodal data.On the other hand,to mitigate image noise,a hybrid approach combining wavelet transform and non-local means is applied for image denoising.Additionally,convolutional neural networks(CNN)and tokens-to-token vision transformer(T2T-ViT)are utilized to extract local and global image features,respectively,avoiding over-reliance on local features and achieving balanced and holistic image feature extraction.Experimental results on benchmark datasets demonstrate that the pro-posed method outperforms existing approaches,the accuracy reached 78.46%on the Twitter-15 dataset and 75.21%on the Twitter-17 dataset,particularly exhibiting superior performance in low-resource scenarios.
李梦晗;仲兆满;徐俊康;陈柯含
江苏海洋大学 计算机工程学院,江苏 连云港 222005江苏海洋大学 计算机工程学院,江苏 连云港 222005||江苏省海洋资源开发研究院,江苏 连云港 222005江苏海洋大学 计算机工程学院,江苏 连云港 222005江苏海洋大学 计算机工程学院,江苏 连云港 222005
信息技术与安全科学
多模态方面级情感分析视觉Transformer特征融合图像去噪注意力机制
multimodal aspect-based sentiment analysisvision Transformerfeature fusionimage denoisingattention mechanism
《山西大学学报(自然科学版)》 2026 (1)
1-14,14
国家自然科学基金(72174079)江苏省"青蓝工程"大数据优秀教学团队(2022-29)连云港市重点研发(产业前瞻与关键核心技术)项目(CG2323)
评论