M2HF:Multi-branch multi-modal hybrid fusion for text-video retrievalOA
Videos contain multi-modal content,and exploring multi-branch cross-modal interactions with natural language queries can be of benefit to the text-video retrieval task(TVR).However,recent methods applying the large-scale pre-trained CLIP model for TVR only focus on visual cues in videos.Furthermore,traditional methods of simply concatenating multimodal features do not exploit fine-grained cross-modal information in videos.In this paper,we propose a multi-branch multi-modal hybrid fusion(M2HF)network to hierarchically explore interaction between text queries and other modality content in videos.Specifically,M2HF first fuses visual features extracted by CLIP with audio and motion features extracted from videos to obtain fused audio-visual features and motion-visual features respectively.The multi-modal completion problem is also considered and solved in this process.Then,visual features,audio-visual features,motion-visual features,and text extracted from the video are used to establish cross-modal relationships with caption text queries using a multibranch approach.The retrieval outputs from all branches are then fused to obtain the final text-video retrieval results.Our framework provides two kinds of training strategies,using an ensemble approach and an end-to-end approach.Moreover,a novel multi-modal loss function is proposed to balance the contributions of each modality for efficient end-to-end training.M2HF allows us to obtain state-of-the-art results on various benchmarks:Rank@1 of 66.0%,68.6%,33.9%,57.4%,and 57.3%on MSR-VTT,MSVD,LSMDC,DiDeMo,and ActivityNet,respectively.
Shuo Liu;Weize Quan;Ming Zhou;Sihong Chen;Jian Kang;Zhe Zhao;Kimmo Yan;Chen Chen;Dong-Ming Yan
MAIS,Institute of Automation,Chinese Academy of Sciences,Beijing 100190,China School of Artificial Intelligence,University of Chinese Academy of Sciences,Beijing 100049,ChinaMAIS,Institute of Automation,Chinese Academy of Sciences,Beijing 100190,China School of Artificial Intelligence,University of Chinese Academy of Sciences,Beijing 100049,ChinaSchool of Information Science and Technology,Donghua University,Shanghai 200051,ChinaTencent,Shenzhen 518057,ChinaTencent,Shenzhen 518057,ChinaTencent,Shenzhen 518057,ChinaTencent,Shenzhen 518057,ChinaTencent,Shenzhen 518057,ChinaMAIS,Institute of Automation,Chinese Academy of Sciences,Beijing 100190,China School of Artificial Intelligence,University of Chinese Academy of Sciences,Beijing 100049,China
信息技术与安全科学
multi-modalitymulti-branchhybrid fusiontext–video retrieval
《Computational Visual Media》 2026 (2)
P.449-471,23
partially supported by the National Natural Science Foundation of China(62102418,62172415)the CCF-Tencent Rhino-Bird Research Fund(RAGR20210124)the Beijing Science and Technology Plan Project(Z231100005923033).
评论