Efficient text-to-video retrieval via multi-modal multi-tagger derived pre-screeningOA
Text-to-video retrieval(TVR)has made significant progress with advances in vision and language representation learning.Most existing methods use real-valued and hash-based embeddings to represent the video and text,allowing retrieval by computing their similarities.However,these methods are often inefficient for large volumes of video,and require significant storage and computing resources.In this work,we present a plug-and-play multi-modal multi-tagger-driven pre-screening framework,which pre-screens a substantial number of videos before applying any TVR algorithms,thereby efficiently reducing the search space of videos.We predict discrete semantic tags for video and text with our proposed multi-modal multi-tagger module,and then leverage an inverted index for space-efficient and fast tag matching to filter out irrelevant videos.To avoid filtering out relevant videos for text queries due to inconsistent tags,we utilize contrastive learning to align video and text embeddings,which are then fed into a shared multi-tag head.Extensive experimental results demonstrate that our proposed method significantly accelerates the TVR process while maintaining high retrieval accuracy on various TVR datasets.
Yingjia Xu;Mengxia Wu;Zixin Guo;Min Cao;Mang Ye;Jorma Laaksonen
School of Computer Science&Technology,Soochow University,Suzhou,215006,ChinaSchool of Computer Science&Technology,Soochow University,Suzhou,215006,ChinaDepartment of Computer Science,Aalto University,Espoo,02150,FinlandSchool of Computer Science&Technology,Soochow University,Suzhou,215006,ChinaSchool of Computer Science,Wuhan University,Wuhan,430072,ChinaDepartment of Computer Science,Aalto University,Espoo,02150,Finland
信息技术与安全科学
Text-to-video retrieval(TVR)Inverted indexPre-screeningContrastive learning(CL)
《Visual Intelligence》 2025 (1)
P.21-33,13
supported by the National Natural Science Foundation of China(No.62476188)the Open Projects Program of the State Key Laboratory of Multimodal Artificial Intelligence Systems,and the Academy of Finland in the USSEE Project(No.345791).
评论