ColorAlignNet:基于参考帧的时间聚合视频着色网络OA
ColorAlignNet:a Reference-Based Video Colorization Network with Temporal Aggregation
视频着色是一项为老旧视频注入新生命的技术.尽管现有的着色方法在静态图像和低动态视频上表现出色,但它们通常难以处理复杂的动态场景.为此,本研究提出了一种基于参考帧的时间聚合视频着色网络 ColorAlignNet.该网络使用源-参考注意力机制,将参考帧中的颜色信息有效传播至灰度帧,确保色彩还原的准确性.同时,通过设计基于可变形卷积的特征对齐模块,对相邻帧进行特征对齐,以提升时序一致性.最后,结合循环 transformer 模块来重构最终的预测结果.大量实验结果表明,ColorAlignNet 在DAVIS 和 Videvo 数据集上取得了优异性能,在感知图像块相似度(learned perceptual image patch similarity,LPIPS)和色彩分布一致性(color distribution consistency,CDC)指标上均优于现有主流方法.
Video colorization is an important technique to breathe life back into old movies.While current colorization methods work well on still images and low-motion video data,they often struggle with complex dynamic scenes.To address this problem,this study proposes ColorAlignNet,a reference-based video colorization network with temporal aggregation.The network uses source-reference attention to propagate color information from reference frames to grayscale frames,guaranteeing color accuracy,and uses deformable convolution to align features of adjacent frames to enhance temporal consistency.Finally,we use the cyclic transformer module to reconstruct the final prediction results.Extensive experimental results demonstrate that ColorAlignNet achieves excellent performance on the DAVIS and Videvo datasets,outperforming other state-of-the-art methods on both the learned perceptual image patch similarity(LPIPS)and color distribution consistency(CDC)metrics.
朱文志;王彤
东华大学 信息科学与技术学院,上海 201620||东华大学 数字化纺织服装技术教育部工程研究中心,上海 201620东华大学 信息科学与技术学院,上海 201620||东华大学 数字化纺织服装技术教育部工程研究中心,上海 201620
信息技术与安全科学
可变形卷积视频着色Swin-transformer
deformable convolutionvideo colorizationSwin-transformer
《东华大学学报(英文版)》 2026 (2)
94-102,9
评论