首页|期刊导航|山西大学学报(自然科学版)|一种基于多尺度内容感知的图像篡改定位方法

一种基于多尺度内容感知的图像篡改定位方法OA

A Multi-scale Content-aware Localization Method for Image Manipulation

中文摘要英文摘要

随着图像编辑技术的快速发展,图像篡改定位面临日益复杂的挑战.现有方法在捕捉篡改区域的多尺度上下文信息方面表现不足,且在多尺度特征融合过程中采用固定采样核,难以关注特征的局部变化,导致定位精度受限.针对这些问题,本文提出一种基于多尺度内容感知的图像篡改定位方法.首先,设计了基于多尺度内容感知的特征融合模块,能够为特征图中的每个位置动态生成自适应的采样核,使模型在粗粒度上定位篡改区域的大致范围,同时在细粒度上识别篡改边缘特征.其次,采用深度可分离卷积解码器替代传统的多层感知机进行预测,进一步提升检测准确性.最后,提出一种结合二元交叉熵损失和Dice损失的联合损失函数,有效增强了模型的鲁棒性和泛化能力.在多个公开数据集上的跨数据集实验结果表明,提出的方法在CASIAv1、Defacto-12k、Coverage和Columbia数据集上,Pixel-level F1 score分别达到了69.4%、25.4%、37.8%和83.2%.相较于主流的MVSS-Net++和最新的IML-ViT,平均提升了10.2%和2.1%,显著提高了图像篡改定位的精度.

With the rapid advancement of image editing technologies,image manipulation localization is facing increasingly complex challenges.Existing methods exhibit limitations in capturing multi-scale contextual information of manipulated regions,and they of-ten rely on fixed sampling kernels during the multi-scale feature fusion process,which makes it difficult to focus on local variations of features,thereby restricting the localization accuracy.To address these issues,we propose an image manipulation localization method based on multi-scale content-aware.Firstly,a feature fusion module based on multi-scale content-aware is designed,which can dynamically generate adaptive sampling kernels for each position in the feature map,enabling the model to locate the approxi-mate range of manipulated regions at a coarse-grained level and identify manipulated edge features at a fine-grained level.Secondly,a depthwise separable convolutional decoder is employed to replace the traditional multilayer perceptron for prediction,further en-hancing detection accuracy.Finally,a joint loss function combining binary cross-entropy loss and Dice loss is proposed,effectively improving the model's robustness and generalization capabilities.Cross-dataset experimental results on multiple public datasets dem-onstrate that,in terms of Pixel-level F1 score,the proposed method achieves 69.4%,25.4%,37.8%,and 83.2%on the CASIAv1,De-facto-12k,Coverage,and Columbia datasets,respectively.Compared to the mainstream MVSS-Net++and the latest IML-ViT,it achieves average improvements of 10.2%and 2.1%,respectively,significantly enhancing the accuracy of image tampering localiza-tion.

张雷;王宝丽;陆晓栋;闫成梁;常敏慧

运城学院 数学与信息技术学院,山西 运城 044000运城学院 数学与信息技术学院,山西 运城 044000运城学院 数学与信息技术学院,山西 运城 044000太原师范学院 计算机科学与技术学院,山西 晋中 030619运城学院 数学与信息技术学院,山西 运城 044000

信息技术与安全科学

特征融合深度可分离卷积Vision Transformer联合损失

feature fusiondepthwise separable convolutionVision Transformerjoint loss

《山西大学学报(自然科学版)》 2026 (2)

209-219,11

国家自然科学基金(61703363)山西省基础研究计划项目(202403021221206)数据挖掘与工业智能应用科研创新团队资助项目(YCXYTD-202402)运城学院科研项目(YQ-2020021)

10.13451/j.sxu.ns.2025093

评论