基于对抗混合专家后训练机制的鲁棒AI生成图像检测方法OA
Adversarial Mixture of Experts Post-Training for Robust AI-Generated Image Detection
AI生成图像技术(AI-Generated Images,AIGI)技术实现了高质量视觉内容的自动化生产,在艺术创作、数字娱乐及虚拟现实等领域展现出巨大的应用潜力.然而,该技术在赋能内容生产的同时,也带来了严峻的安全与伦理挑战.生成模型可能被恶意用于伪造真实人物或事件,进而制造虚假信息、传播深度伪造内容,甚至干扰网络舆论.因此,如何有效识别AI生成图像(AIGI检测),已成为保障数字内容可信性和维护网络空间安全的重要研究课题.然而,现有的AIGI检测器在面对对抗攻击时普遍表现出鲁棒性不足的问题,攻击者仅需向合成图像中添加人眼难以察觉的细微对抗扰动,即可使其绕过检测,导致合成内容被误判为真实图像,且对于此类攻击的防御机制仍鲜有研究.针对该问题,本文首先系统评估了对抗训练在AIGI检测任务上的有效性.理论分析与实验结果表明,其在训练过程中易诱发特征纠缠现象,进而导致检测性能严重退化甚至崩塌.鉴于此,亟需发展一种针对AIGI检测任务有效的专用对抗防御方法.与对抗训练中出现的特征纠缠不同,本文发现在标准训练的检测器中,对抗扰动会致使对抗样本在特征空间中的表示明显偏离于干净样本,从而形成显著的可分离性.基于该观察,本文提出将对抗样本视作独立类别进行建模的策略,并构建了一种后训练防御框架:在保持预训练特征提取器固定的前提下,仅通过学习新的分类边界以拟合对抗样本的特征分布.为增强模型对未知攻击的泛化能力,本文进一步提出一种对抗混合专家后训练机制.该机制利用多个专家模块分别学习特定攻击类型的特征模式,并引入共享专家以捕捉不同攻击间的共性表征,从而实现对多类对抗样本的高效建模与鲁棒识别.实验结果表明,本文方法在ProGAN和Stable Diffusion等主流AIGI数据集上,面对多种典型对抗攻击方式,在不牺牲良性样本检测精度的前提下,其平均对抗准确率相较现有主流防御方法分别提升了18.92%与12.56%,展现出良好的实用性与在实际安全场景中的应用潜力.
AI-generated imagery(AIGI)technology has enabled the automated production of high-quality visual con-tent,demonstrating enormous application potential in fields such as artistic creation,digital entertainment,and virtual reali-ty.However,while empowering content production,this technology also brings serious security and ethical challenges.Gen-erative models can be maliciously used to forge real people or events,thereby creating false information,spreading deep-fake content,and even interfering with online public opinion.Therefore,how to effectively identify AI-generated images(AIGI detection)has become an important research topic for ensuring the credibility of digital content and maintaining cy-berspace security.However,existing AIGI detectors generally exhibit insufficient robustness against adversarial attacks.At-tackers only need to add subtle adversarial perturbations imperceptible to the human eye to the synthesized image to bypass detection,causing the synthesized content to be misclassified as a real image,and defense mechanisms against such attacks are still scarce.To address this issue,this paper first systematically evaluates the effectiveness of adversarial training in AI-GI detection tasks.Theoretical analysis and experimental results show that it is prone to inducing feature entanglement dur-ing training,leading to severe degradation or even collapse of detection performance.Therefore,there is an urgent need to develop a dedicated adversarial defense method effective for AIGI detection tasks.Unlike feature entanglement that occurs in adversarial training,this paper finds that adversarial perturbations in standard-trained detectors cause adversarial exam-ples to deviate significantly from clean examples in the feature space,resulting in significant separability.Based on this ob-servation,this paper proposes a strategy of modeling adversarial examples as independent categories and constructs a post-training defense framework:while keeping the pre-trained feature extractor fixed,it only learns new classification boundar-ies to fit the feature distribution of adversarial examples.To enhance the model's generalization ability to unknown attacks,this paper further proposes an adversarial hybrid expert post-training mechanism.This mechanism utilizes multiple expert modules to learn feature patterns for specific attack types and introduces shared experts to capture common representations among different attacks,thereby achieving efficient modeling and robust identification of multiple classes of adversarial ex-amples.Experimental results show that on mainstream AIGI datasets such as ProGAN and Stable Diffusion,facing various typical adversarial attack methods,the average adversarial accuracy is improved by 18.92%and 12.56%respectively com-pared to existing mainstream defense methods without sacrificing the detection accuracy of benign examples,demonstrating good practicality and application potential in real-world security scenarios.
张睿萱;刁云峰;陆智远;夏海峰;郭治卿;郝孝帅;汪萌
合肥工业大学计算机与信息学院,安徽 合肥 230601合肥工业大学计算机与信息学院,安徽 合肥 230601合肥工业大学计算机与信息学院,安徽 合肥 230601中山大学网络空间安全学院,广东 深圳 518107新疆大学计算机科学与技术学院,新疆 乌鲁木齐 830017小米汽车,北京 100085合肥工业大学计算机与信息学院,安徽 合肥 230601
信息技术与安全科学
AI生成图像检测对抗样本对抗攻击对抗防御混合专家后训练策略
AI-generated image detectionadversarial exampleadversarial attacksadversarial defensemixture of expertspost-training strategy
《电子学报》 2026 (3)
1178-1193,16
国家自然科学基金(No.62302139,No.62406068,No.62302427)中央高校基本科研业务费专项资金项目(No.JZ2025HGTB0227) National Natural Science Foundation of China(No.62302139,No.62406068,No.62302427)Fundamental Research Funds for the Central Universities(No.JZ2025HGTB0227)
评论