首页|期刊导航|佛山科学技术学院学报(自然科学版)|面向多主体语义解耦的图像生成框架研究

面向多主体语义解耦的图像生成框架研究OA

Research on an image generation framework for multi-entity semantic disentanglement

中文摘要英文摘要

针对扩散模型在多主体多属性文本引导生成图像中存在的语义混淆、属性错配及一致性不足等问题,提出一种基于语义引导的多主体图像生成方法.首先,设计文本图像预处理机制,通过词性分割与参考图像潜在表示融合,有效提升模型对多实体属性的解耦能力;其次,引入分层交叉注意力机制,结合区域掩码与实体文本对,实现语义与空间的精确对齐;最后,构建双分支并行去噪结构,联合布局图的结构先验与文本语义,增强生成图像的结构一致性与语义稳定性.基于 Stable Diffusion XL,在 T2I-CompBench 基准的颜色、纹理与形状子集上进行实验,采用 BLIP-VQA 与 VQAScore 两类指标评估语义一致性,分别达到 69.79%、58.43%、52.97%和80.74%,较现有方法在四项指标上提升 1.40%、1.09%、2.05%和 2.04%,验证了所提方法在复杂图文条件下的生成一致性与可控性.

To address issues such as semantic confusion,attribute mismatch,and insufficient consistency in multi-subject and multi-attribute text-guided generation image with diffusion models,this paper proposes a semantic-guided multi-subject image generation method.First,a text-image preprocessing mechanism is designed,which enhances the model's ability to disentangle multiple entities and their attributes through part-of-speech segmentation and fusion with latent representations of reference images.Second,a hierarchical cross-attention mechanism is introduced,combining regional masks with entity-specific text to achieve precise semantic and spatial alignment.Finally,a dual-branch parallel denoising architecture is constructed,integrating structural priors from layout maps with textual semantics to improve structural consistency and semantic stability in the generated images.Experiments are conducted on the color,texture,and shape subsets of the T2I-CompBench benchmark using Stable Diffusion XL.Semantic consistency is evaluated using BLIP-VQA and VQAScore metrics,achieving scores of 69.79%,58.43%,52.97%,and 80.74%,respectively—representing improvements of 1.40%,1.09%,2.05%,and 2.04%over existing methods—demonstrating the effectiveness of the proposed approach in generating consistent and controllable images under complex text-image conditions.

彭俊轩;周燕;陈维德;刘翔宇;周月霞

佛山大学 计算机与人工智能学院,广东 佛山 528225佛山大学 电子信息工程学院,广东 佛山 528225佛山大学 计算机与人工智能学院,广东 佛山 528225佛山大学 电子信息工程学院,广东 佛山 528225佛山大学 电子信息工程学院,广东 佛山 528225

信息技术与安全科学

扩散模型语义属性绑定交叉注意力图像生成

diffusion modelsemantic attribute bindingcross-attentionimage generation

《佛山科学技术学院学报(自然科学版)》 2026 (3)

41-50,10

国家自然科学基金资助项目(61972091)

评论