基于视觉先验驱动的可控文本到图像生成技术综述OA
A Survey of Visual Prior-Driven Controllable Text-to-Image Generation Techniques
以扩散模型为代表的文本到图像生成技术虽已取得突破性进展,但单一文本模态难以精准控制图像的空间结构与外观风格,视觉先验驱动的可控生成技术可以有效解决该问题.现围绕从"单一控制"向"多维协同"的技术演进路径,系统梳理了视觉先验驱动下涵盖经典U-Net架构与新兴扩散Transformer(DiT)范式的潜空间可控图像生成体系.在此基础上,构建了"先验注入—特征冲突—协同消解"的核心分析框架:从几何结构、布局定位与外观风格3个维度,解析单一视觉先验的表征与注入机制;并重点论述基于实例感知与动态裁决、注意力机制调控,以及时步调度与特征层解耦等协同消解策略,以解决多条件融合时的结构-布局、布局-外观及结构-外观之间的交互干扰与特征冲突.此外,归纳常用数据集与量化评价体系,深入剖析了多维可控性评估面临的挑战与新兴细粒度度量范式,梳理图生成技术的场景化应用,并前瞻该领域的未来发展趋势.
While text-to-image generation technologies,epitomized by diffusion models,have achieved groundbreaking advancements,relying solely on a single text modality struggles to precisely control the spatial structure and appearance style of images.Visual prior-driven controllable generation technologies can effectively solve this problem.Centering on the technological evolution path from"single control"to"multi-dimensional synergy,"this paper systematically reviews the visual prior-driven controllable image generation system in the latent space,covering both the classic U-Net architecture and the emerging Diffusion Transformer(DiT)paradigm.Building upon this,a core analytical framework of"Prior Injection-Feature Conflict-Synergistic Resolution"is constructed:from the three dimensions of geometric structure,layout positioning,and appearance style,the representation and injection mechanisms of single visual priors are analyzed;furthermore,synergistic resolution strategies based on instance perception and dynamic adjudication,attention mechanism regulation,as well as timestep scheduling and feature layer disentanglement,are critically discussed to resolve the interactive interference and feature conflicts between structure-layout,layout-appearance,and structure-appearance during multi-condition fusion.In addition,common datasets and quantitative evaluation systems are summarized,the challenges and emerging fine-grained metric paradigms facing multi-dimensional controllability assessment are deeply analyzed,the scenario-based applications of image generation technologies are reviewed,and the future development trends of this field are prospected.
林志豪;陈明举;罗扬铭;段雪阳;解晨
四川轻化工大学 自动化与信息工程学院,四川 宜宾 644000||智能感知与控制四川省重点实验室,四川 宜宾 644000四川轻化工大学 自动化与信息工程学院,四川 宜宾 644000||智能感知与控制四川省重点实验室,四川 宜宾 644000四川轻化工大学 自动化与信息工程学院,四川 宜宾 644000||智能感知与控制四川省重点实验室,四川 宜宾 644000四川轻化工大学 自动化与信息工程学院,四川 宜宾 644000||智能感知与控制四川省重点实验室,四川 宜宾 644000四川轻化工大学 自动化与信息工程学院,四川 宜宾 644000||智能感知与控制四川省重点实验室,四川 宜宾 644000
信息技术与安全科学
可控文本到图像生成视觉先验扩散模型多条件融合特征冲突消解
controllable text-to-image generationvisual priorsdiffusion modelsmulti-condition fusionfeature conflict resolution
《四川轻化工大学学报(自然科学版)》 2026 (3)
86-102,17
四川省自然科学基金项目(2025ZNSFSC04772024NSFSC2042)智能感知与控制四川省重点实验室基金项目(2025IPCY02)四川轻化工大学研究生创新基金资助项目(Y2025089)
评论