首页|期刊导航|佛山科学技术学院学报(自然科学版)|面向人-物交互场景的图像生成框架研究

面向人-物交互场景的图像生成框架研究OA

Research on image generation framework for human-object interaction scenarios

中文摘要英文摘要

针对人-物交互文本图像生成方法在建模细粒度交互关系和处理前景背景融合方面存在明显不足问题,提出一个面向人-物交互场景的图像生成框架.首先,引入文本语义重构机制,利用大语言模型对原始描述进行结构化解析与阶段性重组,从而提升属性表达与交互语义之间的对齐效果.随后,设计主客体先验融合模块,结合文本语义信息与对象几何先验,对主体与客体的掩码与潜在表示进行联合建模,进一步增强交互关系建模能力与结构一致性.实验结果表明,所提文方法在以 T2I-CompBench 为基础构建的交互文本数据集上取得了优异的性能,在 CLIP-Score 与 Verb CLIP-Score 指标上分别达到 28.41 和 21.06,基于用户偏好的 PickScore 得分为 20.22,验证了所提方法在生成质量与交互表达方面的有效性.

Existing Human-Object Interaction(HOI)text-to-image methods remain limited in modeling fine-grained interactions and achieving coherent foreground-background fusion.To address this,we propose a HOI-oriented image generation framework built upon a three-stage architecture.First,a text semantic reconstruction mechanism is introduced,where a large language model performs structured parsing and stage-wise reorganization of the input description to enhance alignment between attribute expression and interaction semantics.Next,a subject-object prior fusion module jointly models masks and latent representations of subjects and objects by integrating textual semantics with geometric priors,improving interaction modeling and structural consistency.Experiments on an interaction dataset constructed from T2I-CompBench demonstrate superior performance,achieving CLIP-Score 28.41,Verb CLIP-Score 21.06,and a user-preference-based PickScore of 20.22,validating the effectiveness of the proposed method in both generation quality and interaction representation.

周燕;陈维德;彭俊轩;刘翔宇;周月霞

佛山大学 电子信息工程学院,广东 佛山 528225佛山大学 计算机与人工智能学院,广东 佛山 528225佛山大学 计算机与人工智能学院,广东 佛山 528225佛山大学 电子信息工程学院,广东 佛山 528225佛山大学 电子信息工程学院,广东 佛山 528225

信息技术与安全科学

人-物交互文本语义重构大语言模型先验融合图像生成

human-object interactiontext semantic reconstructionlarge language modelprior fusionimage generation

《佛山科学技术学院学报(自然科学版)》 2026 (3)

33-40,8

国家自然科学基金资助项目(61972091)

评论