首页|期刊导航|南京大学学报(自然科学版)|残缺条件下基于生成式轨迹建模的鲁棒模仿学习方法研究

残缺条件下基于生成式轨迹建模的鲁棒模仿学习方法研究OA

Robust imitation learning via generative trajectory modeling from incomplete demonstrations

中文摘要英文摘要

高效利用残缺专家演示数据是当前模仿学习领域的重点之一,残缺的演示意味着局部细节的丢失,会导致误差累积.针对演示残缺的场景,提出一种基于生成式轨迹建模的鲁棒模仿学习方法.通过利用决策变换器在残缺演示条件下生成连续状态-动作序列补全轨迹,并通过设定目标分数控制生成轨迹的质量,将补全后的轨迹输入扩散模型进行轨迹生成.同时,引入新的评分机制,用于评价扩散模型生成的轨迹是否符合专家示范中的环境动力学约束.该模型能在良好地补全残缺轨迹的条件下,生成高奖励并符合动力学约束的轨迹,有效减少了误差累积并提高规划精度.实验表明,该方法优于现有的离线强化学习方法,能解决高维连续空间中的规划-现实不匹配问题,并拥有更好的鲁棒性和泛化能力.

Efficiently leveraging imperfect expert demonstration data is one of the key challenges in the field of imitation learning.Imperfect demonstrations,which entail the loss of fine-grained local details,can lead to error accumulation.To address scenarios with incomplete demonstrations,we propose a robust imitation learning method based on generative trajectory modeling.Our approach utilizes a Decision Transformer to generate continuous state-action sequences conditioned on incomplete demonstrations,thereby completing the trajectories.The quality of these generated trajectories is controlled by setting a target return.The completed trajectories are then fed into a diffusion model for further trajectory generation.Additionally,we introduce a novel scoring mechanism to evaluate whether the trajectories generated by the diffusion model conform to the environmental dynamics constraints present in the expert demonstrations.This model can generate high-reward trajectories that adhere to the dynamics constraints while effectively completing imperfect trajectories,thereby significantly reducing error accumulation and improving planning accuracy.Experiments show that our method outperforms existing offline reinforcement learning approaches.It successfully addresses the planning-to-reality mismatch problem in high-dimensional continuous spaces and demonstrates superior robustness and generalization capabilities.

戴领;陈文韬;谭晓阳

模式分析与机器智能工业和信息化部重点实验室,南京航空航天大学计算机科学与技术学院,南京,211106南京大学软件学院,南京,210093模式分析与机器智能工业和信息化部重点实验室,南京航空航天大学计算机科学与技术学院,南京,211106

信息技术与安全科学

离线强化学习模仿学习扩散模型轨迹补全

offline reinforcement learningimitation learningdiffusion modelstrajectory completion

《南京大学学报(自然科学版)》 2026 (4)

551-561,11

国家自然科学基金(62476128)

10.13232/j.cnki.jnju.2026.04.004

评论