首页|期刊导航|北京交通大学学报|网联环境下基于PPO的车队轨迹与信号配时协同优化方法

网联环境下基于PPO的车队轨迹与信号配时协同优化方法OA

Collaborative optimization method for platoon trajectory and signal timing based on proximal policy optimization in connected vehicle environments

中文摘要英文摘要

既有基于深度强化学习的车辆轨迹优化方案在处理网联混行交通时,难以实现车队的动态合并与拆分,而传统的信号配时优化以周期或相位为单位,存在响应滞后问题,二者间的割裂控制限制了交叉口通行效率的进一步提升.针对车辆轨迹灵活控制与信号配时实时响应的协同优化需求,提出一种基于近端策略优化(Proximal Policy Optimization,PPO)的车队轨迹与信号配时协同优化方法(Fleet Trajectory and Signal timing based on Waiting Offsetting,FT-SIWO).该方法首先将车辆轨迹优化问题构建为一个马尔可夫决策过程,设计了一个融合局部车辆动态与全局交通状态变量的状态空间、一个包含跟车与多种加速度模式的动作空间及一个综合了驾驶舒适度、安全性等多维度的奖励函数,使智能体能够全面感知环境并学习高效策略.然后,采用PPO-clip算法进行策略更新,以提升训练过程的稳定性与样本效率.最后,引入一种基于候时对冲(Waiting Offset-ting,WO)的信号配时实时优化机制,通过动态计算并平衡绿灯相位车辆的"时间亏损"与红灯相位车辆的"时间盈余",实现以秒为粒度的绿灯时长动态调整,从而与车辆轨迹优化形成深度协同.研究结果表明:在SUMO与Matlab联合仿真平台上,相较于无优化的基准方案,FT-SIWO使交叉口车辆平均延误减少 28%~32%,平均等待时间降低 30%~35%,同时多种污染物(CO2,CO,HC,NOx,PMx)排放总量降低34%~38%,燃油消耗减少19%;与国内外同类先进的分布式控制或车路协同方案相比,FT-SIWO在控制成效上进一步提升了6%~10%.该方法通过车队智能编组与信号配时的秒级联动,有效提升了绿灯时间利用率和交通流整体运行效率,为智慧城市交通管理系统应对混合交通流、缓解拥堵与降低排放提供了可行的协同优化解决方案.

Existing vehicle trajectory optimization methods based on deep reinforcement learning struggle to achieve the dynamic merging and splitting of platoons in connected,mixed-traffic environ-ments.Simultaneously,conventional signal timing optimization,which operates at the cycle or phase level,is hindered by response lag.The decoupled control between these two elements restricts further improvements in intersection efficiency.To address the necessity for the collaborative optimization of flexible vehicle trajectory control and real-time signal timing,this paper proposes a joint optimization method for platoon trajectories and signal timing based on Proximal Policy Optimization(PPO),termed FT-SIWO(Fleet Trajectory and Signal timing based on Waiting Offsetting).First,the vehicle trajectory optimization problem is formulated as a markov decision process.A state space integrating local vehicle dynamics and global traffic state variables,an action space encompassing car-following and multiple acceleration modes,and a multi-dimensional reward function considering factors such as driving comfort and safety are designed.This formulation enables the agent to comprehensively per-ceive the environment and learn efficient policies.Subsequently,the PPO-clip algorithm is employed for policy updates to enhance both training stability and sample efficiency.Finally,a real-time signal timing optimization mechanism based on Waiting Offsetting(WO)is introduced.By dynamically calcu-lating and balancing the"time loss"of vehicles in the green phase and the"time surplus"of vehicles in the red phase,this mechanism enables the dynamic adjustment of the green light duration at a second-level granularity,thereby achieving deep coordination with the vehicle trajectory optimization.Simula-tion results on a joint SUMO and Matlab platform demonstrate that,compared to a non-optimized baseline scheme,FT-SIWO reduces the average vehicle delay at intersections by 28%-32%and the average wait-ing time by 30%-35%.Additionally,it decreases the total emissions of various pollutants(CO2,CO,HC,NOx,and PMx)by 34%-38%and reduces fuel consumption by 19%.Furthermore,when com-pared to state-of-the-art distributed control or vehicle-infrastructure cooperation schemes,FT-SIWO achieves an additional 6%-10%improvement in control effectiveness.By enabling intelligent platoon formation and a second-level linkage with signal timing,this method effectively enhances green time utilization and the overall operational efficiency of traffic flow.It provides a viable collaborative optimi-zation solution for smart urban traffic management systems to handle mixed traffic flows,alleviate con-gestion,and reduce emissions.

邢凯然

北京交通大学 交通运输学院,北京 100044

交通工程

智能交通协同优化车队轨迹信号配时近端策略优化

intelligent transportationcollaborative optimizationplatoon trajectorysignal timingproximal policy optimization

《北京交通大学学报》 2026 (3)

152-163,12

国家重点研发计划(2022YFB4301305) National Key R&D Plan(2022YFB4301305)

10.11860/j.issn.1673-0291.20250168

评论