首页|期刊导航|交通信息与安全|公交协同优化的时空Transformer-STGNN混合强化学习方法

公交协同优化的时空Transformer-STGNN混合强化学习方法OA

A Spatio-temporal Transformer-STGNN Hybrid Reinforcement Learning Method for Bus Cooperative Optimization

中文摘要英文摘要

为解决公交动态越站与驻站多策略协同中局部决策引发全局连锁延误及多重目标冲突的难题,研究了1种融合Transformer与时空图神经网络的动态协同优化模型.构建双向交互的感知与预知时空特征提取架构.利用时空变换器的空间注意力机制提取公交站点的动态依赖关系,生成自适应权重矩阵并作为动态邻接矩阵输入时空图神经网络.该网络结合图卷积网络与因果扩张卷积,预测客流积压与延误传播风险,并将预知风险评分反向传递至时序注意力层.此闭环反馈机制通过风险驱动调整权重,量化了单一调度决策产生的时空非线性连锁效应.设计了1种嵌套改进遗传算法与深度双Q网络的混合强化学习框架.利用遗传算法的全局广度搜索生成帕累托前沿解集,将其作为深度双Q网络的初始策略空间;同时,在多目标奖励函数中引入时空图神经网络的风险评分作为安全约束惩罚,以此权衡车辆运行效率与乘客出行成本,输出动态越站与驻站协同控制策略.基于佛山市101路公交线路高峰时段运营数据开展仿真实验.结果表明,该策略控制了车头时距波动,避免了公交车辆的连续串车现象.在乘客平均等待时间增加0.83%的前提下,乘客平均在途时间减少约24.7%,同时缩短了车辆整体在途时间并提升了站间行驶速度.相比传统单一深度强化学习算法,该混合算法具备更少的迭代次数与更小的寻优误差.适用于高频发车的城市干线公交实时协同调度,能在控制时空风险传播的前提下实现多调度策略的动态平衡.

To address multi-objective conflicts and global ripple delays caused by local decisions,a dynamic cooper-ative optimization model is investigated.This model integrates a Spatio-Temporal Transformer and a spatiotemporal graph neural network(STGNN).A bidirectional interactive architecture for the perception and anticipation of spatio-temporal features is constructed.Dynamic dependencies of bus stops are extracted using the spatial attention mecha-nism of the Transformer.An adaptive weight matrix is generated and subsequently inputted into the STGNN as a dy-namic adjacency matrix.This network combines graph convolutional networks and causal dilated convolutions to predict risks of passenger accumulation and delay propagation.The anticipated risk scores are reversely transmitted to the temporal attention layer.Driven by risks,this closed-loop feedback mechanism adjusts the weights.Thereby,the nonlinear ripple effects in space and time generated by single scheduling decisions are quantified.A hybrid rein-forcement learning framework is designed.An improved genetic algorithm(GA)and a Deep Double Q-Network(DDQN)are nested in this framework.A Pareto front solution set is generated using the global broad search of the GA.This set serves as the initial strategy space for the DDQN.Simultaneously,the risk scores from the STGNN are introduced into the multi-objective reward function as safety constraint penalties.Consequently,the operational effi-ciency of vehicles and the travel costs of passengers are balanced.Thus,optimal coordinated control strategies for stop-skipping and holding are generated.Simulation experiments are conducted based on the operational data of Fos-han Bus Route 101 during peak hours.The results indicate that the fluctuations in headways are controlled by this strategy.Furthermore,the continuous bus bunching phenomenon is successfully avoided.Under the premise of a 0.83%increase in the average passenger waiting time,the average passenger in-vehicle time is reduced by approxi-mately 24.7%.Meanwhile,the overall vehicle travel time is shortened,and the driving speed between stops is im-proved.Compared with traditional single deep reinforcement learning algorithms,this hybrid algorithm exhibits fewer iteration counts and smaller optimization errors.The proposed model is applicable to the real-time coopera-tive scheduling for urban trunk buses with high-frequency departures.A dynamic balance of multiple scheduling strategies is achieved under the premise of controlling risk propagation.

张韫博;周雪梅;王沛钰;徐傲;戴咏奇

同济大学交通学院 上海 201804||同济大学道路与交通工程教育部重点实验室 上海 201804同济大学交通学院 上海 201804||同济大学道路与交通工程教育部重点实验室 上海 201804同济大学交通学院 上海 201804||同济大学道路与交通工程教育部重点实验室 上海 201804同济大学交通学院 上海 201804||同济大学道路与交通工程教育部重点实验室 上海 201804同济大学交通学院 上海 201804||同济大学道路与交通工程教育部重点实验室 上海 201804

交通工程

智能交通公交运行优化多策略协同混合强化学习时空图神经网络

intelligent transportationbus operation optimizationmulti-strategy coordinationhybrid reinforcement learningspatiotemporal graph neural network

《交通信息与安全》 2026 (1)

101-112,12

国家自然科学基金面上项目(52372318)资助

10.3963/j.jssn.1674-4861.2026.01.009

评论