基于深度强化学习的大型高铁站到发线运用方案快速编制方法OA
A fast approach for station track utilization at large high-speed railway station based on deep reinforcement learning
大型高铁站是高铁网络中的关键枢纽节点,其到发线运用方案的合理安排是车站运输组织工作的核心.为实现大型高铁站到发线运用方案的快速编制,将人工智能领域的深度强化学习方法应用于到发线运用问题中.首先,基于列车过站径路的概念来实现到发线分配与进路排列的协同优化,将到发线运用问题描述为最大独立集问题,以尽可能安排所有列车为目标,形成整数规划模型.然后,构建基于冲突图的强化学习环境模型,将到发线运用方案数学模型转化为马尔可夫决策过程,定义状态表征、动作空间、状态转移和奖励函数,为智能体提供训练数据的交互采集场所.最后,设计基于GCN(graph convolutional network)的到发线运用智能体神经网络架构,基于PPO(proximal policy optimization)算法实现智能体高效预训练,在决策阶段,设计基于2Imp的决策提升算法进一步优化方案编制质量.以北京南高速场和广州南站为案例场景,对比3种不同方法的求解时间与最终目标函数取值.结果表明,经过4.2 h预训练后智能体完成收敛,决策阶段智能算法可以快速求得最优解,相比于运筹学经典算法,求解质量相同的情况下智能算法求解速度可提升5~16倍,1.2 s即可完成广州南站1 133列列车的到发线运用方案编制.此外,该方法具有良好的扩展性和泛化能力,智能体经过一次训练后即可泛化至不同车站、不同规模的场景中进行决策与求解.研究结果可为进一步提升列车运行图整体编制效率提供技术支撑,为铁路运输组织的智能化发展提供方法参考.
Large high-speed railway stations serve as critical hub nodes in high-speed rail networks,where the rational arrangement of station track utilization plans constitutes the core of station transportation organization.To enable rapid generation of track utilization schemes for large high-speed stations,this study proposed applying deep reinforcement learning from artificial intelligence to address track utilization problems.Firstly,based on the concept of train passing routes,this paper achieved collaborative optimization of track allocation and route arrangement.The track utilization problem was formulated as a maximum independent set problem with the objective of accommodating all trains,leading to an integer programming model.Subsequently,a conflict-graph-based reinforcement learning environment model was constructed to transform the mathematical model of track utilization into a Markov Decision Process.This paper defined the state representation,action space,state transition,and reward function to establish an interactive data collection framework for agent training.Finally,a graph convolutional network(GCN)-based intelligent agent architecture was designed,combined with proximal policy optimization(PPO)algorithm for efficient pre-training.During the decision phase,a 2Imp-based decision enhancement algorithm was proposed to improve solution quality.The case studies of Beijing South High-Speed Field and Guangzhou South Station were conducted to compare the solving time and final objective function values of three different methods.The results demonstrate that after 4.2 hours of pre-training,the agent achieves convergence and rapidly obtains optimal solutions during the decision phase.Compared to classical operations research algorithms,the proposed intelligent algorithm can achieve equivalent solution quality with 5 to 16 times faster,completing the track utilization planning for 1 133 trains at Guangzhou South Station within 1.2 seconds.Moreover,the proposed method can exhibit strong scalability and transferability,enabling the agent to be directly deployed to stations of different scales and configurations for decision-making and problem-solving after a single training phase.These findings can provide technical support for enhancing the overall efficiency of train timetable preparation.The results can offer methodological references for the intelligent development of railway transportation organization.
王岩;周黎;范家铭;徐辉章;张新;李博
中国铁道科学研究院,北京 100081||中国铁路列车运行图技术中心,北京 100081||中国铁道科学研究院集团有限公司 运输及经济研究所,北京 100081中国铁道学会,北京 100081中国铁路列车运行图技术中心,北京 100081||中国铁道科学研究院集团有限公司 运输及经济研究所,北京 100081中国铁路列车运行图技术中心,北京 100081||中国铁道科学研究院集团有限公司 运输及经济研究所,北京 100081中国铁路列车运行图技术中心,北京 100081||中国铁道科学研究院集团有限公司 运输及经济研究所,北京 100081中国铁路列车运行图技术中心,北京 100081||中国铁道科学研究院集团有限公司 运输及经济研究所,北京 100081
交通工程
大型高铁站到发线运用深度强化学习智能体PPO算法神经网络
large high-speed railway stationstation track utilizationdeep reinforcement learningagentPPO algorithmneural networks
《铁道科学与工程学报》 2026 (7)
3099-3109,11
中国国家铁路集团有限公司科技研究开发计划课题(P2024X002)中国铁道科学研究院集团有限公司科研项目(2024YJ154)
评论