具有指数图信息通信的大规模情境多智能体强化学习OA
Large-scale Episodic Multi-agent Reinforcement Learning With Exponential Graph Information Communication
多智能体强化学习(MARL)在协同任务中展现出卓越的性能.然而,在具有复杂交互关系的大规模多智能体系统(MAS)中,传统的 MARL 算法由于缺乏高效的通信机制,性能往往受到限制.为提升 MARL 在大规模 MAS 中的性能,本文提出一种具有指数图信息通信的情境 MARL 算法(EMAGIC).首先,设计基于单点指数图的通信拓扑结构,每个智能体在每个时间步仅与一个智能体进行通信,消息通过循环通信链路传递给所有智能体.其次,构建图信息通信机制,利用门控循环单元编码多个时间步的消息,并通过最大化同一时间步不同智能体间消息的互信息来优化消息的编码特征.最后,构建独立情境记忆(EM)模块,建立平均回报与全局状态的对应关系以构建记忆库,利用 EM 目标与个体价值均值的误差来构建损失函数.在多个大规模多智能体环境中的实验结果表明,EMAGIC始终优于最先进的MARL基线方法.
Multi-agent reinforcement learning(MARL)demonstrates excellent performance in cooperative tasks.However,in large-scale multi-agent systems(MAS)with complex interaction relationships,traditional MARL al-gorithms perform poorly due to the lack of efficient communication mechanisms.To enhance the performance of MARL in large-scale MAS,this paper proposes an episodic MARL algorithm with exponential graph information communication(EMAGIC).First,this paper designs a one-peer exponential graph-based communication topology,where each agent communicates with only one other agent at each time step and transmits messages to all agents through a cyclic communication link.Second,this paper constructs a graph information communication mechanism that uses a gated recurrent unit to encode messages across multiple time steps and optimizes the encoded features of the messages by maximizing the mutual information between messages of different agents at the same time step.Fi-nally,this paper builds an independent episodic memory(EM)module to establish the correspondence between av-erage returns and global states for constructing a memory bank,and constructs the loss function by using the error between the EM target and the mean of individual values.Experimental results in multiple large-scale multi-agent environments show that EMAGIC consistently outperforms advanced MARL baseline methods.
李方昱;刘金溢;孙浩源;韩红桂
北京工业大学信息科学技术学院 北京 100124||数字社区教育部工程研究中心 北京 100124北京工业大学信息科学技术学院 北京 100124||数字社区教育部工程研究中心 北京 100124北京工业大学信息科学技术学院 北京 100124||数字社区教育部工程研究中心 北京 100124北京工业大学信息科学技术学院 北京 100124||数字社区教育部工程研究中心 北京 100124
多智能体系统强化学习多智能体通信指数图
multi-agent systemsreinforcement learningmulti-agent communicationexponential graph
《自动化学报》 2026 (6)
1304-1318,15
国家重点研发计划(2023YFB3307300),国家自然科学基金(62373014,92467205,62522302,62473011),北京市科技新星项目(20250484938),北京市自然科学基金-小米创新联合基金(L253010)资助 Supported by National Key Research and Development Pro-gram of China(2023YFB3307300),National Natural Science Foundation of China(62373014,92467205,62522302,62473011),Beijing Nova Program of Science and Technology(20250484938),and Beijing Natural Science Foundation-Xiaomi Innovation Joint Fund(L253010)
评论