An Efficient Multi-Agent Policy Self-Play Learning Method Aiming at Seize-Control ScenariosOA
An Efficient Multi-Agent Policy Self-Play Learning Method Aiming at Seize-Control Scenarios
Huaqing Zhang;Hongbin Ma;Xiaofei Zhang;Li Wang;Minglei Han;Hui Chen;Ao Ding
School of Automation,Beijing Institute of Technology,Beijing 100081,P.R.ChinaSchool of Automation,Beijing Institute of Technology,Beijing 100081,P.R.ChinaSchool of Vehicle and Mobility,Tsinghua University,Beijing 100084,P.R.ChinaSchool of Mechanical Engineering,Beijing Institute of Technology,Beijing 100081,P.R.ChinaSchool of Automation,Beijing Institute of Technology,Beijing 100081,P.R.ChinaSchool of Automation,Beijing Institute of Technology,Beijing 100081,P.R.ChinaSchool of Automation,Beijing Institute of Technology,Beijing 100081,P.R.China
Self-playcooperative confrontationdeep reinforcement learningpolicy evaluationwargame
Self-playcooperative confrontationdeep reinforcement learningpolicy evaluationwargame
《无人系统(英文)》 2025 (4)
987-1004,18
This work was partially funded by the National Key Re-search and Development Plan of China(No.2018AAA0101000)and the National Natural Science Foundation of China under grant 62076028.This work is also funded by Innovation Fund of Qiyuan Lab.
评论