基于离线强化学习的自动协商方法OA
An Offline Reinforcement Learning Based Automated Negotiation Approach
自动协商是实现多智能体系统中合作与协作的关键方式.尽管基于强化学习(Reinforcement Learning,RL)的协商智能体在各种场景中取得了显著的成功,但由于实现环境的限制,它们仍面临挑战.特别是这些智能体需要与对手进行大量的在线互动以进行训练,这在现实世界的应用中往往是不可行和不现实的.因此,需要一种新的方法,以便直接从离线数据集中学习有效的协商策略.此外,在随后的在线协商中,对手可能会因为各种原因(如风险态度的变化或市场条件的变化)而改变其策略.这些因素为自动协商带来了重大挑战.提出了一种新的协商智能体,通过离线—在线RL来提高协商智能体的能力.它使协商智能体能够:① 利用基于 RL 方法的策略与对手交互,提高动态协商环境的适应能力;② 从历史离线数据中学习协商策略,无需大量的在线交互;③ 在在线微调优化过程中,使得学习到的策略快速且稳定地提升性能.基于多种协商场景和最近自动化协商智能体竞赛(Automated Negotiating Agents Competitions,ANAC)中的获胜智能体,提供了广泛的实验结果.结果显示,该智能体的表现超过了最先进的智能体,并且即使对手转换到不同策略时仍保持有效.
Automated negotiation is a key approach to achieving cooperation and collaboration in multi-agent systems.While Rein-forcement Learning(RL)-based negotiating agents have attained remarkable success across various scenarios,they still face constraints imposed by real-world implementation environments.In particular,these agents require extensive online interactions with opponents for training,which is often infeasible and unrealistic in practical applications.Therefore,a novel method is needed to enable the learning of effective negotiation strategies directly from offline datasets.Additionally,during subsequent online negotiations,opponents may al-ter their strategies due to various factors—such as changes in risk attitudes or market conditions—posing significant challenges to auto-mated negotiation.In this work,we propose a new negotiating agent that enhances performance via offline-to-online RL.The proposed agent is able ① to interact with opponents using an RL-based strategy to improve its adaptability to dynamic negotiation environments;② to learn negotiation strategies from historical offline data without the need for extensive online active interactions;and ③ to optimize the online fine-tuning process to facilitate rapid and stable performance improvements of the pre-learned offline strategies.Extensive experi-mental results are presented,based on multiple negotiation scenarios and winning agents from recent Automated Negotiating Agents Com-petitions(ANAC).The results demonstrate that the proposed agent outperforms state-of-the-art alternatives and remains effective even when opponents switch to different strategies.
陈锶奇;熊钊远;汪云飞;王昊杨
重庆交通大学 信息科学与工程学院,重庆 400074重庆交通大学 交通运输学院,重庆 400074重庆交通大学 信息科学与工程学院,重庆 400074重庆交通大学 信息科学与工程学院,重庆 400074
信息技术与安全科学
自动协商深度强化学习智能体电子商务
automated negotiationdeep reinforcement learningagente-commerce
《无线电通信技术》 2026 (1)
1-14,14
国家自然科学基金(61602391)天津市科技计划项目(22JCZDJC00580)National Natural Science Foundation of China(61602391)Tianjin Science and Technology Plan Project(22JCZDJC00580)
评论