首页|期刊导航|网络安全与数据治理|基于DPPO的双智能体协同渗透测试方法研究

基于DPPO的双智能体协同渗透测试方法研究OA

Research on dual-agent collaborative penetration testing method based on DPPO

中文摘要英文摘要

基于强化学习的自动化网络渗透测试方法近年来受到广泛关注,但现有研究多局限于中小规模网络环境,难以应对实际网络中主机数量庞大、拓扑复杂的挑战.提出一种基于 DPPO(Dual-agent Proximal Policy Op-timization)的双智能体协同渗透测试框架,通过分工协作机制提升大规模网络环境下的攻击决策效率.该框架包含目标决策智能体(TargetDecider)与攻击执行智能体(AttackExecutor),分别负责目标节点筛选与具体漏洞利用执行,二者通过共享环境状态进行协同,以协作奖励机制引导,从而实现分工合作.实验在基于 CyberBat-tleSim 构建的仿真网络环境中进行,并与传统的 DQN、PPO、SAC 等智能体方法进行了比较,结果表明 DPPO框架攻陷关键目标所需的步骤数更少、波动更小,显示出更高的攻击效率与策略稳定性.同时,其敏感主机攻陷比率维持在较高水平,表明该框架能可靠完成核心渗透任务.在累计奖励方面,DPPO 亦表现出持续优化与稳健收敛的趋势.实验结果表明,该框架能够有效适应大规模静态仿真网络中的渗透测试任务.该框架对动态防御和真实网络环境的适用性仍需在后续研究中进一步验证.

The automated network penetration testing method based on reinforcement learning has received widespread attention in recent years.However,the existing research is mostly limited to small and medium-sized network environments,making it difficult to address the challenges posed by the vast number of hosts and complex topology in real-world networks.This study proposes a dual-agent collaborative pene-tration testing framework based on DPPO(Dual-agent Proximal Policy Optimization),which enhances the efficiency of attack decision-making in large-scale network environments through a division of labor and collaboration mechanism.The framework comprises a target decision-mak-ing agent(TargetDecider)and an attack execution agent(AttackExecutor),responsible for target node selection and specific vulnerability ex-ploitation execution,respectively.The two agents collaborate through shared environmental states,guided by a collaborative reward mecha-nism,to achieve division of labor and cooperation.Experiments were conducted in a simulated network environment built based on CyberBat-tleSim,and compared with traditional agent methods such as DQN,PPO,and SAC.The results show that the DPPO framework requires fe-wer steps and exhibits less fluctuation to compromise key targets,demonstrating higher attack efficiency and strategic stability.At the same time,its sensitive host compromise rate remains at a high level,indicating that the framework can reliably complete core penetration tasks.In terms of cumulative reward,DPPO also shows a trend of continuous optimization and robust convergence.The results indicate that the frame-work can effectively adapt to penetration testing tasks in large-scale static simulated networks,while its applicability to dynamic defense sce-narios and real network environments still requires further validation.

李科;霍朝宾;贺敏超;杨继

华北计算机系统工程研究所,北京 100083华北计算机系统工程研究所,北京 100083华北计算机系统工程研究所,北京 100083华北计算机系统工程研究所,北京 100083

信息技术与安全科学

自动化渗透测试双智能体协作深度强化学习

automated penetration testingdual-agent collaborationdeep reinforcement learning

《网络安全与数据治理》 2026 (6)

1-8,8

10.19358/j.issn.2097-1788.2026.06.001

评论