基于注意力增强DDPG的短波协同测向与定位方法OA
A Cooperative Direction-Finding and Localization Method for Shortwave Based on Attention-Enhanced DDPG
复杂场景下的短波定位在战场通信中发挥着至关重要的作用,而现有基于深度学习的短波智能测向方法主要依赖人工标注且优质短波信号样本较少.针对上述问题,在短波信号空时高分辨时频图基础上,提出了一种基于注意力增强深度确定性策略梯度(DDPG)算法的短波协同测向与定位方法.该方法首先设计了基于空时高分辨时频图的强化学习环境和具有混合动作空间的智能体,再依据估计定位结果设计奖励函数,以此建立可观测马尔可夫决策过程.其次,利用注意力机制增强深度强化学习的多层次非精细化评估,实现自动测向定位,可以减少人工标注,实现边工作边学习的自主进化,逐步提升短波信号的智能测向定位能力.实验结果表明,所提算法在保证定位精度与参数估计算法性能相当的情况下,使得测向定位时间缩短了77.7%.
Shortwave localization in complex scenarios plays a crucial role in battlefield communica-tions.Existing deep-learning-based intelligent shortwave direction-finding methods mainly rely on manual annotation,and there is a scarcity of high-quality shortwave signal samples.To address the aforementioned challenges,a shortwave autonomous direction-finding and localization method is pro-posed,based on a space-time high-resolution time-frequency representation of shortwave signals and an attention-enhanced deep deterministic policy gradient(DDPG)algorithm.A reinforcement learning environment is first designed based on space-time high-resolution time-frequency representation,along with an agent possessing a hybrid action space.Then,a reward function is designed based on the esti-mated localization results,thereby establishing an observable Markov decision process.Subsequently,by enhancing deep reinforcement learning with an attention mechanism for multi-level non-fine evalua-tion,automatic direction finding and positioning are achieved.This design reduces the need for manual labeling and enables autonomous evolution through simultaneous work and learning,gradually improv-ing the intelligent direction-finding and localization capabilities for shortwave signals.Experimental results show that,comparable localization accuracy and parameter estimation performance are main-tained,while the direction-finding and localization time is improved by approximately 77.7%.
冯祺玥;唐涛;张昀普;赵排航;陈晓云
信息工程大学,河南 郑州 450001信息工程大学,河南 郑州 450001信息工程大学,河南 郑州 450001信息工程大学,河南 郑州 450001信息工程大学,河南 郑州 450001
信息技术与安全科学
测向定位马尔可夫决策过程空时高分辨混合动作空间深度确定性策略梯度算法
direction findingMarkov decision processspace-time high-resolutionhybrid action spacedeep deterministic policy gradient algorithm
《信息工程大学学报》 2026 (3)
259-266,274,9
评论