船载多源感知系统及深度学习在其中的应用研究进展OA
Research progress on deep learning in shipborne multi-source perception systems
远洋渔业的高质量发展离不开渔船和装备的智能化升级,集成多源传感器与深度学习算法的船载多源感知系统是实现渔船智能化管理的重要途径.系统论述了船载多源感知系统及深度学习在其中的应用研究进展,首先,详细阐述了船载多源感知系统的硬件组成,包括视频感知单元、化学与环境感知单元及作为核心的多源数据处理终端,并以一套已成功研制的成套装备为例,说明了其作为最小基础系统的工程实现.其次,重点分析了深度学习在系统中的核心技术应用:在视频感知方面,梳理了从两阶段检测器到端到端Transformer模型的目标检测技术演进,以及从双流网络到基于骨架模型的行为识别方法对比;在多模态融合方面,评述了联合表示与协同表示两类主流融合范式的技术原理与典型应用.最后,总结了当前系统面临的专用数据稀缺、环境干扰强、计算资源受限、模型可解释性不足等核心挑战,并对构建高质量数据集、发展自适应感知模型、推动边缘智能与轻量化、增强决策透明性等未来研究方向进行了展望.
Distant-water fisheries are vital to China's marine resource development and food security,with approximately 1 500 vessels operating on the high seas under extreme conditions including heavy fog,rain,snow,and dynamic sea states.Although electronic monitoring(EM)and vessel monitoring systems(VMS)have been widely deployed,they remain limited by single-modal data reliance,insufficient intelligence for automated analysis,and prohibitive satellite communication costs for raw video transmission.To address these challenges,shipborne multi-source perception systems integrating heterogeneous sensors with deep learning algorithms have emerged as a promising solution.This review systematically examines the hardware architecture,core algorithms,and future directions of such systems.The hardware system comprises visible and infrared cameras for all-weather video perception,conductivity-temperature-depth(CTD)sensors for in-situ seawater parameters measurement,and a multi-source data processing terminal serving as the central hub for data aggregation,intelligent fusion,and reliable communication.A complete engineering implementation of"shipborne multi-source perception and AI-integrated intelligent terminal equipment suite"is presented,which achieves edge-based real-time analysis through embedded AI computing platforms and supports multi-modal communication including GNSS and satellite links.In object detection,the technological evolution from two-stage detectors(Faster R-CNN,Cascade R-CNN)to single-stage detectors(YOLO series,SSD),anchor-free methods(CenterNet,FCOS),and end-to-end Transformer architectures(DETR,DyFish-DETR,FMRFT)is reviewed.Empirical comparisons show that YOLOv8 achieves 97.1%mAP@0.5 for vessel detection and 95.7%for fish detection at 62 FPS,significantly outperforming Faster R-CNN(88.4%and 81.2%at 18 FPS).A fundamental cause analysis reveals that two-stage detectors are more vulnerable to hull-induced geometric distortion because their region proposal networks rely on spatial feature consistency,while single-stage detectors perform dense regression directly on feature maps with higher tolerance.In action recognition,the progression from two-stream networks and 3D convolutional networks to temporal segment networks and skeleton-based models is analyzed.Skeleton-based models(ST-GCN,PoseConv3D)demonstrate unique advantages through three synergistic mechanisms:1)Physical-level robustness against illumination and background interference-PoseConv3D maintains 0.3%accuracy loss under 50%keypoint dropout versus 5.5%for GCN;2)Representation-level invariance to global motion from vessel rolling,which can be learned through data augmentation in graph convolutional networks;3)Computational-level efficiency with input dimensions of 50-75 compared to 150 000 for RGB frames,significantly reducing overfitting risk in data-scarce fishery scenarios.In multimodal fusion,joint representation methods(feature-level,model-level,decision-level fusion)and collaborative representation methods are compared,with applications in vessel detection and fish behavior analysis.Two novel paradigms for shipborne scenarios are proposed:physics-informed neural network-guided fusion that embeds the TEOS-10 equation of state as a physical constraint,and environment-adaptive dynamic fusion that adjusts visible,infrared,and CTD modality weights based on real-time sea state classification using lightweight MobileNetV3 classifiers.Key challenges are identified:1)Scarcity of specialized fishery datasets with extreme weather conditions and rare safety events;2)Systematic violation of fundamental computer vision assumptions——Brightness constancy is disrupted by sea surface specular reflection,temporal consistency is violated by six-degree-of-freedom vessel motion,and input clarity is degraded by sea fog and salt spray;3)Strict computational constraints on fishing vessels limiting model complexity;4)Lack of interpretability in safety-critical deep learning models.Future directions include digital twin-based synthetic data generation,domain adaptation for robust perception across varying sea conditions,and knowledge distillation combined with edge-cloud collaborative inference for model lightweighting on resource-constrained platforms.
周彦翔;程田飞;方辉;戴阳
中国水产科学研究院东海水产研究所,上海 200090中国水产科学研究院东海水产研究所,上海 200090中国水产科学研究院东海水产研究所,上海 200090中国水产科学研究院东海水产研究所,上海 200090
农业科技
船载多源感知系统深度学习目标检测行为识别多模态融合渔业智能化
shipborne multi-source perception systemdeep learningobject detectionaction recognitionmultimodal fusionfishery intelligence
《海洋渔业》 2026 (3)
257-273,17
全球船联网工程体系构建(LSKJ202201800)
评论