首页|期刊导航|青岛大学学报(自然科学版)|基于细粒度文本特征融合的跨模态单目深度估计模型

基于细粒度文本特征融合的跨模态单目深度估计模型OA

Cross-modal Monocular Depth Estimation Model Based on Fine-grained Textual Feature Fusion

中文摘要英文摘要

复杂场景中,针对人物或者动物等不规则目标的图像进行单目深度估计时,存在深度不连续或者边缘模糊等问题,为此,设计了融合多个细粒度文本特征的单目深度估计(Mul-tiple Fine-Grained Textual Features Fusion for Monocular Depth Estimation,MFGT-Depth)模型,利用高斯混合模型对文本特征空间中的潜在语义层次及其复杂关系进行概率建模,再通过跨模态模块自适应地融合文本特征与图像编码器获得的图像特征,在潜在嵌入空间中实现跨模态对齐.实验结果表明,MFGT-Depth模型在训练数据集上的主要指标得到了显著提升,在KITTI数据集上评估的RMSE指标降低了4.8%,具有很强的泛化能力,特别是对于包含不规则目标的图像,相较于基础模型Depth Anything V2在深度估计的精度和鲁棒性上取得了显著改善.

In the field of monocular depth estimation for complex scenes,a key problem is the occurrence of depth discontinuities and blurred edges caused by irregular objects such as people or animals.To address this issue,a model named Multiple Fine-Grained Textual Features Fusion for Monocular Depth Estimation(MFGT-Depth)was designed.The Gaussian Mixture Model(GMM)was used to probabilistically model latent semantic hier-archy and complex relationships among descriptions.For efficient cross-modal alignment,the cross-modal module then adaptively fused the text features and the image features ex-tracted by the image encoder.Experiments show that MFGT-Depth achieves significant improvements in major metrics on training datasets,reducing RMSE by 4.8%on the KITTI dataset,and exhibits strong generalization with significantly improved accuracy and robustness for depth estimation in scenes with irregular objects relative to Depth Any-thing V2.

谭莉;王文龙;周萌萌;刘小娇;杨杰

青岛大学机电工程学院,青岛 266071青岛大学机电工程学院,青岛 266071青岛联合创智科技有限公司,青岛 266100海逸恒安项目管理有限公司青岛事业部,青岛 250011青岛大学机电工程学院,青岛 266071

信息技术与安全科学

单目深度估计文本特征融合高斯混合模型细粒度跨模态深度学习计算机视觉

monocular depth estimationtextual feature fusionGaussian Mixture Modelfine-grainedcross-modaldeep learningcomputer vision

《青岛大学学报(自然科学版)》 2026 (2)

40-47,8

山东省自然科学基金(批准号:ZR2021MF025)资助.

10.3969/j.issn.1006-1037.2026.02.06

评论