首页|期刊导航|Computational Visual Media|BoostPoint:Boosting point cloud backbones with image pre-training for 3D understanding

BoostPoint:Boosting point cloud backbones with image pre-training for 3D understandingOA

中文摘要

Nowadays,pre-training models on largescale datasets and fine-tuning models on task-specific datasets have become common paradigms,achieving impressive success in natural language processing and 2D vision.Nonetheless,the potential of this paradigm has not been fully explored in 3D vision due to the scale of the datasets.To overcome this,we propose BoostPoint,a novel pipeline that uses large-scale rendered images as 3D point cloud model inputs for pretraining and uses general 3D tasks for fine-tuning.In BoostPoint,we propose a novel learning-free image-topoint(I2P)module to transform raw pixels into required inputs.Specifically,we view pixels as unorganized points,including essential raw features(e.g.,color)and positional information(e.g.,coordinates).Employing simple linear iterative clustering(SLIC),the I2P module effectively groups these unorganized points into superpixels,facilitating point cloud backbone pretraining.Furthermore,we employ a modality-agnostic debiasing mechanism during pre-training to prevent negative transfer in downstream tasks.Extensive finetuning experiments show that BoostPoint provides significant improvements to 3D point cloud backbones for 3D point cloud classification and part segmentation.

Honggu Zhou;Yakai Zhang;Haohan Li;Xiaoling Gu;Ming Zeng;Zizhao Wu

School of Digital Media Technology,Hangzhou Dianzi University,Hangzhou 310018,ChinaSchool of Digital Media Technology,Hangzhou Dianzi University,Hangzhou 310018,ChinaSchool of Digital Media Technology,Hangzhou Dianzi University,Hangzhou 310018,ChinaSchool of Computer Science,Hangzhou Dianzi University,Hangzhou 310018,ChinaSchool of Informatics,Xiamen University,Xiamen 361005,ChinaSchool of Digital Media Technology,Hangzhou Dianzi University,Hangzhou 310018,China

信息技术与安全科学

3D point cloudtransfer learningcrossmodal learningrepresentation learning

《Computational Visual Media》 2026 (1)

P.141-158,18

supported by the Zhejiang Province Leading Geese Science and Technology Program(No.2025C02155)the National Natural Science Foundation of China(No.62471168).

10.26599/CVM.2025.9450425

评论