追踪提示结合视觉大模型微调的猪只实例分割方法OA
Pig Instance Segmentation Based on Tracking Prompts and Fine-tuned Large Vision Models
针对当前主流猪只感知算法存在追踪与分割任务缺乏协同、依赖高成本人工标注的难题,提出了一种端到端猪只追踪与分割框架.该方法将追踪算法输出的锚框作为视觉大模型 SAM(Segment anything model)的动态提示,以生成时序连贯且身份一致的个体掩模序列.通过结合低秩自适应(Low rank adaptation,LoRA)与猪只多尺度特征适配器(PF Adapter),在小规模数据集上对 SAM 进行高效微调,使该模型仅需少量标注样本,即可在性能上超越传统的有监督学习算法,有效克服了对高昂人工标注的依赖.实验结果表明,微调后的 SAM 模型在两种不同的数据集上的平均像素准确率、交并比和Dice 相似指数分别达到了91.24%、85.06%和86.18%,参数量相较于SAM 基准模型只增加了3%.改进模型输出的猪只体态和运动轨迹信息可用于猪只评估和母猪分娩预警等,为智慧养殖提供技术支撑.
Aiming to address the challenges of lack of coordination between tracking and segmentation tasks and reliance on costly manual annotation in current mainstream pig perception algorithms,a novel end-to-end framework was proposed.Specifically,the proposed method ingeniously leveraged the bounding boxes output by the multi-object tracking algorithm as dynamic spatial prompts for the segment anything model(SAM).This integration facilitated the generation of temporally coherent and identity-consistent individual mask sequences across video frames,bridging the gap between localization and pixel-level segmentation.To adapt the vision foundation model to the specific agricultural domain,the low-rank adaptation(LoRA)was combined with a custom-designed pig multi-scale feature adapter(PF Adapter).The model required only a minimal number of annotated samples to achieve robust feature extraction,ultimately surpassing the performance of traditional fully supervised learning algorithms and effectively overcoming the bottleneck of expensive data annotation.Comprehensive experimental results demonstrated the superiority of the proposed framework.Evaluated across two distinct datasets,the fine-tuned SAM achieved impressive performance metrics,with pixel accuracy(PA)of 91.24%,an intersection over union(IoU)of 85.06%,and a dice similarity coefficient of 86.18%.Remarkably,this significant performance enhancement was achieved with merely a 3%increase in number of parameters compared with the baseline SAM architecture.Furthermore,the pig body posture and movement trajectory information output by the improved model can be used for pig evaluation and sow farrowing warning,providing technical support for smart farming.
孙立博;张泽昀;秦文虎
东南大学仪器科学与工程学院,南京 210096东南大学仪器科学与工程学院,南京 210096东南大学仪器科学与工程学院,南京 210096
信息技术与安全科学
猪只实例分割追踪提示视觉大模型(SAM)LoRAPF Adapter
piginstance segmentationtracking promptslarge vision model(SAM)LoRAPF Adapter
《农业机械学报》 2026 (15)
36-45,10
江苏现代农业产业单项技术研发项目(CX(23)3120)
评论