基于相互近邻和证据理论的密度峰值聚类算法OA
Density peaks clustering algorithm based on mutual nearest neighbors and Dempster-Shafer theory
密度峰值聚类(Density Peaks Clustering,DPC)是一种基于密度的聚类算法,能够在无需预设聚类数量的条件下自动识别任意形状的类簇.然而,DPC算法在处理具有密度差异的类簇时,会在密集类簇中误选多个聚类中心,而忽略稀疏类簇的聚类中心.此外,DPC算法的单步链式分配策略易引发"多米诺效应",即单个样本的错误分配导致后续样本的分配出现连锁偏差.针对上述问题,提出一种基于相互近邻和证据理论的密度峰值聚类算法(Density Peaks Clustering Algorithm Based on Mutual Nearest Neighbors and Dempster-Shafer Theory,MDS-DPC).首先,融合样本的距离相似性与邻域凝聚度,重新定义局部密度的计算方式,有效平衡簇间密度差异.其次,综合考虑样本局部和全局分布特征,基于局部密度峰与相互近邻图优化相对距离度量,更准确地表征样本间的相对位置关系.最后,引入证据理论,采用多阶段分配和跨簇连接样本分层微调策略,提高样本分配的准确性.在九个合成数据集与12个真实数据集上,将所提算法与六个优秀的聚类算法进行对比,实验结果表明,MDS-DPC算法的聚类效果更优.
Density Peaks Clustering(DPC)is a density-based clustering algorithm capable of automatically identifying clusters of arbitrary shapes without requiring the number of clusters to be pre-specified.However,when processing datasets containing clusters with significant density variations,DPC tends to erroneously select multiple cluster centers within dense clusters while overlooking the true centers in sparse clusters.Furthermore,DPC's single-step chained allocation strategy is prone to a"domino effect",where the misallocation of a single data point can trigger a cascade of erroneous assignments for subsequent points.To address these issues,this paper proposes a novel Density Peaks Clustering algorithm based on Mutual Nearest Neighbors and Dempster-Shafer Theory,termed MDS-DPC.First,by integrating a sample's distance-based similarity with its neighborhood cohesion,we redefine the calculation of local density to effectively mitigate inter-cluster density disparities.Second,to more accurately characterize the relative positional relationships between samples,we refine the relative distance metric by leveraging both local and global distribution characteristics,specifically through the integration of local density peaks and a mutual nearest neighbor graph.Finally,we introduce Dempster-Shafer theory and employ a multi-stage assignment strategy coupled with a hierarchical fine-tuning mechanism for cross-cluster connected samples to enhance the overall accuracy of sample allocation.We evaluated the proposed MDS-DPC algorithm against six state-of-the-art clustering algorithms on nine synthetic and twelve real-world datasets.The experimental results demonstrate that MDS-DPC achieves superior clustering performance.
张巳杨;张清华;周新然;邓偲;程云龙
重庆邮电大学计算机科学与技术学院,重庆,400065||重庆邮电大学计算智能重庆市重点实验室,重庆,400065||网络空间大数据智能安全教育部重点实验室,重庆,400065重庆邮电大学计算智能重庆市重点实验室,重庆,400065||网络空间大数据智能安全教育部重点实验室,重庆,400065重庆邮电大学计算机科学与技术学院,重庆,400065||重庆邮电大学计算智能重庆市重点实验室,重庆,400065||网络空间大数据智能安全教育部重点实验室,重庆,400065重庆邮电大学计算机科学与技术学院,重庆,400065||重庆邮电大学计算智能重庆市重点实验室,重庆,400065||网络空间大数据智能安全教育部重点实验室,重庆,400065重庆邮电大学计算智能重庆市重点实验室,重庆,400065
信息技术与安全科学
密度峰值聚类相互近邻局部密度峰证据理论多阶段分配
density peaks clusteringmutual nearest neighborslocal density peaksDempster-Shafer theorymulti-stage assignment
《南京大学学报(自然科学版)》 2026 (4)
592-606,15
国家重点研发计划(2026YFE0201300),国家自然科学基金(62576056),重庆市自然科学基金创新发展联合基金(CSTB2023-NSCQ-LZX0164),重庆市教委科学技术研究(KJZD-K202300613)
评论