差分隐私HADPK-means++聚类算法OA
Differential Privacy HADPK-means++Clustering Algorithm
为解决差分隐私k-means聚类算法在迭代过程中因噪声累积导致簇中心偏离,进而影响聚类可用性的问题,本文提出一种高可用性的差分隐私 HADPK-means++算法.该方法通过基于逆序排序的初始簇中心选择以提升初始中心质量,引入结合簇内与簇间相似度的新度量以优化样本划分,并利用差分隐私的变换不变性对加噪后的簇中心进行修正,防止其偏离有效数据范围.在Iris、Wine等多个真实数据集上的实验表明,在相同隐私保护预算下,本算法的F值与标准互信息(NMI)均优于现有主流差分隐私k-means算法.HADPK-means++算法能有效抑制簇中心偏离,提升聚类的可用性与鲁棒性.
To address the issue of cluster center deviation caused by noise accumulation during iterations in differentially private k-means clustering,which severely compromises clustering utility,this paper proposes a highly available differential privacy HADPK-means++algorithm.The method enhances initial center quality through reverse-order-sorting-based initial cluster center selection,optimizes sample partitioning by introducing a new metric combining intra-cluster and inter-cluster similarity,and corrects the noisy cluster centers using the transformation invariance of differential privacy to prevent them from deviating from the valid data range.Experiments on multiple real-world datasets,including Iris and Wine,demonstrate that under the same privacy budget,the proposed algorithm outperforms existing mainstream differentially private k-means algorithms in terms of both F-measure and Normalized Mutual Information(NMI).The HADPK-means++algorithm effectively suppresses cluster center deviation,thereby improving clustering utility and robustness.
徐富国;李磊;陈涛
商丘师范学院信息技术学院 河南 商丘 476000商丘师范学院信息技术学院 河南 商丘 476000商丘师范学院信息技术学院 河南 商丘 476000
信息技术与安全科学
聚类K-means算法差分隐私
ClusteringK-Means AlgorithmDifferential Privacy
《福建电脑》 2026 (2)
7-15,9
本文得到河南省哲学社会科学教育强省研究项目(No.2025JYQS1084)、商丘市哲学社会科学规划项目(商社规办[2025]3号)资助.
评论