首页|期刊导航|福建电脑|差分隐私HADPK-means++聚类算法

差分隐私HADPK-means++聚类算法OA

Differential Privacy HADPK-means++Clustering Algorithm

中文摘要英文摘要

为解决差分隐私k-means聚类算法在迭代过程中因噪声累积导致簇中心偏离,进而影响聚类可用性的问题,本文提出一种高可用性的差分隐私 HADPK-means++算法.该方法通过基于逆序排序的初始簇中心选择以提升初始中心质量,引入结合簇内与簇间相似度的新度量以优化样本划分,并利用差分隐私的变换不变性对加噪后的簇中心进行修正,防止其偏离有效数据范围.在Iris、Wine等多个真实数据集上的实验表明,在相同隐私保护预算下,本算法的F值与标准互信息(NMI)均优于现有主流差分隐私k-means算法.HADPK-means++算法能有效抑制簇中心偏离,提升聚类的可用性与鲁棒性.

To address the issue of cluster center deviation caused by noise accumulation during iterations in differentially private k-means clustering,which severely compromises clustering utility,this paper proposes a highly available differential privacy HADPK-means++algorithm.The method enhances initial center quality through reverse-order-sorting-based initial cluster center selection,optimizes sample partitioning by introducing a new metric combining intra-cluster and inter-cluster similarity,and corrects the noisy cluster centers using the transformation invariance of differential privacy to prevent them from deviating from the valid data range.Experiments on multiple real-world datasets,including Iris and Wine,demonstrate that under the same privacy budget,the proposed algorithm outperforms existing mainstream differentially private k-means algorithms in terms of both F-measure and Normalized Mutual Information(NMI).The HADPK-means++algorithm effectively suppresses cluster center deviation,thereby improving clustering utility and robustness.

徐富国;李磊;陈涛

商丘师范学院信息技术学院 河南 商丘 476000商丘师范学院信息技术学院 河南 商丘 476000商丘师范学院信息技术学院 河南 商丘 476000

信息技术与安全科学

聚类K-means算法差分隐私

ClusteringK-Means AlgorithmDifferential Privacy

《福建电脑》 2026 (2)

7-15,9

本文得到河南省哲学社会科学教育强省研究项目(No.2025JYQS1084)、商丘市哲学社会科学规划项目(商社规办[2025]3号)资助.

10.16707/j.cnki.fjpc.2026.02.002

评论