首页|期刊导航|沈阳工业大学学报|高维稀疏电力负荷数据无监督挖掘算法

高维稀疏电力负荷数据无监督挖掘算法OA

Unsupervised mining algorithm for high-dimensional sparse power load data

中文摘要英文摘要

[目的]在电力系统中,负荷数据分析对电网调度、规划和管理至关重要.然而,随着电力系统的复杂化与智能化程度的加深,电力负荷数据呈现高维度、稀疏性等特点,导致传统数据分析方法在处理效率和捕捉负荷变化内在信息方面面临较大挑战.本文提出一种高效无监督数据挖掘算法,旨在提升高维稀疏电力负荷数据的处理效率与信息提取能力.[方法]首先,采用基于信息熵的特征排序法确定特征重要度.通过计算互信息、开展中心化、标准化处理等完成数据初始化,选择互信息最大的特征扩充特征集合,通过计算相关信息熵筛选特征子集,以支持向量机(SVM)分类器为基准模型优化子集筛选过程,引入改进粒子群算法进行特征二次选择,同时借助SVM分类器完成特征初步筛选.然后,引入主成分分析(PCA)实施降维.对样本矩阵进行中心化处理,构建协方差矩阵,获取特征值与特征向量,选择特征向量构建新的矩阵以实现降维.最后,引入基于无监督学习的自编码网络开展数据无监督挖掘.编码阶段将输入数据转化为特征表示,解码阶段完成数据恢复,通过设定数据、执行聚类操作、筛选数据点、开展数据均衡处理、获取训练模型分类界面等步骤,实现隐藏特征提取与网络调节.[结果]本文算法在整个测试过程中兰德指数一直大于0.60,呈现较高的聚类准确性.在60 次迭代实验中,最大内存开销占比约为8.3%,表明本文算法的计算资源利用率较高.与其他传统算法相比,本文算法在处理高维稀疏电力负荷数据时,能够表现出更高的处理效率和更优的挖掘效果.[结论]无监督挖掘算法在高维稀疏电力负荷数据分析中表现优异,本文算法通过特征选择与降维处理减少计算量,并借助自编码网络挖掘非线性特征,显著提升了数据挖掘的准确性与效率,具有很强的适用性与可行性.本文算法的创新之处在于,融合信息熵特征排序、支持向量机、改进粒子群、主成分分析与自编码网络等多种方法,从特征处理到数据挖掘形成完整体系,既能有效应对高维稀疏电力负荷数据的挖掘难题,又为电力系统负荷数据分析提供了新的有效手段,因而对推动电力系统智能化发展具有重要意义.

[Objective]In power systems,load data analysis is crucial for power grid dispatching,planning,and management.However,with the deepening of the complexity and intelligence of power systems,power load data exhibit characteristics of high dimensionality and sparsity,which poses significant challenges to traditional data analysis methods in terms of processing efficiency and ability to capture the intrinsic information of load changes.An efficient unsupervised data mining algorithm was proposed in this paper,which was aimed at improving the processing efficiency and information extraction capability of high-dimensional sparse power load data.[Methods]Firstly,a feature ranking method based on information entropy was adopted to determine feature importance.Data initialization was completed by calculating mutual information and conducting centralization and standardization.Features with the maximum mutual information were selected to expand the feature set,and feature subsets were screened by calculating relevant information entropy.The subset screening process was optimized using support vector machine(SVM)classifier as the benchmark model,and an improved particle swarm optimization algorithm was introduced for secondary feature selection.Meanwhile,the SVM classifier was used to complete the preliminary feature screening.Secondly,the principal component analysis(PCA)was introduced for dimensionality reduction.The sample matrix was centralized,and the covariance matrix was established.Eigenvalues and eigenvectors were obtained,and eigenvectors were selected to construct a new matrix to achieve dimensionality reduction.Finally,an autoencoder network based on unsupervised learning was introduced to conduct unsupervised mining.In the encoding stage,input data were converted into feature representations.In the decoding stage,data recovery was completed.Through steps such as data setting,clustering execution,data point screening,data balancing processing,and model training to obtain a classification interface,hidden feature extraction and network adjustment were realized.[Results]When the algorithm in this paper is applied,the Rand index values all exceed 0.60,indicating high clustering accuracy.In 60 iterations of experiments,the maximum memory overhead ratio is about 8.3%,demonstrating the algorithm's high efficiency in computing resource utilization.Compared with other traditional methods,this algorithm can achieve higher processing efficiency and better mining results when dealing with high-dimensional sparse power load data.[Conclusions]The unsupervised mining algorithm performs excellently in the analysis of high-dimensional sparse power load data.By reducing computational complexity through feature selection and dimensionality reduction,and mining nonlinear features with the autoencoder network,it significantly improves the accuracy and efficiency of data mining,and has strong applicability and feasibility.Its innovation lies in integrating multiple methods such as information entropy-based feature ranking,SVM,improved particle swarm optimization,PCA,and autoencoder network to form a complete system from feature processing to data mining.This system can not only effectively address the challenges in mining high-dimensional sparse power load data but also provide a new and effective means for load data analysis in power systems,which is of great significance for promoting the intelligent development of power systems.

丁业豪;杨月;马保全

华南理工大学电子与信息学院,广东 广州 510640||广东电网有限责任公司 广东电网能源投资有限公司,广东 广州 510308广东电网有限责任公司 广东电网能源投资有限公司,广东 广州 510308广东电网有限责任公司 清远供电局,广东 广州 510308

信息技术与安全科学

电力负荷数据特征选择与降维自编码网络无监督挖掘主成分分析改进粒子群算法支持向量机

power load datafeature selection and dimensionality reductionautoencoder networkunsupervised miningprincipal component analysisimproved particle swarm optimization algorithmsupport vector machine

《沈阳工业大学学报》 2026 (2)

57-64,8

广东省基础与应用基础研究基金项目(2023A1515011598)南方电网公司科技项目(GDKJXM20200507,031800KK52200004).

评论