首页|期刊导航|煤矿安全|基于深度神经网络的煤矿数据中台元数据特征提取与清洗策略研究

基于深度神经网络的煤矿数据中台元数据特征提取与清洗策略研究OA

Research on metadata feature extraction and cleaning strategy of coal mine data middle platform based on deep neural network

中文摘要英文摘要

为解决煤矿数据中台多源异构数据质量差异大、元数据管理粗放、清洗规则依赖人工且无法动态演化、冷热数据存储与业务特征不匹配等问题,提出了一种基于深度神经网络的元数据特征提取与清洗策略自演化方法:①在传统数据中台嵌入智能清洗与质量反馈层,构建了"采集—清洗—融合—存储—分析—服务"六位一体的全链路治理架构;②设计了卷积神经网络与双向长短期记忆网络融合注意力机制(Convolutional Neural Network-Bidirectional Long Short-Term Memory-Attention Mechanism,CNN-BiLSTM-Attention)的深度学习模型,从字段名称与统计特征双维度提取元数据深层语义,自动生成字段级清洗规则候选集;③建立了 4 级质量标签体系,通过 Q 学习(Q-learning)实现清洗规则权重的动态调整与策略自演化;④构建了数据温度感知的冷热分层存储模型,将热数据、温数据、冷数据分别存储至 Redis、SQL Server、HBase 中;⑤提出了 5 维数据质量指数(Data Quality Index,DQI)与质量报告自动生成框架,形成"清洗—评价—反馈—优化"闭环.在某千万吨级煤矿的应用表明:深度学习模型 F1 值达91.7%,较正则匹配提升了 24.3 个百分点;规则迭代周期由 7 天缩短至 4 h,规则失效平均修复时间由 4.2 h 缩短至 0.5 h;数据质量合格率由 72.6%提升至 94.3%,DQI 由 64.7 升至 91.2;平均查询延迟由 3.8 s 缩短至 1.2 s,存储成本降低了 47%.该方法实现了煤矿数据治理模式的根本转变,即由人工经验驱动转向深度神经网络驱动,由静态规则管控转向自演化智能清洗.

To address the issues of significant quality differences in multi-source heterogeneous data,coarse metadata management,manual dependency and inability to dynamically evolve cleaning rules,and mismatch between cold and hot data storage and busi-ness characteristics in coal mine data middle platforms,a self-evolving metadata feature extraction and cleaning strategy based on deep neural networks is proposed:①embed an intelligent cleaning and quality feedback layer in the traditional data platform to build a six-in-one full-link governance architecture of"collection-cleaning-integration-storage-analysis-service";②design a deep learning model integrating convolutional neural networks,bidirectional long short-term memory networks,and attention mechan-isms(CNN-BiLSTM-Attention)to extract deep semantic features from field names and statistical features,and automatically gener-ate candidate sets of field-level cleaning rules;③establish a four-level quality label system and use Q-learning to dynamically adjust the weights of cleaning rules and achieve self-evolution of strategies;④build a cold and hot data stratified storage model based on data temperature perception,storing hot,warm,and cold data in Redis,SQL Server,and HBase respectively;⑤propose a five-di-mensional data quality index(DQI)and an automatic quality report generation framework to form a"cleaning-evaluation-feed-back-optimization"closed loop.Application in a coal mine with an annual output of over 10 million tons shows that the F1 value of the deep learning model reaches 91.7%,a 24.3 percentage points increase compared to regular matching;the rule iteration cycle is re-duced from 7 days to 4 hours,and the average repair time for rule failure is decreased from 4.2 hours to 0.5 hours;the data quality pass rate increases from 72.6%to 94.3%,and the DQI rises from 64.7 to 91.2;the average query delay is reduced from 3.8 seconds to 1.2 seconds,and storage costs are cut by 47%.This method realizes a fundamental transformation in coal mine data governance from manual experience-driven to deep neural network-driven and from static rules to self-evolving intelligent cleaning.

李翔;芦胜金;阳升;张正清

国家能源集团宁夏煤业有限责任公司,宁夏 银川 750002国家能源集团宁夏煤业有限责任公司,宁夏 银川 750002中煤科工集团重庆研究院有限公司,重庆 400039国能数智科技开发(北京)有限公司,北京 100010

矿业与冶金

煤矿智能化煤矿数据中台元数据特征提取CNN-BiLSTM-Attention清洗策略自演化冷热分层存储数据质量指数

coal mine intellectualizationcoal mine data middle platformmetadata feature extractionCNN-BiLSTM-Attentionself-evolution of cleaning strategiescold and hot stratified storagedata quality index(DQI)

《煤矿安全》 2026 (6)

25-35,11

国家科技重大专项资助项目(2025ZD1700805)天地科技股份有限公司科技创新创业资金专项资助项目(2024-TDZD013-05)

10.13347/j.cnki.mkaq.20260303

评论