首页|期刊导航|数据与计算发展前沿|基于深度学习的Lustre文件系统异常行为智能检测方法

基于深度学习的Lustre文件系统异常行为智能检测方法OA

A Deep Learning-Based Intelligent Anomaly Detection Method for the Lustre File System

中文摘要英文摘要

[背景]Lustre文件系统是科学计算中的重要基座之一,随着存储容量和用户业务量不断扩大,存储系统的负载持续升高.单一用户的异常读写请求往往容易导致存储系统的卡顿,影响其余用户的使用体验.传统的异常处理方式是由运维人员浏览日志,排查定位异常用户或异常读写请求,最后实施故障排除策略恢复系统访问速度.[目的]为提升异常诊断效率,替代依赖运维人员人工筛查日志的低效模式,本研究将深度学习引入运维过程,以实现对异常读写行为的智能诊断.[方法]本研究设计了面向Lustre文件系统的异常行为智能检测系统,涉及用户行为数据采集、时序数据处理、模型搭建及训练和部署验证等过程.该系统将用户读写信息转换为时间序列数据,通过搭建长短时记忆网络,使用无监督学习的方式训练模型.[结论]本研究分别在Lustre的MDT和OST数据上完成了模型的训练和验证.实验结果表明,所提方法能够显著提升异常识别精确率,并有效降低异常误报率.本文提出的方法能够降低运维人员的错误定位时间成本,提升科学计算环境中文件存储系统的异常处理效率.

[Background]The Lustre file system is a crucial foundation for scientific computing.With the continuous expansion of storage capacity and user workload,the load on storage systems is con-stantly increasing.Abnormal read and write requests from a single user often lead to storage cluster lag,affecting the overall user experience.Traditional anomaly handling typically in-volves maintainers browsing logs,identifying and locating abnormal users or read/write re-quests,and finally implementing troubleshooting strategies to restore system access speed.[Purpose]To improve the efficiency of anomaly diagnosis and replace the inefficient mode of relying on manual screening of logs by maintainers,this study introduces deep learning into the operation and maintenance process to achieve intelligent diagnosis of abnormal read and write behavior.[Method]This study builds an intelligent abnormal behavior de-tection system for Lustre file systems,involving user behavior data collection,time-series data processing,model construction,training,and deployment verification.The system converts user read and write information into time series data,builds a long short-term memory network and trains the model using unsupervised learning.[Conclusions]This study successfully trained and validated the model on Lustre's MDT and OST data.Experi-mental results show that the proposed method can significantly improve the accuracy of anomaly detection and ef-fectively reduce the false alarm rate.The proposed method can reduce the time cost for error localization for maintainers and improve the efficiency of anomaly handling in file systems within scientific computing environ-ments.

侯思琦;程垚松;程耀东;李海波;毕玉江;姚秋玲

中国科学院高能物理研究所,北京 100049中国科学院高能物理研究所,北京 100049中国科学院高能物理研究所,北京 100049||中国科学院大学,北京 100049中国科学院高能物理研究所,北京 100049||中国科学院大学,北京 100049中国科学院高能物理研究所,北京 100049中国科学院高能物理研究所,北京 100049

智能运维存储系统时序数据处理

AIOpsstorage systemtime-series data processing

《数据与计算发展前沿》 2026 (3)

29-39,11

国家重点研发计划2023YFC2206401原初引力波望远镜智能数据传输与运行监控系统浪潮存储青蓝基金广域网分布式文件系统与基于SPDK高性能文件系统中国科学院青年创新促进会(2023013)

10.11871/jfdc.issn.2096-742X.2026.03.003

评论