基于国产异构平台的光前哈密顿量对角化优化OA
Optimization of light-front Hamiltonian diagonalization on domestic heterogeneous platform
为解决基矢光前量子化BLFQ 方法在求解大规模稀疏哈密顿矩阵特征值时面临的性能瓶颈,提出了一种基于国产异构计算平台的优化方案.该方案采用CPU与类 GPU加速卡相结合的异构计算架构,并基于 MPI实现进程间通信,将核心计算任务卸载至加速卡执行.通过对 ARPACK库相关接口的移植,使其适用于加速卡环境,设计了实现稀疏矩阵-向量乘法的并行算法.主要改进包括:将原有CPU端线性代数函数替换为加速卡优化函数,利用 NCCL进行加速卡卡间通信,采用线程合并访存和负载均衡等策略优化并行性能.实验结果显示,优化方案在多种规模测试中取得显著加速效果,加速比在3.7~10.3,通信开销显著降低,负载分配更为均衡.所提方案有效提升了BLFQ在强子物理计算中的性能,并为其他应用大规模稀疏矩阵的科学计算提供了有价值的参考.
To address the performance bottleneck encountered by the basis light-front quantization(BLFQ)method when solving the eigenvalue problems of large-scale sparse Hamiltonian matrices,this paper proposes an optimization scheme based on a domestic heterogeneous computing platform.This scheme employs a heterogeneous computing architecture that combines CPUs and GPU-like accelerator cards,utilizing message passing interface(MPI)for inter-process communication,offloading core com-putational tasks to the accelerator cards for execution.By porting the relevant interfaces of the AR-PACK library to make it compatible with the accelerator card environment,designs a parallel algorithm for sparse matrix-vector multiplication.The key improvements include replacing the original CPU-side linear algebra functions with accelerator-optimized functions,utilizing NVIDIA collective communication library(NCCL)for inter-accelerator communication,and adopting strategies such as thread coalescing memory access and load balancing to optimize parallel performance.Experimental results demonstrate that the optimization scheme achieves significant speedups across various test scales,with acceleration ratios ranging from 3.7 to 10.3.Additionally,communication overhead is significantly reduced,and load distribution becomes more balanced.The proposed scheme effectively enhances the performance of BLFQ in hadron physics computations and provides valuable insights for scientific computing applications involving large-scale sparse matrices in other domains.
陶晨博;韩雨萌;武鹏;徐思琦;赵行波;刘杰;戴荣;曹武迪
郑州大学计算机与人工智能学院,河南 郑州 450001郑州大学计算机与人工智能学院,河南 郑州 450001曙光信息产业(北京)有限公司,北京 100193中国科学院近代物理研究所,甘肃 兰州 730000中国科学院近代物理研究所,甘肃 兰州 730000曙光信息产业(北京)有限公司,北京 100193曙光信息产业(北京)有限公司,北京 100193曙光信息产业(北京)有限公司,北京 100193
信息技术与安全科学
基矢光前量子化哈密顿矩阵稀疏矩阵-向量乘法异构计算通信优化负载均衡合并访存
basis light-front quantization(BLFQ)Hamiltonian matrixsparse matrix-vector multi-plication(SpMV)heterogeneous computingcommunication optimizationload balancingcoalesced memory access
《计算机工程与科学》 2026 (6)
997-1007,11
国家自然科学基金(12375143)国家重点研发计划(2021YFB0300200)河南省重大科技专项(221100210600)
评论