TY - GEN
T1 - Large-Scale Parallelization and Optimization of Lattice QCD on Tianhe New Generation Supercomputer
AU - Chen, Junlin
AU - Liu, Chaojing
AU - Luana, Zhongzhi
AU - Gong, Ming
AU - Li, Qingfeng
AU - Qian, Depei
N1 - Publisher Copyright:
© 2023 IEEE.
PY - 2023
Y1 - 2023
N2 - Lattice quantum chromodynamics (Lattice QCD) is a systematic approach to the study of hadron physics using large-scale numerical simulation methods. Conventional computing platforms have difficulty meeting the needs of large-scale, high-precision computations of Lattice QCD. Tianhe New Generation Supercomputer, which is based on MT-3000 heterogeneous multi-zone processor, provides a new platform for computations of Lattice QCD. However, there are many difficulties in implementing and optimizing large-scale parallel computations of Lattice QCD on a new unique hardware architecture. In this paper, we design a parallel acceleration method for Lattice QCD on the MT-3000 processor, which includes optimization of data distribution, data transmission and vectorization. Finally, we perform performance tests and large-scale parallel experiments on this method. The new method achieves 31 times performance improvement compared to the original serial program and still maintains good performance when eventually using 4521984 cores for large-scale parallelization.
AB - Lattice quantum chromodynamics (Lattice QCD) is a systematic approach to the study of hadron physics using large-scale numerical simulation methods. Conventional computing platforms have difficulty meeting the needs of large-scale, high-precision computations of Lattice QCD. Tianhe New Generation Supercomputer, which is based on MT-3000 heterogeneous multi-zone processor, provides a new platform for computations of Lattice QCD. However, there are many difficulties in implementing and optimizing large-scale parallel computations of Lattice QCD on a new unique hardware architecture. In this paper, we design a parallel acceleration method for Lattice QCD on the MT-3000 processor, which includes optimization of data distribution, data transmission and vectorization. Finally, we perform performance tests and large-scale parallel experiments on this method. The new method achieves 31 times performance improvement compared to the original serial program and still maintains good performance when eventually using 4521984 cores for large-scale parallelization.
KW - Lattice Quantum Chromodynamics
KW - Parallel Computing
KW - Performance Optimization
KW - Tianhe New Generation Supercomputer
UR - https://www.scopus.com/pages/publications/85189862134
U2 - 10.1109/HPCC-DSS-SmartCity-DependSys60770.2023.00074
DO - 10.1109/HPCC-DSS-SmartCity-DependSys60770.2023.00074
M3 - 会议稿件
AN - SCOPUS:85189862134
T3 - Proceedings - 2023 IEEE International Conference on High Performance Computing and Communications, Data Science and Systems, Smart City and Dependability in Sensor, Cloud and Big Data Systems and Application, HPCC/DSS/SmartCity/DependSys 2023
SP - 499
EP - 506
BT - Proceedings - 2023 IEEE International Conference on High Performance Computing and Communications, Data Science and Systems, Smart City and Dependability in Sensor, Cloud and Big Data Systems and Application, HPCC/DSS/SmartCity/DependSys 2023
A2 - Chen, Jinjun
A2 - Yang, Laurence T.
PB - Institute of Electrical and Electronics Engineers Inc.
T2 - 25th IEEE International Conferences on High Performance Computing and Communications, 9th International Conference on Data Science and Systems, 21st IEEE International Conference on Smart City and 9th IEEE International Conference on Dependability in Sensor, Cloud and Big Data Systems and Applications, HPCC/DSS/SmartCity/DependSys 2023
Y2 - 13 December 2023 through 15 December 2023
ER -