TY - GEN
T1 - Accelerating the Cryo-EM Structure Determination in RELION on Modern Many-Core CPU
AU - Lei, Kelun
AU - Yang, Hailong
AU - Yuan, Jia
AU - Du, Shaokang
AU - Luan, Zhongzhi
AU - Liu, Yi
AU - Qian, Depei
N1 - Publisher Copyright:
© 2025 IEEE.
PY - 2025
Y1 - 2025
N2 - RELION is a widely-used software suite for cryoelectron microscopy (cryo-EM) single-particle analysis (SPA), yet its performance optimization has primarily focused on x86 CPUs and NVIDIA GPUs. In this work, we present the first systematic effort to optimize RELION on modern many-core CPUs. Through detailed performance analysis, we identify critical bottlenecks across RELION's major computational stages. We then apply a set of software- and hardware-aware optimizations, including vectorization optimization, process and thread configurations tuning, algorithm optimization, lock optimization, memory affinity optimization, and computation redundancy optimization. Our optimized version achieves significant speedups and exhibits better scalability than the original RELION across all stages. Notably, it outperforms a single NVIDIA A100 GPU on the complete SPA workflow, achieving a 2.22 × speedup on the SPA dataset and a 1.13 × speedup on the RELION Benchmark dataset. Validation experiments further confirm that our optimizations preserve the reconstruction accuracy, demonstrating the potential of specific CPU architectures as a competitive and efficient platform for cryo-EM data processing.
AB - RELION is a widely-used software suite for cryoelectron microscopy (cryo-EM) single-particle analysis (SPA), yet its performance optimization has primarily focused on x86 CPUs and NVIDIA GPUs. In this work, we present the first systematic effort to optimize RELION on modern many-core CPUs. Through detailed performance analysis, we identify critical bottlenecks across RELION's major computational stages. We then apply a set of software- and hardware-aware optimizations, including vectorization optimization, process and thread configurations tuning, algorithm optimization, lock optimization, memory affinity optimization, and computation redundancy optimization. Our optimized version achieves significant speedups and exhibits better scalability than the original RELION across all stages. Notably, it outperforms a single NVIDIA A100 GPU on the complete SPA workflow, achieving a 2.22 × speedup on the SPA dataset and a 1.13 × speedup on the RELION Benchmark dataset. Validation experiments further confirm that our optimizations preserve the reconstruction accuracy, demonstrating the potential of specific CPU architectures as a competitive and efficient platform for cryo-EM data processing.
KW - Many-core CPU
KW - Optimization
KW - RELION
UR - https://www.scopus.com/pages/publications/105022714626
U2 - 10.1109/HPCC67675.2025.00023
DO - 10.1109/HPCC67675.2025.00023
M3 - 会议稿件
AN - SCOPUS:105022714626
T3 - Proceedings - 2025 27th IEEE International Conference on High Performance Computing and Communications, 11th IEEE International Conference on Data Science and Systems, 23rd IEEE International Conference on Smart City, 11th IEEE International Conference on Dependability in Sensor, Cloud, and Big Data Systems and Applications and 21st IEEE International Conference on Embedded Software and Systems, HPCC/DSS/SmartCity/DependSys/ICESS 2025
SP - 26
EP - 33
BT - Proceedings - 2025 27th IEEE International Conference on High Performance Computing and Communications, 11th IEEE International Conference on Data Science and Systems, 23rd IEEE International Conference on Smart City, 11th IEEE International Conference on Dependability in Sensor, Cloud, and Big Data Systems and Applications and 21st IEEE International Conference on Embedded Software and Systems, HPCC/DSS/SmartCity/DependSys/ICESS 2025
A2 - Hu, Jia
A2 - Min, Geyong
A2 - Wang, Haozhe
A2 - Miao, Wang
A2 - Xu, Lexi
A2 - Georgalas, Nektarios
A2 - Zhao, Zhiwei
A2 - Jin, Rui
A2 - Pang, Guangyao
A2 - Han, Wei
A2 - Hao, Fei
PB - Institute of Electrical and Electronics Engineers Inc.
T2 - 27th IEEE International Conference on High Performance Computing and Communications, HPCC 2025
Y2 - 13 August 2025 through 15 August 2025
ER -