跳到主要导航 跳到搜索 跳到主要内容

Re-running large-scale parallel programs using two nodes

  • Beihang University

科研成果: 书/报告/会议事项章节会议稿件同行评审

摘要

With the increasing of scale and complexity of high performance computing (HPC) systems, the programming, debugging, and tuning of large-scale parallel programs face a series of challenges, one of which is that programmers often need to repeatedly run their programs with large number of processes on HPC systems to identify sources of errors and performance bottlenecks in their programs, which means large amounts of resource consumptions. Furthermore, since most HPC systems use job scheduling system to manage their resources and schedule multiple jobs from different users, programmers cannot interact with their programs during the execution of programs, which further increases complexities of debugging and tuning. To address this challenge, this paper proposes a system that re-runs large-scale MPI parallel programs using two nodes. According to an approach of one real-execution + multiple emulation-executions, the parallel program is firstly executed with desired number of processes on an HPC system, which is referred as real-execution, and during the execution, our system records MPI messages transmitted among processes as well as control information of processes; after that, one or more processes can be re-run on a two-node local system under the scale the same with the real-execution. In the meantime, programmers can interact with their programs by attaching the GDB, a commonly used debugger, to the re-running process. Therefore, not only can our system reduce resource-consumptions in debugging and tuning of large-scale parallel programs significantly, but also support interactions between developers and their programs during the execution of the programs, which makes programmers easier to identify sources of the errors and performance bottlenecks in their parallel programs.

源语言英语
主期刊名Proceedings - 16th IEEE International Symposium on Parallel and Distributed Processing with Applications, 17th IEEE International Conference on Ubiquitous Computing and Communications, 8th IEEE International Conference on Big Data and Cloud Computing, 11th IEEE International Conference on Social Computing and Networking and 8th IEEE International Conference on Sustainable Computing and Communications, ISPA/IUCC/BDCloud/SocialCom/SustainCom 2018
编辑Jinjun Chen, Laurence T. Yang
出版商Institute of Electrical and Electronics Engineers Inc.
485-492
页数8
ISBN(电子版)9781728111414
DOI
出版状态已出版 - 2 7月 2018
活动16th IEEE International Symposium on Parallel and Distributed Processing with Applications, 17th IEEE International Conference on Ubiquitous Computing and Communications, 8th IEEE International Conference on Big Data and Cloud Computing, 11th IEEE International Conference on Social Computing and Networking and 8th IEEE International Conference on Sustainable Computing and Communications, ISPA/IUCC/BDCloud/SocialCom/SustainCom 2018 - Melbourne, 澳大利亚
期限: 11 12月 201813 12月 2018

出版系列

姓名Proceedings - 16th IEEE International Symposium on Parallel and Distributed Processing with Applications, 17th IEEE International Conference on Ubiquitous Computing and Communications, 8th IEEE International Conference on Big Data and Cloud Computing, 11th IEEE International Conference on Social Computing and Networking and 8th IEEE International Conference on Sustainable Computing and Communications, ISPA/IUCC/BDCloud/SocialCom/SustainCom 2018

会议

会议16th IEEE International Symposium on Parallel and Distributed Processing with Applications, 17th IEEE International Conference on Ubiquitous Computing and Communications, 8th IEEE International Conference on Big Data and Cloud Computing, 11th IEEE International Conference on Social Computing and Networking and 8th IEEE International Conference on Sustainable Computing and Communications, ISPA/IUCC/BDCloud/SocialCom/SustainCom 2018
国家/地区澳大利亚
Melbourne
时期11/12/1813/12/18

学术指纹

探究 'Re-running large-scale parallel programs using two nodes' 的科研主题。它们共同构成独一无二的学术指纹。

引用此