TY - GEN
T1 - HPC-SFI
T2 - 15th IFIP International Conference on Network and Parallel Computing, NPC 2018
AU - Wang, Yanqi
AU - Zhang, Qi
AU - Liu, Yi
AU - Qian, Depei
N1 - Publisher Copyright:
© IFIP International Federation for Information Processing 2018.
PY - 2018
Y1 - 2018
N2 - Resilience/fault-tolerance has become a key challenge for large-scale parallel systems. To ensure reliability of high performance computing systems, various kinds of techniques have been proposed, such as hardware-level fault-tolerance, checkpointing, replication, algorithm-base fault-tolerance, etc. There are also many software systems to monitor and handle system-failures, e.g. management and job-scheduling system of HPC systems. To evaluate the effectiveness of these systems, it is necessary to provide some kind of tool to inject failures in a HPC system. This paper proposes HPC-SFI, a system-level fault injection tool for HPC systems. Basically, HPC-SFI can generate three kinds of system-failures in a HPC system including in-node faults, failure in the interconnection network and failure of storage/parallel-file system. In addition, HPC-SFI can inject system-faults in pseudo-random model according to pre-defined parameters and probabilities. Preliminary experimental results demonstrate effectiveness of the tool.
AB - Resilience/fault-tolerance has become a key challenge for large-scale parallel systems. To ensure reliability of high performance computing systems, various kinds of techniques have been proposed, such as hardware-level fault-tolerance, checkpointing, replication, algorithm-base fault-tolerance, etc. There are also many software systems to monitor and handle system-failures, e.g. management and job-scheduling system of HPC systems. To evaluate the effectiveness of these systems, it is necessary to provide some kind of tool to inject failures in a HPC system. This paper proposes HPC-SFI, a system-level fault injection tool for HPC systems. Basically, HPC-SFI can generate three kinds of system-failures in a HPC system including in-node faults, failure in the interconnection network and failure of storage/parallel-file system. In addition, HPC-SFI can inject system-faults in pseudo-random model according to pre-defined parameters and probabilities. Preliminary experimental results demonstrate effectiveness of the tool.
UR - https://www.scopus.com/pages/publications/85059683035
U2 - 10.1007/978-3-030-05677-3_9
DO - 10.1007/978-3-030-05677-3_9
M3 - 会议稿件
AN - SCOPUS:85059683035
SN - 9783030056766
T3 - Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics)
SP - 103
EP - 113
BT - Network and Parallel Computing - 15th IFIP WG 10.3 International Conference, NPC 2018, Proceedings
A2 - Snir, Marc
A2 - Valero, Mateo
A2 - Zhang, Feng
A2 - Kasahara, Hironori
A2 - Zhai, Jidong
A2 - Jin, Hai
PB - Springer Verlag
Y2 - 29 November 2018 through 1 December 2018
ER -