TY - GEN
T1 - Event-driven fault tolerance for building nonstop active message programs
AU - Li, Chao
AU - Zhao, Changhai
AU - Yan, Haihua
AU - Zhang, Jianlei
PY - 2014
Y1 - 2014
N2 - With the decreasing Mean Time Between Failures (MTBF) of high performance computing systems, process failure has become a normal phenomenon rather than an exception. The failures in high frequency lead to fault tolerance, a key feature of high performance applications. To provide fault tolerance interfaces for active message programs, this paper proposes a novel model called event-driven fault tolerance. The model converts each detected process failure into an event containing detailed failure information of the execution context, and then schedules the event up to application layer by executing user-directed event handlers to drive the program to recover from faults. Based on events, the model can provide applications with dynamic process groups and fault tolerant communication interfaces. We present an implementation of the model called EDFT (Event Driven Fault Tolerance) and describe its architecture, principle, components and application programming interfaces (API). To evaluate this model, we incorporate EDFT into a scientific application, PreStack Depth Migration (PSDM). Experiments are conducted by injecting various kinds of faults into PSDM when it is running. Experimental results show that for active message programs that demand high performance, event-driven fault tolerance model promises strong robustness, low overhead and high scalability.
AB - With the decreasing Mean Time Between Failures (MTBF) of high performance computing systems, process failure has become a normal phenomenon rather than an exception. The failures in high frequency lead to fault tolerance, a key feature of high performance applications. To provide fault tolerance interfaces for active message programs, this paper proposes a novel model called event-driven fault tolerance. The model converts each detected process failure into an event containing detailed failure information of the execution context, and then schedules the event up to application layer by executing user-directed event handlers to drive the program to recover from faults. Based on events, the model can provide applications with dynamic process groups and fault tolerant communication interfaces. We present an implementation of the model called EDFT (Event Driven Fault Tolerance) and describe its architecture, principle, components and application programming interfaces (API). To evaluate this model, we incorporate EDFT into a scientific application, PreStack Depth Migration (PSDM). Experiments are conducted by injecting various kinds of faults into PSDM when it is running. Experimental results show that for active message programs that demand high performance, event-driven fault tolerance model promises strong robustness, low overhead and high scalability.
KW - active messages
KW - event-driven
KW - fault tolerance
KW - nonstop programs
KW - reliability
UR - https://www.scopus.com/pages/publications/84903995177
U2 - 10.1109/HPCC.and.EUC.2013.62
DO - 10.1109/HPCC.and.EUC.2013.62
M3 - 会议稿件
AN - SCOPUS:84903995177
SN - 9780769550886
T3 - Proceedings - 2013 IEEE International Conference on High Performance Computing and Communications, HPCC 2013 and 2013 IEEE International Conference on Embedded and Ubiquitous Computing, EUC 2013
SP - 382
EP - 390
BT - Proceedings - 2013 IEEE International Conference on High Performance Computing and Communications, HPCC 2013 and 2013 IEEE International Conference on Embedded and Ubiquitous Computing, EUC 2013
PB - IEEE Computer Society
T2 - 15th IEEE International Conference on High Performance Computing and Communications, HPCC 2013 and 11th IEEE/IFIP International Conference on Embedded and Ubiquitous Computing, EUC 2013
Y2 - 13 November 2013 through 15 November 2013
ER -