摘要
Reduction is one of the most commonly used collective communication operations for parallel applications. There are two problems for the existing reduction algorithms: First, they cannot adapt to complex environment. When interferences appear in computing environment, the efficiency of reduction degrades significantly. Second, they are not fault tolerant. The reduction operation is interrupted when a node failure occurs. To solve these problems, this paper proposes a task-based parallel high-performance distributed reduction framework. Firstly, each reduction operation is divided into a series of independent computing tasks. The task scheduler is adopted to guarantee that ready tasks will take precedence in execution and each task will be scheduled to the computing node with better performance. Thus, the side effect of slow nodes on the whole efficiency can be reduced. Secondly, based on the reliability storage for reduction data and fault detecting mechanism, fault tolerance can be implemented in tasks without stopping the application. The experimental results in complex environment show that the distributed reduction framework promises high availability and, compared with the existing reduction algorithm, the reduction performance and concurrent reduction performance of distributed reduction framework are improved by 2.2 times and 4 times, respectively.
| 投稿的翻译标题 | A fault tolerant high-performance reduction framework in complex environment |
|---|---|
| 源语言 | 繁体中文 |
| 页(从-至) | 2115-2124 |
| 页数 | 10 |
| 期刊 | Beijing Hangkong Hangtian Daxue Xuebao/Journal of Beijing University of Aeronautics and Astronautics |
| 卷 | 44 |
| 期 | 10 |
| DOI | |
| 出版状态 | 已出版 - 1 10月 2018 |
关键词
- Collective communication
- Complex environment
- Fault tolerance
- Interference
- Parallel computing
- Reduction
指纹
探究 '一种在复杂环境中支持容错的高性能规约框架' 的科研主题。它们共同构成独一无二的指纹。引用此
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver