TY - GEN
T1 - Large scale data centers simulation based on baseline test model
AU - Lei, Fei
AU - Yu, Lei
AU - Shao, Bing
AU - Teng, Fei
AU - Zhou, Bo
N1 - Publisher Copyright:
© 2018 IEEE.
PY - 2018/8/3
Y1 - 2018/8/3
N2 - With the growing increase of data processing and Hadoop data center construction requirements, simulations of large scale Hadoop data centers with high precision are becoming a great challenge. In this paper, a new simulator which integrates baseline test and multi-layered network model is introduced and implemented. With baseline test, we can predict precisely the execution time for each MapReduce task. The network model considers the complexity of data center network and it can provide more accurate prediction of data transfer delay. The experimental test is implemented in a large scale data center to evaluate the performance of our simulator. Three experimental environments with 35 nodes, 47 nodes and 80 nodes are configured in the data center, and Terasort, Wordcount and Hive are selected as benchmarks with maximum 100 TB input data. The experiments show that the error comparing between the simulator results and experimental environment results in most cases is less than 10%. The comparison with YARNsim confirms that our simulator is capable to achieve precise simulation for a large scale data center and the baseline test model is suitable to optimize the execution time simulation for any type of Hadoop applications.
AB - With the growing increase of data processing and Hadoop data center construction requirements, simulations of large scale Hadoop data centers with high precision are becoming a great challenge. In this paper, a new simulator which integrates baseline test and multi-layered network model is introduced and implemented. With baseline test, we can predict precisely the execution time for each MapReduce task. The network model considers the complexity of data center network and it can provide more accurate prediction of data transfer delay. The experimental test is implemented in a large scale data center to evaluate the performance of our simulator. Three experimental environments with 35 nodes, 47 nodes and 80 nodes are configured in the data center, and Terasort, Wordcount and Hive are selected as benchmarks with maximum 100 TB input data. The experiments show that the error comparing between the simulator results and experimental environment results in most cases is less than 10%. The comparison with YARNsim confirms that our simulator is capable to achieve precise simulation for a large scale data center and the baseline test model is suitable to optimize the execution time simulation for any type of Hadoop applications.
KW - Baseline Test
KW - Hadoop 2 System
KW - Multi layered Network
KW - Simulation
UR - https://www.scopus.com/pages/publications/85052210572
U2 - 10.1109/IPDPSW.2018.00018
DO - 10.1109/IPDPSW.2018.00018
M3 - 会议稿件
AN - SCOPUS:85052210572
SN - 9781538655559
T3 - Proceedings - 2018 IEEE 32nd International Parallel and Distributed Processing Symposium Workshops, IPDPSW 2018
SP - 57
EP - 68
BT - Proceedings - 2018 IEEE 32nd International Parallel and Distributed Processing Symposium Workshops, IPDPSW 2018
PB - Institute of Electrical and Electronics Engineers Inc.
T2 - 32nd IEEE International Parallel and Distributed Processing Symposium Workshops, IPDPSW 2018
Y2 - 21 May 2018 through 25 May 2018
ER -