TY - JOUR
T1 - Accelerating in-memory transaction processing using general purpose graphics processing units
AU - Gao, Lan
AU - Xu, Yunlong
AU - Wang, Rui
AU - Yang, Hailong
AU - Luan, Zhongzhi
AU - Qian, Depei
N1 - Publisher Copyright:
© 2019 Elsevier B.V.
PY - 2019/8
Y1 - 2019/8
N2 - High throughput is critical for on-line transaction processing (OLTP) applications with a large amount of users. With massive parallel processing units and high memory bandwidth, GPUs are suitable for accelerating OLTP transactions. However, it is challenge to implement transaction execution on GPUs, due to (1) the branch divergences caused by the single instruction multiple threads (SIMT) execution paradigm, and (2) the lack of fine-grained synchronization mechanism and pointer-based dynamic data structures in the GPU ecosystem. In this paper, we present a high-performance in-memory transaction processing system on GPUs to accelerate OLTP applications, named GPU-TPS. Firstly, we propose a transaction execution model to improve GPU hardware utilization and perform synchronization among transactions. Secondly, we optimize the indexing data structures that used extensively in OLTP systems (i.e., hash table for unordered store, and b+ tree for ordered store) for fast storing on GPUs. To evaluate GPU-TPS, we apply it to two popular OLTP workloads (SmallBank and TPCC), and compare it with the state-of-the-art hardware transactional memory based CPU OLTP system (DrTM) and a GPU OLTP system (GPUTx). The experimental results show that GPU-TPS outperforms the CPU implementation by 3.8X for SmallBank and by 1.9X for TPCC, and outperforms the GPU implementation by 1.6X for SmallBank and by 1.8X for TPCC.
AB - High throughput is critical for on-line transaction processing (OLTP) applications with a large amount of users. With massive parallel processing units and high memory bandwidth, GPUs are suitable for accelerating OLTP transactions. However, it is challenge to implement transaction execution on GPUs, due to (1) the branch divergences caused by the single instruction multiple threads (SIMT) execution paradigm, and (2) the lack of fine-grained synchronization mechanism and pointer-based dynamic data structures in the GPU ecosystem. In this paper, we present a high-performance in-memory transaction processing system on GPUs to accelerate OLTP applications, named GPU-TPS. Firstly, we propose a transaction execution model to improve GPU hardware utilization and perform synchronization among transactions. Secondly, we optimize the indexing data structures that used extensively in OLTP systems (i.e., hash table for unordered store, and b+ tree for ordered store) for fast storing on GPUs. To evaluate GPU-TPS, we apply it to two popular OLTP workloads (SmallBank and TPCC), and compare it with the state-of-the-art hardware transactional memory based CPU OLTP system (DrTM) and a GPU OLTP system (GPUTx). The experimental results show that GPU-TPS outperforms the CPU implementation by 3.8X for SmallBank and by 1.9X for TPCC, and outperforms the GPU implementation by 1.6X for SmallBank and by 1.8X for TPCC.
KW - Graphics processing units
KW - Hashing
KW - Online transaction processing
KW - Synchronization
KW - b+ tree
UR - https://www.scopus.com/pages/publications/85063480461
U2 - 10.1016/j.future.2019.03.034
DO - 10.1016/j.future.2019.03.034
M3 - 文章
AN - SCOPUS:85063480461
SN - 0167-739X
VL - 97
SP - 836
EP - 848
JO - Future Generation Computer Systems
JF - Future Generation Computer Systems
ER -