TY - GEN
T1 - A practical approach of Curved Ray Prestack Kirchhoff Time Migration on GPGPU
AU - Shi, Xiaohua
AU - Li, Chuang
AU - Wang, Xu
AU - Li, Kang
PY - 2009
Y1 - 2009
N2 - We introduced four prototypes of General Purpose GPU solutions by Compute Unified Device Architecture (CUDA) on NVidia GeForce 8800GT and Tesla C870 for a practical Curved Ray Prestack Kirchhoff Time Migration program, which is one of the most widely adopted imaging methods in the seismic data processing industry. We presented how to re-design and re-implement the original CPU code to efficient GPU code step by step. We demonstrated optimization methods, such as how to reduce the overhead of memory transportation on PCI-E bus, how to significantly increase the kernel thread numbers on GPU cores, how to buffer the inputs and outputs of CUDA kernel modules, and how to utilize the memory streams to overlap GPU kernel execution time, etc., to improve the runtime performance on GPUs. We analyzed the floating point errors between CPUs and GPUs. We presented the images generated by CPU and GPU programs for the same real-world seismic data inputs. Our final approach of Prototype-IV on NVidia GeForce 8800GT is 16.3 times faster than its CPU version on Intel's P4 3.0G.
AB - We introduced four prototypes of General Purpose GPU solutions by Compute Unified Device Architecture (CUDA) on NVidia GeForce 8800GT and Tesla C870 for a practical Curved Ray Prestack Kirchhoff Time Migration program, which is one of the most widely adopted imaging methods in the seismic data processing industry. We presented how to re-design and re-implement the original CPU code to efficient GPU code step by step. We demonstrated optimization methods, such as how to reduce the overhead of memory transportation on PCI-E bus, how to significantly increase the kernel thread numbers on GPU cores, how to buffer the inputs and outputs of CUDA kernel modules, and how to utilize the memory streams to overlap GPU kernel execution time, etc., to improve the runtime performance on GPUs. We analyzed the floating point errors between CPUs and GPUs. We presented the images generated by CPU and GPU programs for the same real-world seismic data inputs. Our final approach of Prototype-IV on NVidia GeForce 8800GT is 16.3 times faster than its CPU version on Intel's P4 3.0G.
KW - CUDA
KW - General purpose GPU
KW - Prestack Kirchhoff Time Migration
UR - https://www.scopus.com/pages/publications/70350637462
U2 - 10.1007/978-3-642-03644-6_13
DO - 10.1007/978-3-642-03644-6_13
M3 - 会议稿件
AN - SCOPUS:70350637462
SN - 3642036430
SN - 9783642036439
T3 - Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics)
SP - 165
EP - 176
BT - Advanced Parallel Processing Technologies - 8th International Symposium, APPT 2009, Proceedings
T2 - 8th International Symposium on Advanced Parallel Processing Technologies, APPT 2009
Y2 - 24 August 2009 through 25 August 2009
ER -