Benchmark-Driven GPU Performance Optimization for Medical Imaging, Genomics, and Large-Scale AI Workloads
Main Article Content
Abstract
This paper illustrates how benchmarking is useful in achieving optimization of the workloads that can be accelerated using GPUs in clinical imaging, genomics studies, and generative AI training. We tested High-Performance Linpack (HPL) tuning, memory throughput optimization, NCCL communication optimization and GPU health validation to clusters of multiple GPUs. Peak floating-point performance of 12.3 TFLOPS to 34.7 TFLOPS was attained in various GPUs. Memory optimizations boosted performance in effective bandwidth up to 1.8-2.2x. In distributed AI workloads NCCL optimization helped to cut communication latency by 35- 42, and memory virtualization trained large models, including VGG-16 (batch size 256), with only 18 percent loss in performance on 12 GB of a GPU. In medical imaging, when 2.1 -3.3 times less time was spent on reconstruction, there was no quality loss, and this was due to the use of the GPU. Genomics processes were almost 166X faster in identifying microRNAs than on a CPU. These findings demonstrate that the optimizations through benchmarking can lead to a reduction in the time-to-diagnosis, training, and cluster utilization in healthcare and AI
Article Details
Section
How to Cite
References
[1] Srikanth, A., Trigila, C., & Roncali, E. (2024). GPU optimization techniques to accelerate optiGAN— a particle simulation GAN. Machine Learning Science and Technology, 5(2), 027001. https://doi.org/10.1088/2632-2153/ad51c9
[2] Mittal, S., & Vaishay, S. (2019). A survey of techniques for optimizing deep learning on GPUs. Journal of Systems Architecture, 99, 101635. https://doi.org/10.1016/j.sysarc.2019.101635
[3] Das, R. S., & Gupta, V. (2024). A Systematic Literature Review on Graphics Processing Unit Accelerated Realm of High-Performance Computing. International Journal of Computing and Engineering, 5(3), 10–21. https://doi.org/10.47941/ijce.1813
[4] Hijma, P., Heldens, S., Sclocco, A., Van Werkhoven, B., & Bal, H. E. (2022). Optimization techniques for GPU programming. ACM Computing Surveys, 55(11), 1–81. https://doi.org/10.1145/3570638
[5] Lee, S., & Lee, J. (2024). Collective Communication Performance Evaluation for distributed Deep learning training. Applied Sciences, 14(12), 5100. https://doi.org/10.3390/app14125100
[6] Hidayetoglu, M., Garcia, D. G. S., Slaughter, E., Surana, P., Hwu, W., Gropp, W., & Aiken, A. (2024). HICCL: a Hierarchical Collective Communication Library. arXiv (Cornell University). https://doi.org/10.48550/arxiv.2408.05962
[7] Cámara, J., Cuenca, J., Galindo, V., Vicente, A., & Boratto, M. (2024). An autotuning approach to select the inter-GPU communication library on heterogeneous systems. The Journal of Supercomputing, 81(1). https://doi.org/10.1007/s11227-024-06794-3
[8] Li, T., Narayana, V. K., & El-Ghazawi, T. (2015). Efficient resource sharing through GPU virtualization on accelerated high performance computing systems. arXiv (Cornell University). https://doi.org/10.48550/arxiv.1511.07658
[9] Zhou, K., Adhianto, L., Anderson, J., Cherian, A., Grubisic, D., Krentel, M., Liu, Y., Meng, X., & Mellor-Crummey, J. (2021). Measurement and analysis of GPU-accelerated applications with HPCToolkit. Parallel Computing, 108, 102837. https://doi.org/10.1016/j.parco.2021.102837
[10] Wang, P., & Yu, Z. (2023). RayBench: an advanced NVIDIA-Centric GPU rendering benchmark suite for optimal performance analysis. Electronics, 12(19), 4124. https://doi.org/10.3390/electronics12194124
[11] Liu, Z., Zhang, S., Garrigus, J., & Zhao, H. (2023). Genomics-GPU: A Benchmark Suite for GPU- accelerated Genome Analysis. Genomics-GPU: A Benchmark Suite for GPU-accelerated Genome Analysis, 178–188. https://doi.org/10.1109/ispass57527.2023.00026
[12] HPC-AI benchmarks - A comparative overview of high-performance computing hardware and AI benchmarks across domains. (n.d.). https://joaiar.org/articles/AIR-1017.html
[13] Madougou, S., Varbanescu, A., De Laat, C., & Van Nieuwpoort, R. (2016). The landscape of GPGPU performance modeling tools. Parallel Computing, 56, 18–
33. https://doi.org/10.1016/j.parco.2016.04.002
[14] Després, P., & Jia, X. (2017). A review of GPU-based medical image reconstruction. Physica Medica, 42, 76–92. https://doi.org/10.1016/j.ejmp.2017.07.024
[15] Cavicchioli, R., Hu, J. C., Piccolomini, E. L., Morotti, E., & Zanni, L. (2020). GPU acceleration of a model-based iterative method for Digital Breast Tomosynthesis. Scientific Reports, 10(1), 43. https://doi.org/10.1038/s41598-019-56920-y
[16] Yang, S., Zhou, J., Guo, H., Wang, L., & Xu, M. (2024). GPU-accelerated OCT imaging: Real-time data processing and artifact suppression for enhanced monitoring of 3D bioprinted tissues and vascular-like networks. Journal of Innovative Optical Health Sciences, 17(06). https://doi.org/10.1142/s1793545824500135
[17] Chen, K., Wang, C., Xiong, J., & Xie, Y. (2018). GPU based parallel acceleration for fast C-arm cone- beam CT reconstruction. BioMedical Engineering OnLine, 17(1), 73. https://doi.org/10.1186/s12938-018-0506-4
[18] Wang, H., Peng, H., Chang, Y., & Liang, D. (2018). A survey of GPU-based acceleration techniques in MRI reconstructions. Quantitative Imaging in Medicine and Surgery, 8(2),196–208. https://doi.org/10.21037/qims.2018.03.07
[19] Price, D. C., Clark, M. A., Barsdell, B. R., Babich, R., & Greenhill, L. J. (2015). Optimizing performance-per-watt on GPUs in high performance computing. Computer Science - Research and Development, 31(4), 185–193. https://doi.org/10.1007/s00450-015-0300-5
[20] Rhu, M., Gimelshein, N., Clemons, J., Zulfiqar, A., & Keckler, S. W. (2016). VDNN: Virtualized Deep Neural Networks for Scalable, Memory-Efficient Neural Network Design. arXiv (Cornell University). https://doi.org/10.48550/arxiv.1602.08124
[21] Zheng, X., Jin, J., Wang, Y., Yuan, M., & Qiang, S. (2023). Research on the application and performance optimization of GPU parallel computing in concrete temperature control simulation. Buildings, 13(10), 2657. https://doi.org/10.3390/buildings13102657