[1] D. Q. Ren, "Algorithm level power efficiency optimization for CPU–GPU processing element in data intensive SIMD/SPMD computing," Journal of Parallel and Distributed Computing, vol. 71, no. 2, pp. 245-253, 2011, doi: 10.1016/j.jpdc.2010.10.007.
[2] Y. Oh et al., "Adaptive cooperation of prefetching and warp scheduling on GPUs," IEEE Transactions on Computers, vol. 68, no. 4, pp. 609-616, Apr. 2019, doi: 10.1109/TC.2018.2878671.
[3] T. L. Falch and A. C. Elster, "Machine learning-based auto-tuning for enhanced performance portability of OpenCL applications," Concurrency and Computation: Practice and Experience, vol. 29, no. 8, p. e4029, 2017, doi: 10.1002/cpe.4029.
[4] H. Wu, G. Diamos, J. Wang, S. Cadambi, S. Yalamanchili, and S. Chakradhar, "Optimizing data warehousing applications for GPUs using kernel fusion/fission," 2012 IEEE 26th International Parallel and Distributed Processing Symposium Workshops and PhD Forum, Shanghai, China, 2012, pp. 2433-2442, doi: 10.1109/IPDPSW.2012.300.
[5] M. Korch and T. Werner, "Exploiting limited access distance for kernel fusion across the stages of explicit one-step methods on GPUs," 2018 30th International Symposium on Computer Architecture and High Performance Computing, Lyon, France, 2018, pp. 148-157, doi: 10.1109/CAHPC.2018.8645892.
[6] W. Sun, A. Li, S. Stuijk, and H. Corporaal, "How much can we gain from tensor kernel fusion on GPUs?," IEEE Access, vol. 12, pp. 126135–126144, 2024, doi: 10.1109/ACCESS.2024.3411473.
[7] P. Hijma, S. Heldens, A. Sclocco, B. Van Werkhoven, and H. E. Bal, "Optimization techniques for GPU programming," ACM Computing Surveys, vol. 55, no. 11, pp. 1–81, 2023.
[8] J. Filipovič et al., "Optimizing CUDA code by kernel fusion: application on BLAS," Journal of Supercomputing, vol. 71, pp. 3934–3957, 2015, doi: 10.1007/s11227-015-1483-z.
[9] H. Zhao et al., "Adaptive kernel fusion for improving the GPU utilization while ensuring QoS," IEEE Transactions on Computers, vol. 74, no. 2, pp. 386–400, Feb. 2025, doi: 10.1109/TC.2024.3477995.
[10] M. Wahib and N. Maruyama, "Scalable kernel fusion for memory-bound GPU applications," SC 2014: Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis, New Orleans, LA, USA, 2014, pp. 191–202, doi: 10.1109/SC.2014.21.
[11] G. Wang, "Coordinate strip-mining and kernel fusion to lower power consumption on GPU," 2011 Design, Automation & Test in Europe, Grenoble, France, 2011, pp. 1–4, doi: 10.1109/DATE.2011.5763317.
[12] A. Li, B. Zheng, G. Pekhimenko and F. Long, "Automatic Horizontal Fusion for GPU Kernels," 2022 IEEE/ACM International Symposium on Code Generation and Optimization (CGO), Seoul, Korea, Republic of, 2022, pp. 14-27, doi: 10.1109/CGO53902.2022.9741270.
[13] J. Fousek, J. Filipovič, and M. Madzin, "Automatic fusions of CUDA-GPU kernels for parallel map," ACM SIGARCH Computer Architecture News, vol. 39, no. 4, pp. 98–99, 2011, doi: 10.1145/2082156.2082183.
[14] B. Qiao et al., "Automatic kernel fusion for image processing DSLs," in Proceedings of the 21st International Workshop on Software and Compilers for Embedded Systems, May 2018, pp. 76–85, doi: 10.1145/3207719.3207723.
[15] J. Fukuhara and M. Takimoto, "Automated kernel fusion for GPU based on code motion," in Proceedings of the 23rd ACM SIGPLAN/SIGBED International Conference on Languages, Compilers, and Tools for Embedded Systems, Jun. 2022, pp. 151–161, doi: 10.1145/3519941.3535078.
[16] N. D. Gai, "Highly efficient and accurate deep learning–based classification of MRI contrast on a CPU and GPU," Journal of Digital Imaging, vol. 35, no. 3, pp. 482–495, 2022, doi: 10.1007/s10278-022-00583-1.
[17] J. Lin, J. Liu, E. F. Y. Young, and M. D. F. Wong, "GAMER: GPU-Accelerated Maze Routing," IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems, vol. 42, no. 2, pp. 583–593, Feb. 2023, doi: 10.1109/TCAD.2022.3184281.
[18] A. Riahi, A. Savadi, and M. Naghibzadeh, "Comparison of analytical and ML-based models for predicting CPU–GPU data transfer time," Computing, vol. 102, pp. 2099–2116, 2020, doi: 10.1007/s00607-019-00780-x.
[19] M. Fang, J. Fang, W. Zhang, H. Zhou, J. Liao, and Y. Wang, "Benchmarking the GPU memory at the warp level," Parallel Computing, vol. 71, pp. 23–41, 2018, doi: 10.1016/j.parco.2017.11.003.
[20] D.-H. Kim, "Evaluation of the performance of GPU global memory coalescing," Evaluation, vol. 4, no. 4, pp. 1–5, 2017.
[21] NVIDIA C. "NVIDIA’s Next Generation CUDA Compute Architecture: Fermi", NVIDIA Corp, 2009.
[22] NVIDIA C. "Whitepaper NVIDIA GeForce GTX 1080", NVIDIA Corp, 2016.
[23] NVIDIA C. "Whitepaper NVIDIA TESLA V100 GPU ARCHITECTURE", NVIDIA Corp, 2017.
[24] Huang, Jen-Cheng, et al. "GPUMech: GPU performance modeling technique based on interval analysis.", in 2014 47th Annual IEEE/ACM International Symposium on Microarchitecture, IEEE, pp. 268-279, 2014, doi: 10.1109/MICRO.2014.59.
[25] NVIDIA C. "CUDA C Programming Guide, Version 12.9", NVIDIA Corporation, 2025.
[26] S. Tabik, G. Ortega, and E. M. Garzón, "Performance evaluation of kernel fusion BLAS routines on the GPU: iterative solvers as case study," Journal of Supercomputing, vol. 70, pp. 577–587, 2014, doi: 10.1007/s11227-014-1102-4.
[27] T. P. C. Benchmark and H. T. M., "Standard Specification Revision 1.3.0," Transaction Processing Performance Council (TPC), 1999.
[28] Y. N. Khalid et al., "FusionCL: a machine-learning based approach for OpenCL kernel fusion to increase system performance," Computing, vol. 103, pp. 2171–2202, 2021, doi: 10.1007/s00607-021-00958-2.
[29] S. M. Atif et al., "Multi-Kernel Fusion for RBF Neural Networks," Neural Processing Letters, vol. 55, pp. 1045–1069, 2023, doi: 10.1007/s11063-022-10925-3.
[30] A. Ashari et al., "On optimizing machine learning workloads via kernel fusion," ACM SIGPLAN Notices, vol. 50, no. 8, pp. 173–182, 2015.
[31] S. K. Shekofteh, H. Noori, M. Naghibzadeh, H. S. Yazdi, and H. Fröning, "Metric selection for GPU kernel classification," ACM Transactions on Architecture and Code Optimization, vol. 15, no. 4, pp. 1–27, 2019, doi: 10.1145/3295690.
[32] NVIDIA C. "Profiler Release 12.9", NVIDIA, 2025.
[33] A. Riahi, A. Savadi, and M. Naghibzadeh, "Many-BSP: an analytical performance model for CUDA kernels," Computing, vol. 106, no. 5, pp. 1519–1555, 2024, doi: 10.1007/s00607-023-01255-w.
[34] NVIDIA C. "PARALLEL THREAD EXECUTION ISA v6.5", NVIDIA, 2019.
[35] M. A. Hall, Practical Machine Learning Tools and Techniques, 3rd ed., Morgan Kauffman, 2011.
[36] A. Riahi, A. Savadi, and M. Naghibzadeh, "Performance prediction of regular CUDA programs by machine learning methods," Soft Computing Journal, 2024, doi: 10.22052/scj.2024.252862.1145.